API integration fails in predictable ways, and the fix is always the same shape: you cannot guarantee a message is delivered exactly once, so you make processing exactly once instead. Idempotency handles duplicates, retries handle minutes of failure, replay handles hours, and reconciliation handles everything else. An integration missing any of those four will eventually produce a double charge, a lost order, or two systems quietly disagreeing.
This guide covers the failure modes, the patterns that contain them, how to price an integration honestly, and what to define before any code is written.
Key takeaways
- Assume at-least-once delivery: Providers resend; your handler must produce the same result on the second attempt as the first.
- Return fast, process later: Accept the payload, acknowledge it, and do the work in a background worker.
- Build reconciliation before you need it: A daily comparison against the source catches what retries never will.
- Price each integration separately: A single line item for "integrations" is the most reliable predictor of an overrun.
Define these four things before writing code
Source of truth. For every field that exists on both sides, which system wins when they disagree. Without this written down, the answer becomes whichever code ran last.
Direction and trigger. Push or pull, real time or batch, and what initiates it. Most integrations that "stopped working" were scheduled jobs nobody was monitoring.
Failure behaviour. What the product does when the other system is slow, unavailable, or returns something unexpected. This is a business decision: do you block the user, queue the work, or proceed and reconcile later?
Owner. A named person on each side who is called when it breaks, and a runbook that does not require the original developer.
The failure modes, and what each needs
| Failure | What it looks like | The fix |
|---|---|---|
| Duplicate delivery | Same event processed twice; double charge or duplicate record | Idempotency key stored and checked before any write |
| Transient outage | 500s or timeouts for a few minutes | Exponential backoff with jitter, capped attempts |
| Rate limiting | 429 responses under load | Honour backpressure; slow down rather than retrying harder |
| Schema change | Fields renamed or removed without notice | Contract tests; validate on receipt and alert rather than silently drop |
| Silent divergence | Both systems working, totals disagree | Scheduled reconciliation against the source |
| Permanent failure | Retries exhausted | Dead-letter queue with an owner and a review process |
The most common design mistake is what practitioners call policy flattening: collapsing every error into one retry path. A quota error and a transient server error need opposite responses, and treating them identically turns a rate limit into a retry storm.
Make every write idempotent
Providers guarantee at-least-once delivery, which means duplicates are normal operation rather than an incident. Every database write triggered by an incoming event must produce the same result when it runs twice with the same event identifier.
In practice: store the event ID on receipt, check it before processing, and make the check and the write atomic. Without that, one redelivered webhook becomes a duplicate order, a second charge, or a second confirmation email to a customer who is now calling support.
Test it deliberately. Most providers offer a replay feature, send the same event three times and confirm the outcome is identical. A test suite that only covers first delivery will pass on the day duplicates start arriving.
Accept fast, process in the background
The pattern that holds up under load is: receive the payload, write it to a queue, return a success response immediately, and process from the queue in a worker.
Doing the work inline before responding couples your processing time to the provider's timeout. Slow processing then looks like a failure to the sender, which triggers a retry, which arrives while the first is still running, and now you have a concurrency problem on top of a performance one.
The queue also gives you replay. When a bug corrupts a day of processing, you reprocess from stored payloads rather than asking the provider to resend.
Build reconciliation from the start
Retries cover minutes. Replay covers hours. Reconciliation covers everything else, including the failures you never detected.
Schedule a job that compares counts and identifiers against the provider's API for the same time window, reports gaps, and backfills them. Run it daily at minimum, and alert on the difference rather than requiring someone to read a report.
This is the piece teams build after their first serious incident. Building it first costs a fraction and is the difference between finding a discrepancy the next morning and finding it at month-end close.
Instrument it in business terms
Technical metrics tell you the integration is running. Business metrics tell you it is correct.
Track orders written per minute, payments cleared per minute, records synchronised per hour, whatever the integration exists to move, and alert when the number deviates from its normal range. An integration returning 200s while writing nothing is the failure that hurts most, because every technical dashboard says it is healthy.
Classify errors rather than counting them: authentication and signature, rate limiting, schema, destination, and business-rule rejection each need a different response.
Price each connection individually
Every integration has its own data model, authentication scheme, failure behaviour and reconciliation requirement. Estimating them as one line hides that entirely.
A straightforward connection typically runs $2,500–$8,000; one involving bidirectional sync, complex mapping or an undocumented legacy system runs well beyond it. The variables that move the number are whether the API is documented and stable, whether it supports webhooks or requires polling, and whether you must write back or only read.
Ask any supplier to price connections separately and to name the one they are least confident about. Integration is where software projects overrun most often, and a single combined figure removes your ability to see it coming.
Security is part of the contract
Verify webhook signatures on every request and reject anything unsigned. Store credentials in a secret manager rather than configuration, and rotate them on a schedule someone owns.
Send the minimum data the other system needs. Integrations accumulate fields over time because adding one is easy, and each becomes a field you are responsible for protecting and, under GDPR and Turkish data protection law, justifying.
Plan for the provider changing
APIs change. Versions are deprecated, fields are removed, rate limits are tightened, and pricing is restructured.
Subscribe to the provider's changelog, pin an API version explicitly rather than tracking latest, and validate incoming payloads against an expected shape so a change surfaces as an alert rather than a silent data-quality problem months later.
Related guides
- Alongside API Integration: Planning, Patterns, and Common Failures, continue with AI Workflow Automation: Use Cases, Risks, and Roadmap.
- Alongside API Integration: Planning, Patterns, and Common Failures, continue with Admin Dashboard Development: Features, Architecture, and Cost Drivers.
Should API Integration: Planning, Patterns, and Common Failures lead to a platform the next three years can be built on, Discuss Your Platform Foundation.
Ali Boran Gazel