Integrations fail for a short list of reasons, and almost all of them are decided before any code is written: nobody agreed which system owns each field, nobody defined what happens when the other side is unavailable, and nobody built a way to detect silent divergence. The technical failures, timeouts, rate limits, schema changes, are the easy half.
This guide covers the causes in order of how often they cause real damage, and what prevents each one.
Key takeaways
- Undefined ownership is the root cause: Without a source of truth per field, the answer is whichever code ran last.
- Silent failure beats loud failure for damage: A connection returning success while writing nothing is the worst case.
- Monitor in business terms: Records synchronised per hour tells you more than uptime.
- Nearly every failure is recoverable if detected early: Which is what reconciliation is for.
The causes, ranked by damage
1. No agreed source of truth. Two systems hold the same customer address and both allow edits. Whichever syncs last wins, and neither is wrong from its own perspective. This produces the disagreements that erode trust in all your data, and it is a governance failure rather than a technical one.
2. Undefined failure behaviour. The other system is slow or down. Does the user wait, does the work queue, or does the transaction proceed and reconcile later? Unanswered, this becomes whatever the code happened to do, usually a timeout the user sees as a broken product.
3. Silent success. The integration returns 200 and writes nothing, or writes to the wrong place. Every technical dashboard shows healthy while the business quietly diverges. This is the most damaging failure mode because nothing alerts.
4. Duplicate processing. Providers guarantee at-least-once delivery, so the same event arrives twice. Without idempotency that becomes a double charge or a duplicate record, and the customer notices before you do.
5. Schema drift. The other side renames or removes a field. Without validation on receipt, the integration keeps running and silently stops carrying that data.
6. Policy flattening. Every error handled by the same retry path, so a rate-limit response triggers the same aggressive retry as a transient outage, turning a throttle into a storm.
7. Nobody owns it. The developer who built it has moved on, there is no runbook, and the first failure becomes an archaeology exercise.
What prevents each
| Cause | Prevention |
|---|---|
| No source of truth | A written field-level ownership map, agreed before build |
| Undefined failure behaviour | A decision per integration: block, queue, or proceed and reconcile |
| Silent success | Business-level monitoring and daily reconciliation |
| Duplicate processing | Idempotency keys checked before every write |
| Schema drift | Validate payloads on receipt; alert rather than drop |
| Policy flattening | Distinct handling per error class |
| No owner | A named owner each side and a runbook that does not need the original developer |
Design the failure states as business decisions
What happens when the other system is unavailable is not a technical detail. It determines whether a customer can complete a purchase, whether an order is lost, and whether staff have to reconcile something by hand tomorrow.
Answer per integration: is this synchronous or can it be queued, what does the user see, what is the maximum acceptable delay, and who is told when the delay is exceeded. Write it down alongside the mapping, that document is what makes an integration supportable by someone who did not build it.
Monitor what the integration exists to do
Uptime tells you the endpoint responds. It does not tell you it is working.
Track the business quantity: orders written per minute, records synchronised per hour, payments cleared. Alert on deviation from the normal range rather than on absolute failure, because the failures that cost most do not produce errors.
Classify errors instead of counting them, authentication, rate limit, schema, destination, business-rule rejection, because each needs a different response and a single error count hides which is happening.
Build reconciliation before the first incident
Retries cover minutes. Replay covers hours. Reconciliation covers everything else.
A daily job comparing counts and identifiers across the boundary, alerting on the gap, is the difference between finding a discrepancy the next morning and finding it at month-end close with three weeks of divergence behind it.
Teams almost always build this after their first serious incident. Building it first costs a fraction.
Plan for the other side changing
APIs are deprecated, fields are removed, rate limits tighten, vendors are acquired. None of these are your decision and all of them are your problem.
Pin an API version explicitly rather than tracking latest, subscribe to the provider's changelog, validate incoming payloads against an expected shape, and keep the integration isolated enough that replacing one provider does not mean rewriting your application.
Test the failures, not just the success
Most integration test suites cover the happy path, which is why most integration failures are discovered in production.
Test deliberately: send the same event twice and confirm the outcome is identical, simulate a timeout and confirm the queue behaves, send a malformed payload and confirm it alerts rather than silently dropping, and exhaust the retries to confirm the dead-letter queue captures it and someone is notified.
Related guides
- Alongside Integration Failures: Common Causes and Prevention, continue with AI Workflow Automation: Use Cases, Risks, and Roadmap.
- Alongside Integration Failures: Common Causes and Prevention, continue with Admin Dashboard Development: Features, Architecture, and Cost Drivers.
If the work prompted by Integration Failures: Common Causes and Prevention leads to a funded initiative that needs product strategy, design, engineering, or integration support, Discuss Your Platform Foundation.
Ali Boran Gazel