Skip to contentAnemo
EN
Contact

Integration Failures: Common Causes and Prevention

· 5 min read

Integrations fail for a short list of reasons, and almost all of them are decided before any code is written: nobody agreed which system owns each field, nobody defined what happens when the other side is unavailable, and nobody built a way to detect silent divergence. The technical failures, timeouts, rate limits, schema changes, are the easy half.

This guide covers the causes in order of how often they cause real damage, and what prevents each one.

Key takeaways

The causes, ranked by damage

1. No agreed source of truth. Two systems hold the same customer address and both allow edits. Whichever syncs last wins, and neither is wrong from its own perspective. This produces the disagreements that erode trust in all your data, and it is a governance failure rather than a technical one.

2. Undefined failure behaviour. The other system is slow or down. Does the user wait, does the work queue, or does the transaction proceed and reconcile later? Unanswered, this becomes whatever the code happened to do, usually a timeout the user sees as a broken product.

3. Silent success. The integration returns 200 and writes nothing, or writes to the wrong place. Every technical dashboard shows healthy while the business quietly diverges. This is the most damaging failure mode because nothing alerts.

4. Duplicate processing. Providers guarantee at-least-once delivery, so the same event arrives twice. Without idempotency that becomes a double charge or a duplicate record, and the customer notices before you do.

5. Schema drift. The other side renames or removes a field. Without validation on receipt, the integration keeps running and silently stops carrying that data.

6. Policy flattening. Every error handled by the same retry path, so a rate-limit response triggers the same aggressive retry as a transient outage, turning a throttle into a storm.

7. Nobody owns it. The developer who built it has moved on, there is no runbook, and the first failure becomes an archaeology exercise.

What prevents each

Cause Prevention
No source of truth A written field-level ownership map, agreed before build
Undefined failure behaviour A decision per integration: block, queue, or proceed and reconcile
Silent success Business-level monitoring and daily reconciliation
Duplicate processing Idempotency keys checked before every write
Schema drift Validate payloads on receipt; alert rather than drop
Policy flattening Distinct handling per error class
No owner A named owner each side and a runbook that does not need the original developer

Design the failure states as business decisions

What happens when the other system is unavailable is not a technical detail. It determines whether a customer can complete a purchase, whether an order is lost, and whether staff have to reconcile something by hand tomorrow.

Answer per integration: is this synchronous or can it be queued, what does the user see, what is the maximum acceptable delay, and who is told when the delay is exceeded. Write it down alongside the mapping, that document is what makes an integration supportable by someone who did not build it.

Monitor what the integration exists to do

Uptime tells you the endpoint responds. It does not tell you it is working.

Track the business quantity: orders written per minute, records synchronised per hour, payments cleared. Alert on deviation from the normal range rather than on absolute failure, because the failures that cost most do not produce errors.

Classify errors instead of counting them, authentication, rate limit, schema, destination, business-rule rejection, because each needs a different response and a single error count hides which is happening.

Build reconciliation before the first incident

Retries cover minutes. Replay covers hours. Reconciliation covers everything else.

A daily job comparing counts and identifiers across the boundary, alerting on the gap, is the difference between finding a discrepancy the next morning and finding it at month-end close with three weeks of divergence behind it.

Teams almost always build this after their first serious incident. Building it first costs a fraction.

Plan for the other side changing

APIs are deprecated, fields are removed, rate limits tighten, vendors are acquired. None of these are your decision and all of them are your problem.

Pin an API version explicitly rather than tracking latest, subscribe to the provider's changelog, validate incoming payloads against an expected shape, and keep the integration isolated enough that replacing one provider does not mean rewriting your application.

Test the failures, not just the success

Most integration test suites cover the happy path, which is why most integration failures are discovered in production.

Test deliberately: send the same event twice and confirm the outcome is identical, simulate a timeout and confirm the queue behaves, send a malformed payload and confirm it alerts rather than silently dropping, and exhaust the retries to confirm the dead-letter queue captures it and someone is notified.

If the work prompted by Integration Failures: Common Causes and Prevention leads to a funded initiative that needs product strategy, design, engineering, or integration support, Discuss Your Platform Foundation.

Frequently asked questions

What should be defined first?

Start by defining the expected result and owner for service boundary. Then follow one real example through data and integration, recording the data used, waiting points, exceptions, and evidence of completion. This creates a more reliable first scope than a screen inventory.

How should success be measured?

Review failure rate, latency, recovery time, and data accuracy together. Give each measure a definition, data source, owner, review cadence, and response when it crosses a threshold. A single speed or usage metric should not hide quality, rework, or abandonment.

Does this work always require new software?

New software is not automatic. If the underlying problem is policy, ownership, training, or an unnecessary approval, fix the process first. Configure an established tool when it supports the critical workflow and data boundary. Consider custom development only when a differentiating rule, integration, or experience creates clear value.

Related services

Related reading

Building the product for what comes next

We would rather deliver one product that holds up than three that have to be rebuilt. That standard is the same on every project, whatever its size.

Ali Boran GazelCEO

Contact us