Follow one request through a customer portal. A client submits a service change with a signed document attached. Billing needs to know whether the account is current. Operations has to approve the requested date. Support needs a ticket if the request is incomplete, and the reporting dashboard should not count the change as active until operations approves it.
The first API call in that workflow proves the portal can reach the billing system when nothing is wrong. The production work is everything that happens when something is: a vendor returns HTTP 429, an OAuth token expires overnight, a webhook arrives twice, or billing and the CRM disagree about the customer's status.
Which system is allowed to change the field?
In that portal, the service-change record starts as pending, carries the document, and moves to active only when operations approves it. The portal can create the record and attach the file. It cannot set the status to active, and the billing sync reads the approved status rather than the requested one. Once that authority is written down per field, an argument about which system is right becomes a lookup.
Make the same decision for each field that billing, operations, support or reporting reads: which system may write it, and which systems only read it. An integration that lets two systems write the same status field produces conflicts no retry policy can settle. This is where API integration and wider software integration work overlap.
Duplicate and out-of-order deliveries
Read the vendor's delivery contract before writing the handler. Stripe's webhook documentation is a clear example: endpoints might occasionally receive the same event more than once, event order is not guaranteed, and in live mode undelivered events are retried for up to three days with exponential backoff. It also notes that the same change can sometimes produce two separate Event objects, so the event ID alone will not catch every duplicate; Stripe suggests using the ID of the underlying object together with the event type. Check the equivalent page for whichever vendor the integration actually uses.
The deduplication record belongs in the same database transaction as the business change. If the handler updates the service-change record and then crashes before logging the event ID, the vendor's retry arrives looking new and the update runs again. Writing the processed event ID, or the object ID and event type, in the same commit as the update closes that gap for the local database. It does not extend to a call the handler makes to another vendor or payment API: that effect is outside the transaction, a rollback does not undo it, and a crash after the call but before the commit leaves the unknown outcome described below. That outbound call needs its own idempotency key or a lookup before it is repeated. Out-of-order delivery needs a separate rule: when an "invoice paid" event arrives before the "invoice created" event, the handler can fetch the current object from the vendor's API instead of assuming the earlier event has already been processed.
Replay is safe for some failures and not others
A support screen with a replay button needs to know why each record failed, because the safe action differs by cause. A request the vendor rejected before executing it, such as a validation error for a missing tax code on an endpoint whose contract says validation runs before any change, is the candidate for replay once the mapping is fixed. That depends on the endpoint: some business APIs apply part of a request before reporting an error. The replay also needs the record's current state and the caller's authorization checked again, rather than resending the stored payload as it was. A request that timed out after it was sent is different. The vendor may have created the invoice, and the integration does not know.
Idempotency keys help here, within limits the vendor defines. Stripe's idempotent requests reference says it saves the status code and body of the first request made with a given key, success or failure, and returns the same result for later requests with that key, including 500 errors. It saves a result only once execution of the endpoint has begun: a request whose parameters fail validation, or that conflicts with another request executing concurrently, has no saved result and can be retried. Reusing a key with different parameters returns an error. Keys can be removed once they are at least 24 hours old, and a key reused after that point is treated as a new request. So for a request that reached execution, resending it with the original key and parameters while the key is retained returns the saved outcome, including a saved 500, instead of executing again. A replay after the key has been pruned may create a second object.
The support screen therefore needs at least two failure states: rejected before any change, and sent with an unknown outcome. The replay button is enabled only for the first. For the second, the integration looks up the vendor's record, or resends with the same key while the vendor still retains it, and marks the record complete or failed from what comes back. The accounting-timeout case in AI-assisted invoice review with human approval applies the same rule after an employee approves an invoice: check whether the accounting system recorded the approval before allowing a retry. For HTTP 429, RFC 6585 says the response may include a Retry-After header; when it does, the retry waits at least that long.
A readiness check before people rely on it
- Which system may write each field that customers, billing, operations or reporting depend on?
- Where is the processed event ID stored, and is it committed with the business change?
- Which outbound vendor or payment calls sit outside that commit, and what key or lookup protects each one from a repeat?
- How are credentials rotated without breaking the nightly job?
- What alerts support when a vendor rate-limits the connection or returns a payload the integration does not recognize?
- Which failure states can support replay, and which require a lookup first?