English · Deutsch
Reference implementations of the handful of patterns that decide whether a system-to-system integration survives contact with reality. TypeScript, zero runtime dependencies, fully tested.
This is not a framework and it is not an HTTP client. It is the small, awkward logic that sits between two systems that were never built to talk to each other. Written out properly, so it can be read, argued with and copied.
An animated walkthrough of each pattern. Watch a duplicate webhook get absorbed. Watch eight clients take down a recovering service, then watch them not take it down. Watch the race that a SELECT-then-INSERT claim actually loses. English and German, light and dark.
pnpm add github:ArsalanRC/integration-patternsInstalled straight from the repository: the
preparescript compiles it on install, so there is no build step for you. pnpm rather than npm is deliberate. npm runs dependencypostinstallscripts by default, which is the main supply-chain attack vector in this ecosystem; pnpm blocks them unless each one is named. Same registry, safer default.
Integrations fail in a narrow and very repetitive set of ways. A webhook is delivered twice. A carrier API times out after it created the label. A downstream system is down for ninety seconds and every client retries at the same instant, so it stays down. An ERP and a warehouse disagree about stock and nobody notices for a week.
None of that is exotic. What makes it expensive is that each one gets solved ad hoc, slightly differently, in every service. Usually under time pressure, and usually after it has already cost someone a shipment.
Each pattern here documents the failure it prevents and, just as important, the way people get it wrong. The code is short. The reasoning is the point.
The problem: a sender times out waiting for you. It does not know whether you processed the message, so it sends again. If your handler charges a card or books a shipment, doing it twice is a real-world incident.
import { withIdempotency, MemoryIdempotencyStore } from "@arsalanrc/integration-patterns";
const store = new MemoryIdempotencyStore();
const result = await withIdempotency(store, event.id, async () => {
return bookShipment(event.payload); // runs at most once per event.id
});- First delivery claims the key, runs the handler, stores the result.
- Redelivery returns the stored result without running anything.
- Concurrent delivery throws
DuplicateInFlightError, so you can reply 409. - Handler throws releases the claim, so a retry genuinely retries.
That last point is the one that bites. It is tempting to mark the key done as soon as you start, so duplicates bounce off. Do that and a single transient failure becomes permanent silent data loss: the work never happened, but every future delivery is treated as a duplicate and dropped. Failure has to release the claim.
The mirror-image mistake is replying 200 to a concurrent delivery. The in-flight attempt might still fail, and you have just told the sender it succeeded, so it will never send again.
claim must be atomic. If two workers can both be told they won the same key, the pattern is decorative.
INSERT INTO idempotency_keys (key, state, claimed_at)
VALUES ($1, 'in_flight', now())
ON CONFLICT (key) DO NOTHING
RETURNING key;ON CONFLICT DO NOTHING is doing the real work. The intuitive version, SELECT then INSERT, has a window between the two statements where both workers read "not present" and both insert. Under normal load you may never see it. Under a redelivery storm, which is precisely when idempotency matters, you will.
Full DDL, including the reaper for claims whose worker died and the retention policy, is in schema.sql.
MemoryIdempotencyStore is for tests and single-process development. It is not a production option: the moment you run two instances, two separate Maps protect you from nothing.
The problem: retries either help or turn a brief outage into a long one. Which one you get depends on two decisions.
import { retry, isRetryableHttpError } from "@arsalanrc/integration-patterns";
const label = await retry((attempt) => carrier.createLabel(shipment), {
attempts: 4,
baseMs: 250,
deadlineMs: 10_000,
isRetryable: isRetryableHttpError,
onRetry: ({ attempt, delayMs, error }) => log.warn({ attempt, delayMs, error }),
});Decision one: know what not to retry. A 500 deserves another attempt. A 422 will fail identically every time, and retrying it burns your budget while delaying the error reaching a human who could fix it. isRetryable is a required option, not an optional one, because there is no safe default and the caller is the only one who knows.
Decision two: jitter, properly. Without it, every client that failed during the same outage retries at the same instant and knocks the recovering service straight back over.
This uses full jitter: the delay is picked uniformly from [0, window), where the window grows exponentially and is capped.
delay = random() * min(maxDelayMs, baseMs * 2^attempt)
Adding a few percent of noise to a fixed delay is not the same thing and does not work. It shifts the herd rather than spreading it. The test suite asserts the distribution, not just the ceiling.
Two smaller details that matter in production:
- No sleep after the final attempt. A trailing wait delays the caller's failure for no benefit.
- The deadline is checked against the wait you are about to take, not just time already spent. Sleeping past your own deadline and noticing afterwards wastes the exact budget the deadline exists to protect.
Retry handles the failures that pass on their own. A malformed payload, a schema the consumer no longer understands, an account deleted between the send and the delivery: none of those get better on the fourth attempt. Retrying them forever is how a queue fills up and how the one real incident ends up buried under a million copies of the same error.
await withDeadLetter(
() => retry(() => deliver(msg), { attempts: 5, isRetryable: isRetryableHttpError }),
{ store, payload: msg, key: msg.id, attempts: 5 },
);The original error always reaches the caller. This wrapper decides nothing about whether a failure is fatal, because that belongs with whoever knows what the message is.
The entry stores why, not just what. Anyone opening this queue is asking why it failed. A payload on its own makes them reproduce the failure to find out. The error class, its message, the stack, the attempt count and both timestamps are recorded with it.
Replay is one at a time and only when asked. Nothing re-runs on a timer.
await replayDeadLetter(store, id, async (entry) => deliver(entry.payload));Replaying the queue in bulk after an outage replays the outage. The entries are there precisely because something broke. If it is still broken, you have turned one incident into a stampede. Read the cause, fix the thing, replay the entry. A failed replay leaves the entry exactly where it was, so nothing is lost by trying.
A dead letter that cannot be written is the worst case in this library. If the store throws, the message was not delivered and not recorded. It is nowhere. Swallowing that, or letting it replace the original error, is how a message disappears with nobody the wiser. It raises DeadLetterWriteError instead, carrying the delivery failure as cause and the store failure as writeCause. The second is the one to page on: the first is an application problem, this is a lost message.
Zero runtime dependencies. This code sits on the path of every message an integration handles. It should not drag a tree of transitive packages onto that path.
Storage-agnostic. Behaviour is defined by interfaces, so the same logic runs over Postgres in production and a Map in a test. There is no database driver in this package.
Everything injectable that a test needs to control. Clock, sleep, and randomness are all parameters. That is why the suite runs in milliseconds instead of genuinely sleeping through backoff windows, and why jitter can be made deterministic when asserting on it.
Errors are classes. A webhook handler has to distinguish "duplicate, reply 200" from "genuinely failed, reply 500 so the sender retries". It cannot do that by matching on a message string.
Built in the open, one pattern at a time, each with its reasoning written down.
| Pattern | State |
|---|---|
| Idempotency | Done |
| Retry, backoff, jitter | Done |
| Dead-letter queue | Done |
| Reconciliation, drift detection between two systems | Planned |
| Circuit breaker | Planned |
Reconciliation is the one I am most interested in getting right. Two systems both believe they are correct. Deciding what to do about the difference is the problem that eats the most time in practice, and the one generic tooling covers least well.
pnpm install
pnpm test # 41 tests
pnpm run type-check
pnpm run build| Strategy | Delay | When to pick it |
|---|---|---|
full (default) |
anywhere in [0, window) |
Almost always. It spreads clients across the whole window, which is what actually stops them returning together |
equal |
window/2 plus anywhere in [0, window/2) |
When a retry is expensive and must never come back almost instantly. Keeps a floor under the delay and still spreads half a window |
none |
exactly window |
Tests and demos, where reproducible timing matters. Never against a shared service: every client that failed together returns together |
attempts limits how many times you try. deadlineMs limits how long you are
willing to wait in total. They answer different questions, and a call that
matters usually wants both: attempts stops a fast-failing dependency spinning,
the deadline stops a slow one holding your request open past the point anybody
is still waiting for it.
Arsalan Khadim · LinkedIn · GitHub
MIT © Arsalan Khadim