A payment was approved, but the merchant was never told
A sub-second race between a cron job and a browser request quietly lost a payment notification. Here is how it happened, and the one-line idea that fixed it.
A customer completed a card payment. The card was approved by the processor. And yet the merchant’s system was never told the order was paid, so the order sat in pending forever. No double charge, no lost money, just a payment that succeeded everywhere except in the one place that mattered to the merchant.
We found out because the merchant messaged us, not because anything alerted us. That alone is worth a section later. But the interesting part is the bug: a textbook check-then-act race condition, hiding behind a slow third-party call, in a system I would have told you was already idempotent.
The setup
Two services are involved. One is our payment-links platform, which owns orders and delivers webhooks to merchants. The other is a small payments microservice that owns the integration with the card processor. Both run on Cloudflare Workers with an edge SQLite database.
When a customer pays, the payment page calls the microservice: “sync this checkout.” The microservice asks the processor for the checkout’s status, stores the resulting transaction, and returns it. Back on the platform, an approved result flips the order from pending to paid exactly once and only then enqueues the payment.success webhook. That “exactly once” already made the confirmation path idempotent, so duplicate confirmations were safe.
There is also a background cron that runs every minute. For each checkout still awaiting payment, it runs the same status check and transaction write as the browser-driven request. Its job is to catch payments that complete after the customer closes their tab. So two callers execute the same code path against the same checkout, and nothing coordinates them.
Two properties of the processor’s API turned out to matter:
- A checkout’s result can be read only once. After the first successful read, the checkout identifier is spent.
- A spent checkout returns a “no payment session” code that belongs to the processor’s input-validation family. It is not a payment outcome, but our code treated it like one.
The race
Here is what happened in the space of about four seconds.
sequenceDiagram
participant B as Browser request
participant C as Cron (every minute)
participant P as Card processor
participant DB as Database
B->>P: read checkout status (slow, ~2s)
C->>P: read checkout status (slow, ~2s)
P-->>C: 000 approved (result consumed)
C->>DB: insert transaction (approved) OK
P-->>B: "no session" (already consumed)
B->>DB: insert transaction (expired) -> UNIQUE violation
Note over B: unhandled error -> HTTP 500
Both callers checked the database, found no transaction yet, and each started a roughly two-second call to the processor. The cron’s read landed first and recorded the approved payment. A fraction of a second later the browser request’s read came back as “no session,” because the result had already been consumed. Our code mapped that to EXPIRED and tried to write its own transaction, which collided with the uniqueness constraint on the table. The error was not caught, so the request returned 500.
That 500 crossed back to the platform, which treats any non-success from the microservice as fatal and stops. So the order was never marked paid and the webhook was never enqueued. And because nothing re-examines payments left in pending, the order never recovered on its own.
The root cause
The underlying fault is a check-then-act race: we decided whether to create a transaction by first checking if one existed, then creating it, with a slow external call in between. That gap is a wide-open window, and two uncoordinated callers walked into it together.
// Before: check, then act, with a slow call in the gap
const existing = await repo.findByReference(ref);
if (!existing) {
const status = await processor.getCheckoutStatus(id); // ~2s, read-once
await repo.create({ ref, status }); // the loser collides here
}
Three things escalated a benign race into a customer-visible failure. The read-once resource meant only one caller could win. Mapping the “no session” validation code to EXPIRED meant the loser was about to record the wrong status for a payment that had actually succeeded. And letting the resulting 500 abort the platform’s confirm-and-notify flow meant a transient collision became a permanently lost webhook.
The fix
The race window exists only because the check and the write are separate steps. Collapse them into one atomic, idempotent write and the window disappears:
// After: one atomic, idempotent write. The database decides the winner.
async createIfNotExists(tx: NewTransaction) {
const [row] = await db
.insert(transactions)
.values(tx)
.onConflictDoNothing({ target: [transactions.integrationId, transactions.referenceId] })
.returning();
// On conflict, nothing is returned; fetch the row the winner wrote.
return row ?? this.findByReference(tx.referenceId);
}
Now the loser does not throw. It gets back the transaction the winner already wrote and returns a normal 200, so the platform completes its sync and fires the webhook. No error-string parsing, no “did the insert fail because of a duplicate or something else” guessing, just the database enforcing uniqueness and telling us who won.
I paired that with a second change: when a read comes back as “no session,” instead of assuming the checkout expired, the microservice queries the processor’s reporting endpoint for the real outcome and only falls back to EXPIRED when there genuinely is no record. A losing reader now recovers the true approved status.
What I took away from it
Idempotency belongs at the data layer, not in a comment. I had reasoned that the confirmation flow was idempotent, and the high-level flow was. But the actual write was a check-then-act, and no amount of careful ordering upstream saves you from that. INSERT ... ON CONFLICT pushes the decision to the one place that can make it atomically.
A uniqueness constraint is a safety net, not a nuisance. The constraint is what turned this into a loud, recoverable 500 instead of a silent duplicate transaction wrongly marked expired. I have seen people remove constraints because they “cause errors in production.” The error was the constraint doing its job.
A validation error is not a domain outcome. Mapping “no session” straight to EXPIRED was the line that would have corrupted data. Error codes from a provider deserve the same scrutiny as their success codes.
A transient upstream error must never be allowed to strand state. One 500 should not permanently lose a webhook. The real gap was the absence of a reconciliation path: a scheduled job that re-syncs payments stuck in pending past a threshold would have healed this without anyone noticing.
Alert on the symptom, not just the cause. We now watch for payments that stay pending too long. You cannot foresee every race, but you can detect the shape of the failure they produce, so the next one is a page instead of a customer message.
The day after this fix, the same service taught me a second lesson about the unhappy path: a slow provider hanging our Workers, and why a retry is only safe when the write it repeats is idempotent.