The SQS Deduplication Trap: Why Naive Webhook Queues Silently Drop Payment Retries
Every senior engineer will tell you: "Just put SQS and Lambda in front of your webhooks in 20 lines of code."
The pattern seems simple enough: an API Gateway endpoint verifies the signature, drops the event onto an Amazon SQS queue, and returns an immediate 200 OK to Stripe. A separate consumer Lambda drains the queue and processes the customer subscription.
Then you hit the duplicate delivery problem. SQS provides at-least-once delivery, meaning network retries or visibility timeouts will inevitably deliver the same invoice twice. Your customers start receiving duplicate billing confirmation emails and double credits.
The "Fix" That Causes Silent Data Loss
To prevent duplicates, developers typically add a unique constraint table in Postgres or DynamoDB:
// The common "quick" deduplication pattern:
async function handleWebhook(event) {
// Try to claim the event upfront
const res = await db.query(
"INSERT INTO processed_events (event_id) VALUES ($1) ON CONFLICT DO NOTHING RETURNING *",
[event.id]
);
if (res.rows.length === 0) {
// Seen before! Skip to prevent duplicate email
return { status: 'duplicate_skipped' };
}
try {
await provisionSubscription(event);
} catch (err) {
// Rollback marker so retry can try again!
await db.query("DELETE FROM processed_events WHERE event_id = $1", [event.id]);
throw err;
}
}
Notice the assumption: you assume the
catch block will always execute to clean up the marker if something goes wrong. In reality, production serverless environments fail hard:
- Lambda Out-Of-Memory (OOM): Process is killed by the runtime immediately. No catch block runs.
- Hard Execution Timeout: Lambda reaches its 15-minute or API Gateway 29-second ceiling. Process terminated.
- Database Pool Exhaustion: The downstream query timed out because Postgres ran out of connections. The
DELETEstatement fails with the exact same connection error!
The Silent Disappearance of Payment Events
When the process dies without executing the DELETE cleanup, the event_id marker remains locked in the database.
Stripe's exponential backoff kicks in and delivers the webhook again 1 hour later. The worker receives the retry, queries INSERT ... ON CONFLICT DO NOTHING, sees the stale marker from the earlier crashed attempt, concludes it was already successfully handled, and returns 200 OK!
The event is gone forever. The customer was charged, your database has no subscription record, and no alert was fired.
The Solution: Status-Keyed Ingress Deduplication
A resilient webhook gateway must decouple receipt from downstream execution status. In HookArmor, deduplication operates under three strict invariants:
- Only Deduplicate Successful Deliveries: Deduplication checks only suppress if the existing record in storage has
status === 'delivered'. If an event failed, timed out, or is quarantined in the DLQ, re-deliveries are never suppressed. - Single Logical Event Identity: An incoming retry for a failed event touches and re-queues the existing row instead of creating duplicate competing rows.
- Destination Concurrency Limiting: Rather than letting 100 simultaneous invoices flood your database and exhaust the connection pool, HookArmor's ingress buffers the burst and relays requests under a configurable concurrency cap (e.g., 5 concurrent connections).
Bulletproof Webhooks Without The AWS Glue Code
HookArmor provides safe status deduplication, concurrency protection, and 1-click dead-letter replay in a single self-hostable binary.
Star on GitHub