Gunjan Sharma

System Design

If async workers fail, a Dead Letter Queue (DLQ) is enough. It'll scale fine

· Updated

A Staff Engineer looked me in the eye last week and said:

"We have 8 microservices doing CRUD on one database. If async workers fail, a Dead Letter Queue (DLQ) is enough. It'll scale fine."

I had to politely disagree.

Here is why that architecture is a ticking production time bomb—and the exact failure modes that will explode at scale:


𝟭. 𝗧𝗵𝗲 𝗦𝗵𝗮𝗿𝗲𝗱 𝗗𝗕 𝗜𝗹𝗹𝘂𝘀𝗶𝗼𝗻 (𝗧𝗵𝗲 𝗗𝗲𝗮𝗱𝗹𝗼𝗰𝗸 𝗗𝗲𝗮𝘁𝗵 𝗦𝗽𝗶𝗿𝗮𝗹)

At 20 req/min in staging: Everything looks green.  
At 2,000 req/sec in production: Physics takes over.

When 8 independent services run concurrent CRUD workers (e.g., UPDATE accounts SET balance = balance - X):

1. Multiple workers mutate the same ledger rows simultaneously.  
2. Database engines (InnoDB/WiredTiger) take Exclusive Write Locks (X-locks).  
3. Worker A holds a lock on Row 1 waiting for Row 2. Worker B holds Row 2 waiting for Row 1. 

𝗧𝗵𝗲 𝗗𝗼𝗺𝗶𝗻𝗼 𝗘𝗳𝗳𝗲𝗰𝘁:
1. Lock wait queues saturate the connection pool.  
2. ERROR 1205: Lock wait timeout & ERROR 1213: Deadlock found.  
3. Failed workers trigger exponential retries.  
4. Retries create a Thundering Herd onto the DB. 

𝗧𝗵𝗲 𝗥𝗲𝘀𝘂𝗹𝘁: DB CPU hits 100% purely arbitrating lock contention. Throughput drops to 0%. Production freezes. 

𝗧𝗵𝗲 𝗙𝗶𝘅: A scalable ledger must be 𝗶𝗺𝗺𝘂𝘁𝗮𝗯𝗹𝗲 𝗮𝗻𝗱 𝗮𝗽𝗽𝗲𝗻𝗱-𝗼𝗻𝗹𝘆. Never UPDATE a balance row. Only INSERT double-entry journal items DEBIT / CREDIT). Appends eliminate row-update lock contention completely.


𝟮. 𝗧𝗵𝗲 "𝗗𝗟𝗤 𝗪𝗶𝗹𝗹 𝗦𝗮𝘃𝗲 𝗨𝘀" 𝗙𝗮𝗹𝗹𝗮𝗰𝘆

Relying on a Dead Letter Queue as your safety net confuses a transport pipe with a system of record. A queue moves ephemeral payloads; it is NOT a ledger.

𝗪𝗵𝘆 𝗮 𝗗𝗟𝗤 𝘄𝗶𝗹𝗹 𝗳𝗮𝗶𝗹 𝘆𝗼𝘂:
𝗕𝗿𝗼𝗸𝗲𝗿 𝗖𝗿𝗮𝘀𝗵𝗲𝘀: If a node drops before disk replication, the message vanishes before ever reaching the DLQ.
  
𝗡𝗲𝘁𝘄𝗼𝗿𝗸 𝗣𝗮𝗿𝘁𝗶𝘁𝗶𝗼𝗻𝘀: A worker writes to the DB, but the network drops before sending an ACK. The broker redelivers → Duplicate processing & double-spends. 

𝗧𝗧𝗟 𝗘𝘅𝗽𝗶𝗿𝘆: Under severe consumer lag, messages silently expire and get dropped.
  
𝗭𝗲𝗿𝗼 𝗔𝘂𝗱𝗶𝘁𝗮𝗯𝗶𝗹𝗶𝘁𝘆: You cannot construct an accounting proof from an error queue. A DLQ is an error trash can, not a financial audit log. 

𝗧𝗵𝗲 𝗙𝗶𝘅: The 𝗧𝗿𝗮𝗻𝘀𝗮𝗰𝘁𝗶𝗼𝗻𝗮𝗹 𝗢𝘂𝘁𝗯𝗼𝘅 𝗣𝗮𝘁𝘁𝗲𝗿𝗻. Write your business entity and task event into an Outbox table within the same 𝗹𝗼𝗰𝗮𝗹 𝗔𝗖𝗜𝗗 𝘁𝗿𝗮𝗻𝘀𝗮𝗰𝘁𝗶𝗼𝗻. 

Even if your message broker dies for 3 hours, 0 transactions are lost. An asynchronous CDC worker (like Debezium) streams events out the second the broker recovers.


𝗧𝗵𝗲 𝗛𝗮𝗿𝗱 𝗧𝗿𝘂𝘁𝗵
A queue cannot solve a database lock problem.  
And an error queue is not an audit log. 

Architect for failure modes on Day 1, or production will teach you at 03:00 AM.

If async workers fail, a Dead Letter Queue (DLQ) is enough. It'll scale fine | Gunjan Sharma