A Staff Engineer looked me in the eye last week and said:
"We have 8 microservices doing CRUD on one database. If async workers fail, a Dead Letter Queue (DLQ) is enough. It'll scale fine."
I had to politely disagree.
Here is why that architecture is a ticking production time bomb—and the exact failure modes that will explode at scale:
𝟭. 𝗧𝗵𝗲 𝗦𝗵𝗮𝗿𝗲𝗱 𝗗𝗕 𝗜𝗹𝗹𝘂𝘀𝗶𝗼𝗻 (𝗧𝗵𝗲 𝗗𝗲𝗮𝗱𝗹𝗼𝗰𝗸 𝗗𝗲𝗮𝘁𝗵 𝗦𝗽𝗶𝗿𝗮𝗹)
At 20 req/min in staging: Everything looks green.
At 2,000 req/sec in production: Physics takes over.
When 8 independent services run concurrent CRUD workers (e.g., UPDATE accounts SET balance = balance - X):
1. Multiple workers mutate the same ledger rows simultaneously.
2. Database engines (InnoDB/WiredTiger) take Exclusive Write Locks (X-locks).
3. Worker A holds a lock on Row 1 waiting for Row 2. Worker B holds Row 2 waiting for Row 1.
𝗧𝗵𝗲 𝗗𝗼𝗺𝗶𝗻𝗼 𝗘𝗳𝗳𝗲𝗰𝘁:
1. Lock wait queues saturate the connection pool.
2. ERROR 1205: Lock wait timeout & ERROR 1213: Deadlock found.
3. Failed workers trigger exponential retries.
4. Retries create a Thundering Herd onto the DB.
𝗧𝗵𝗲 𝗥𝗲𝘀𝘂𝗹𝘁: DB CPU hits 100% purely arbitrating lock contention. Throughput drops to 0%. Production freezes.
𝗧𝗵𝗲 𝗙𝗶𝘅: A scalable ledger must be 𝗶𝗺𝗺𝘂𝘁𝗮𝗯𝗹𝗲 𝗮𝗻𝗱 𝗮𝗽𝗽𝗲𝗻𝗱-𝗼𝗻𝗹𝘆. Never UPDATE a balance row. Only INSERT double-entry journal items DEBIT / CREDIT). Appends eliminate row-update lock contention completely.
𝟮. 𝗧𝗵𝗲 "𝗗𝗟𝗤 𝗪𝗶𝗹𝗹 𝗦𝗮𝘃𝗲 𝗨𝘀" 𝗙𝗮𝗹𝗹𝗮𝗰𝘆
Relying on a Dead Letter Queue as your safety net confuses a transport pipe with a system of record. A queue moves ephemeral payloads; it is NOT a ledger.
𝗪𝗵𝘆 𝗮 𝗗𝗟𝗤 𝘄𝗶𝗹𝗹 𝗳𝗮𝗶𝗹 𝘆𝗼𝘂:
𝗕𝗿𝗼𝗸𝗲𝗿 𝗖𝗿𝗮𝘀𝗵𝗲𝘀: If a node drops before disk replication, the message vanishes before ever reaching the DLQ.
𝗡𝗲𝘁𝘄𝗼𝗿𝗸 𝗣𝗮𝗿𝘁𝗶𝘁𝗶𝗼𝗻𝘀: A worker writes to the DB, but the network drops before sending an ACK. The broker redelivers → Duplicate processing & double-spends.
𝗧𝗧𝗟 𝗘𝘅𝗽𝗶𝗿𝘆: Under severe consumer lag, messages silently expire and get dropped.
𝗭𝗲𝗿𝗼 𝗔𝘂𝗱𝗶𝘁𝗮𝗯𝗶𝗹𝗶𝘁𝘆: You cannot construct an accounting proof from an error queue. A DLQ is an error trash can, not a financial audit log.
𝗧𝗵𝗲 𝗙𝗶𝘅: The 𝗧𝗿𝗮𝗻𝘀𝗮𝗰𝘁𝗶𝗼𝗻𝗮𝗹 𝗢𝘂𝘁𝗯𝗼𝘅 𝗣𝗮𝘁𝘁𝗲𝗿𝗻. Write your business entity and task event into an Outbox table within the same 𝗹𝗼𝗰𝗮𝗹 𝗔𝗖𝗜𝗗 𝘁𝗿𝗮𝗻𝘀𝗮𝗰𝘁𝗶𝗼𝗻.
Even if your message broker dies for 3 hours, 0 transactions are lost. An asynchronous CDC worker (like Debezium) streams events out the second the broker recovers.
𝗧𝗵𝗲 𝗛𝗮𝗿𝗱 𝗧𝗿𝘂𝘁𝗵
A queue cannot solve a database lock problem.
And an error queue is not an audit log.
Architect for failure modes on Day 1, or production will teach you at 03:00 AM.
System Design
If async workers fail, a Dead Letter Queue (DLQ) is enough. It'll scale fine
· Updated