This article explores architectural decisions for handling failed jobs, specifically reservation expiry in a healthtech context. It compares using a message queue with a Dead-Letter Queue (DLQ) and redrive mechanisms against a manual database polling approach. The core trade-off revolves around latency requirements, operational costs, and the complexity of ensuring idempotency, backoff, and poison-message handling.
Read original on Dev.to #architectureThe article addresses a common system design challenge: how to reliably process asynchronous jobs, particularly those requiring retries upon failure. It uses the example of expiring reservations in a healthtech application, where timely and accurate processing is crucial. The central dilemma presented is choosing between a queue-based system and a database polling mechanism, each with its own set of trade-offs and implications for system reliability and operational overhead.
The article contrasts two primary approaches for managing failed jobs:
Invariants for Robust Job Processing
The article highlights four critical invariants to guide the design of a reliable retry mechanism: 1. A reservation changes state from `held` to `expired` at most once. 2. A retry never extends the original hold duration. 3. A lost transport message can be reconstructed from durable application state. 4. An acknowledged message is not the audit record. This emphasizes that business audit logs should be stored in a durable application-owned history table, not solely rely on transient queue retention.