Menu
Dev.to #architecture·October 9, 2026

Designing Reliable Job Retries with Queues vs. Database Polling

This article explores architectural decisions for handling failed jobs, specifically reservation expiry in a healthtech context. It compares using a message queue with a Dead-Letter Queue (DLQ) and redrive mechanisms against a manual database polling approach. The core trade-off revolves around latency requirements, operational costs, and the complexity of ensuring idempotency, backoff, and poison-message handling.

Read original on Dev.to #architecture

The article addresses a common system design challenge: how to reliably process asynchronous jobs, particularly those requiring retries upon failure. It uses the example of expiring reservations in a healthtech application, where timely and accurate processing is crucial. The central dilemma presented is choosing between a queue-based system and a database polling mechanism, each with its own set of trade-offs and implications for system reliability and operational overhead.

Queue-Based System with DLQ vs. Database Polling

The article contrasts two primary approaches for managing failed jobs:

  • Queue with DLQ and Redrive: This method involves sending job messages to a queue. If a job fails after multiple attempts, it's moved to a Dead-Letter Queue (DLQ). A redrive mechanism allows these messages to be returned to the main queue for reprocessing, typically after manual intervention or a fix. This approach naturally handles backoff, poison messages, and offers better visibility into stuck jobs.
  • Manual Database Polling: Here, a database table stores job states (pending, claimed, retry, terminal). A poller periodically queries the database for jobs to process. This approach requires manual implementation of features like backoff, concurrent claim safety (preventing multiple workers from processing the same job), and recovery for stuck jobs. It can seem simpler initially for small teams already familiar with database operations.

Key Invariants for Reliable Processing

ℹ️

Invariants for Robust Job Processing

The article highlights four critical invariants to guide the design of a reliable retry mechanism: 1. A reservation changes state from `held` to `expired` at most once. 2. A retry never extends the original hold duration. 3. A lost transport message can be reconstructed from durable application state. 4. An acknowledged message is not the audit record. This emphasizes that business audit logs should be stored in a durable application-owned history table, not solely rely on transient queue retention.

job processingmessage queuesdead-letter queueidempotencyretry mechanismsdatabase pollingsystem reliabilitydistributed tasks

Comments

Loading comments...