3 October 2026
TypeScript Job Queue: How to Choose for Retries, Concurrency, and Recovery

TL;DR
A production TypeScript job queue must preserve work, coordinate separate workers, bound retries, and make failures repairable. Compare an integrated application queue such as db3.ai, a dedicated queue library such as BullMQ, and a PostgreSQL-centered library such as pg-boss against your existing infrastructure and the failure behavior you can actually operate. Whichever you choose, worker concurrency does not prevent duplicate business effects; use a stable business identity and test recovery before launch.
A promise that runs after an HTTP handler returns can feel like a job queue. But if the process exits, there may be no durable work to recover. I would choose a queue by walking through the complete path from accepted request to stored job, worker claim, external side effect, completion, and possible replay. The important question is not just whether a library can process an asynchronous function. It is what your users and operators can prove when any step fails. Microsoft's background-job guidance distinguishes asynchronous work from the mechanisms needed to observe its completion.
This guide is for product engineers with a TypeScript backend choosing where jobs live and how workers operate. It is not a duplicate setup tutorial. Once you have chosen db3.ai, the Queue guide supplies the complete installation, migration, job registration, dispatch, and worker steps.
Table of contents
- First decide whether you need a durable queue
- Three approaches, compared on the same decisions
- Test the failure path before the throughput path
- Concurrency is three different limits
- A compact db3.ai implementation path
- A production acceptance checklist
- Frequently asked questions
- Sources
- Recommended Reads
First decide whether you need a durable queue
An in-process task list coordinates work only while its process is alive. It can be appropriate for disposable work whose loss is acceptable. It is not a substitute for stored jobs when a completed API request commits the product to producing a report or delivering an event.
A durable job queue stores a work identity and payload independently of the producer. Workers claim available jobs, acknowledge completed attempts, and apply a defined policy to failed or abandoned attempts. Storage alone is not sufficient: ask how claims expire, whether old workers can acknowledge a job after losing ownership, and how the queue records terminal failures. A queue ID confirms enqueueing, not a successful business result. The db3.ai Queue lifecycle documentation makes that distinction explicit.
A scheduler determines when to create work. A workflow engine coordinates related steps and their state. A message broker can distribute messages to consumers with routing patterns beyond one background task. These overlap in products, but they answer different questions. If you need a daily report, identify both the due occurrence and the queued report job; if you need a multi-step approval, decide whether a single retried handler is sufficient. Microsoft's background-job architecture guidance separates event-driven and schedule-driven triggers, and db3.ai distinguishes Queue, Scheduler, Events, and Flows.
Three approaches, compared on the same decisions
For a product team choosing a conventional background-job path, I would shortlist three architectures rather than rank every tool with a queue-shaped API. The criteria are: where work persists, who operates workers, how attempts recover, and how capacity is limited. Each option still needs application-owned idempotency and an alerting plan.
| Approach | Persistence and worker boundary | Retry and recovery path | Capacity and operational responsibility |
|---|---|---|---|
| db3.ai Queue | Application database driver by default, or Redis driver; a separately started worker consumes named queues. | Per-dispatch attempts, backoff and retry window; terminal failures are retained and can be replayed as linked replacement jobs. | Deploy and supervise workers, choose queue-to-worker allocation, and supply business-level limits and monitoring. |
| BullMQ | Redis backend by default, with a documented optional PostgreSQL backend; independent Worker processes. | Configurable attempts and fixed or exponential backoff; failed-job retention and manual retry depend on queue configuration. | Worker-local concurrency and documented queue-wide rate limiting; operate the chosen backend, workers, and alerting. |
| pg-boss | PostgreSQL-backed jobs and polling workers; enqueue can participate in a supported database transaction. | Retry policies, terminal or dead-letter handling, and operator retry or redrive. | Size database connections, worker concurrency and queue policies; operate PostgreSQL maintenance and worker processes. |
These are architecture choices, not interchangeable guarantees. The rows summarize the db3.ai Queue guide, BullMQ's backend and retry documentation, BullMQ's retry guide, and the pg-boss project documentation. Check the selected backend and product version before adopting a specific option.

Choose the integrated path when your application already uses @db3.ai/app and sharing its database, services, job registrations, and worker bootstrap is valuable. Its named queues let you separate a slow report workload from another worker pool. That is routing and capacity isolation, not a built-in promise of a provider-wide API rate limit. Its database and Redis drivers share a Queue API, but you still need migrations or backend configuration, worker supervision, and a policy for durable business outcomes. The Queue guide's driver and worker sections document those boundaries.
Choose the dedicated library path when a queue is a separately owned subsystem and your team wants direct control over worker deployment and retry behavior. BullMQ's default backend is Redis, while its documentation also describes an optional PostgreSQL backend. Do not reduce this decision to the old shorthand that BullMQ can only use Redis; instead test the backend you actually plan to operate. Its worker concurrency setting is local to one worker, while additional processes increase total capacity. BullMQ's PostgreSQL backend and worker concurrency documentation describe those distinctions.
Choose the PostgreSQL-centered path when keeping queue data in an existing PostgreSQL operational domain matters, particularly if enqueueing must participate in a supported database transaction. pg-boss documents retries, dead-letter redrive, and worker concurrency controls. PostgreSQL reduces the number of distinct datastores in this case, but queue volume, connections, retention, and monitoring still consume database capacity. Its project documentation and worker reference explain the available controls and transaction limits.
There is no universal winner: the smallest infrastructure diagram is not necessarily the simplest recovery procedure. If work consists of coordinated steps with independent progress, compare a workflow system separately rather than forcing the entire process into one queue handler.
Test the failure path before the throughput path
Imagine a report job that writes an artifact and then loses its worker lease before the queue records success. Another worker may run the same logical report. A durable claim can protect the queue record without undoing an already completed file write or external API request. I would test the following transitions with a disposable job, not infer reliability from a successful enqueue. Microsoft recommends designing background jobs for repeated execution, and db3.ai documents lease loss as an ambiguous outcome.

Persist before acknowledging. If a request saves a business record and then dispatches a job, identify what happens when either operation succeeds alone. A shared database does not automatically make every application write and enqueue one transaction. db3.ai's report-outside-the-request example calls for an application outbox or another explicit handoff when atomicity is required; pg-boss documents sending through supported transaction adapters. The design you choose must match the actual write boundary, including any external storage.
Classify a failure before retrying. A temporary unavailable provider may merit backoff; invalid input usually needs repair. Set both an attempt budget and a usefulness deadline. Check the units when comparing products: db3.ai dispatch retry options use seconds, while BullMQ's illustrated retry backoff delays use milliseconds. In db3.ai, an explicit QueueRetryLaterError can defer provider backpressure without consuming an ordinary attempt, so your application must enforce a separate deadline or the job could defer indefinitely. db3.ai's retry example and BullMQ's retry documentation show the different policies.
Make the business effect repeatable. Key a generated report by an application-owned report revision and replace or conditionally publish its result. A queue job ID is not a durable business key across replay: db3.ai's retryFailed() creates a replacement linked to the earlier failed record. For email or payments, a timed-out provider call may already have succeeded; use the provider's documented idempotency or reconciliation mechanism before trying again. An application flag alone cannot resolve that ambiguous outcome. The db3.ai report example illustrates repeatable output without claiming exactly-once execution.
Keep terminal failure inspectable. Decide who can view payloads and errors, how the original failure stays auditable, what repair is required, and who authorizes replay. Do not turn every failure into automatic requeueing. db3.ai exposes failed-job inspection and linked replay through its Queue API; BullMQ's retry behavior depends partly on retention settings; pg-boss documents job retry and redrive operations.
Concurrency is three different limits
Worker concurrency caps jobs a process handles simultaneously. Queue allocation decides which workload receives worker capacity. Downstream limits constrain calls to a database, tenant, or external provider across all worker processes. A setting of five per worker does not become a global limit of five after you deploy several workers. BullMQ explicitly distinguishes local worker concurrency from multiple workers; pg-boss documents both local and database-coordinated group controls. BullMQ worker concurrency and pg-boss worker options are useful references when sizing a fleet.
For example, suppose an indexing job waits on a provider and each worker has four in-flight slots. Three replicas could have twelve requests in flight unless another coordination layer constrains the total. That is an upper-bound illustration, not a measured throughput prediction. I would separate provider-bound jobs into their own queue, set a fleet-wide rate or concurrency policy where required, and test what happens when the provider starts rejecting calls. Check fairness too: one tenant's backlog should not silently starve urgent work. A named queue helps partition workers; it does not by itself provide per-tenant fairness.
CPU-bound transformations change the answer. Increasing asynchronous concurrency within one Node.js process is most useful when handlers spend time waiting on I/O; it does not create additional CPU cores. BullMQ recommends separate processes or sandboxed processors for CPU-intensive jobs in its concurrency guide. Whichever queue you choose, measure queue wait, processing duration, provider errors, and database connection pressure before increasing worker counts.
A compact db3.ai implementation path
If you choose db3.ai, first complete the Queue setup guide: configure a persistent driver, migrate QueuedJob and FailedJob for the database driver, create GenerateReportJob, register it in the bootstrap shared by producer and worker, and run a worker listening on reports. The guide's initial handler logs a report ID; it does not generate a report until you implement your own repeatable reporting logic. Work from an installed app with its documented configuration. The following dispatch belongs inside a booted route or application service, not at module top level:
import { app } from '@db3.ai/app/server';
import { GenerateReportJob } from './jobs/GenerateReportJob';
const jobId = await app().queue.dispatch(
new GenerateReportJob({ reportId: 'weekly-v1' }),
{
queue: 'reports',
maxTries: 5,
backoff: {
strategy: 'exponential',
initialSeconds: 15,
maxSeconds: 900,
jitter: true,
},
retryUntilSeconds: 3_600,
},
);
jobId identifies queued work; it is not the report's completion status or an authorization token for a user-facing status endpoint. A reports worker must be running, and it must register the same job class. The dispatch options above follow the documented Queue guide and public dispatch contract. I have not executed this snippet in your application; verify it with your installed package and your own report handler.
After the underlying cause of a terminal failure has been repaired, an authorized operator action can replay a chosen failed record. This helper assumes your application has already verified that the supplied ID refers to the intended failed report and that repeating its effect is safe; it does not supply the access check or the repair itself:
import { app } from '@db3.ai/app/server';
import type { QueueJobId } from '@db3.ai/app/queue';
export async function replayRepairedReport(
failedJobId: QueueJobId,
): Promise<QueueJobId> {
return app().queue.retryFailed(failedJobId, { queue: 'reports' });
}
The returned ID belongs to a replacement job, not a retroactive success of the original. Check its outcome and preserve the earlier failure record. The queued-work example shows the failure and replay lifecycle in a complete SQL lab; the Queue API reference documents retryFailed() and the identifier type. Test the same path with the driver and deployment configuration you intend to run.
A production acceptance checklist
Before selecting a library, I would ask the on-call engineer to perform one controlled recovery drill and record the answers:
- Durability: Stop the producer immediately after dispatch. Can a separately started worker claim the job, and can the team identify any business-record-to-enqueue gap?
- Retry budget: Produce a transient error and an invalid payload. Which one retries, with what backoff and deadline, and where does each terminal error appear?
- Duplicate safety: Stop a worker after a harmless side effect but before completion. Does replay preserve one logical business result?
- Capacity: Increase replicas while a provider is rate limiting. Are the limits fleet-wide or per process, and can a noisy queue delay critical work?
- Visibility: Can you distinguish waiting, claimed, deferred, failed, replayed, and successfully completed business work? Correlate logs by business key, queue ID, attempt, and worker, then alert on old waiting work and terminal failures.
- Deployment: Drain workers on shutdown, retain compatibility with old queued payloads during releases, and verify that a worker on the wrong named queue does not falsely signal that all work is done.
These checks turn feature claims into observable behavior. For a db3.ai deployment, the Queue guide documents leases, named workers, retry outcomes, and startup registration, while the report lab exercises a separately booted producer and worker. Neither guide removes the need to test your own external effects and production operations.
Frequently asked questions
Does a durable TypeScript job queue guarantee exactly-once effects?
No. A worker can perform an effect and lose the ability to record completion; another attempt may repeat the effect. Keep a stable business identity, make local writes repeatable where possible, and reconcile uncertain remote operations before replaying. Microsoft's idempotency guidance and db3.ai's lease discussion explain why queue ownership alone is insufficient.
Should scheduled jobs use the same queue as request-triggered jobs?
They can, but only when the workloads can share worker capacity and retry priorities. A schedule identifies when work becomes due; a queue manages its execution. If a nightly import could block urgent notifications, dispatch it to a separately staffed named queue and decide how missed occurrences are detected. db3.ai's Queue guide describes named queues and the Scheduler boundary.
Next step: Pick one consequential but harmless job and run the recovery checklist against your preferred architecture. If you already use db3.ai, start with the complete Queue walkthrough and then adapt the policy and replay examples above to a business key your application actually owns.
Sources
- Microsoft Learn: Best practices for background jobs
- db3.ai: Queue guide
- db3.ai: Queue API reference
- db3.ai: Build and operate queued work
- db3.ai: Move a report out of the request
- db3.ai: Retry later without using an attempt
- BullMQ: PostgreSQL backend
- BullMQ: Retrying failing jobs
- BullMQ: Worker concurrency
- pg-boss: Project documentation
- pg-boss: Job API
- pg-boss: Worker API
Recommended Reads
- 5 Background Job Frameworks Compared: Queues, Retries, Scheduling, and Recovery for a wider set of framework and platform choices.
- Background Task Scheduling: A Reliability Guide for Backend Jobs for missed-run, overlap, and schedule-to-queue decisions.