30 September 2026
5 Background Job Frameworks Compared: Queues, Retries, Scheduling, and Recovery

A background job framework should do more than run code after an HTTP response. It needs to preserve work, coordinate workers, expose failure states, and give operators a way to recover. The best choice depends on whether you want to operate a queue datastore yourself, use a managed execution platform, or keep jobs inside a broader application runtime.
We compare five real options for product engineers building TypeScript applications: db3.ai, BullMQ, pg-boss, Trigger.dev, and Inngest. This is a framework-selection guide, not a claim that all five have identical storage models or execution guarantees. Microsoft’s background job design guidance highlights the questions that matter after launch: duplicate execution, missed schedules, completion visibility, and independent scaling.
Table of contents
- How we compared the frameworks
- 1. db3.ai: Background jobs inside an application runtime
- 2. BullMQ: Queue primitives with a choice of backends
- 3. pg-boss: PostgreSQL-centered job processing
- 4. Trigger.dev: Hosted or self-hosted task execution
- 5. Inngest: Event-driven functions with durable steps
- Test the failure path before you choose
- Which approach fits your application?
- Frequently asked questions
How we compared the frameworks
We selected options that can handle background work in a Node.js or TypeScript product and represent three operational approaches: an integrated application runtime, self-operated queue libraries, and coordinated execution platforms. Each listing addresses the same questions where the product documentation supplies an answer:
| Criterion | Ask this before choosing |
|---|---|
| Queue and payload | Where does work persist, how is it identified, and what can the payload safely contain? |
| Workers and isolation | Who runs the code, and can a slow queue be separated from urgent work? |
| Retries and recovery | What counts as an attempt, where do terminal failures go, and can an operator replay one? |
| Scheduling | Does a recurring trigger create durable work, and how are missed or duplicate occurrences handled? |
| Visibility and operations | Can you inspect attempts, waiting time, failures, and deployment or infrastructure requirements? |

An event, a scheduler, and a queue are different components. A scheduler decides when work is due; a queue persists work until a worker claims it. A platform may hide one or both components behind a function API. Whatever the interface, ask what happens between a successful external side effect and a worker crash. A retry may repeat that side effect, so business-level idempotency still matters. Microsoft explains the duplicate-execution risk.
1. db3.ai: Background jobs inside an application runtime

Best fit: A TypeScript product that wants application services, durable jobs, and scheduling to share an app bootstrap rather than connecting unrelated libraries. The db3.ai Queue guide documents database and Redis queue drivers, named queues, JSON-safe QueueableJob payloads, explicit worker registration, delayed dispatch, and retry policies. It also makes an important distinction: dispatch returns a stored job ID, not proof that the work has finished.
On failure, the queue can apply linear or exponential backoff with a retry window. It retains terminal failures for inspection and supports replay as a new job linked to the original failure. A deliberate QueueRetryLaterError can defer work without spending an ordinary attempt, but the application must enforce a deadline to prevent endless deferral. The queue API reference defines the attempt outcomes, while the queued-work example exercises replay and backpressure. These are documented features, not an exactly-once side-effect guarantee.
Here is a small selection test: can we express a report as a durable identity rather than serializing a database connection or a request object? In an installed db3 app, a job can carry a report ID and resolve services when a worker handles it:
// server/jobs/GenerateReportJob.ts
import { QueueableJob } from '@db3.ai/app/queue';
import { app } from '@db3.ai/app/server';
export class GenerateReportJob extends QueueableJob<{ reportId: string }> {
static readonly jobName = 'reports.generate.v1';
constructor(data: { reportId: string }) {
if (!data.reportId?.trim()) throw new Error('reportId is required');
super(data);
}
async handle(): Promise<void> {
app().log.info({ reportId: this.data.reportId }, 'Generate report');
// Replace logging with an idempotent report writer keyed by reportId.
}
}
Register GenerateReportJob in the shared application factory used by both producer and worker. Then, from a route or service after boot, dispatch it with an explicit policy:
import { app } from '@db3.ai/app/server';
import { GenerateReportJob } from './jobs/GenerateReportJob';
const jobId = await app().queue.dispatch(
new GenerateReportJob({ reportId: 'weekly' }),
{
queue: 'reports',
maxTries: 5,
backoff: {
strategy: 'exponential',
initialSeconds: 15,
maxSeconds: 900,
jitter: true,
},
retryUntilSeconds: 3600,
},
);
This illustration logs rather than generating a report. It also assumes queue tables have been migrated and a reports worker is running. Use the complete Queue setup for those steps, rather than treating the snippet as a standalone deployment.
For recurring work, the Scheduler guide describes named daily, hourly, and minute schedules that dispatch queue jobs and record occurrences. It does not document arbitrary cron or weekly/monthly schedules. Durable catch-up across scheduler restarts requires an application-owned checkpoint; operators should reconcile a claimed occurrence if dispatch fails between the claim and enqueue. This is the chief trade-off of the integrated approach: fewer separate application APIs, but you still own migrations, supervised scheduler and worker processes, capacity planning, and idempotent business behavior. The package documentation describes preview installation through matching packages or maintainer-supplied tarballs, so check the current installation route before committing to a rollout.
2. BullMQ: Queue primitives with a choice of backends

Best fit: A team that wants a dedicated Node.js queue API and controls how its workers are deployed. The BullMQ overview documents queues, workers, delayed jobs, retries, concurrency, and dependent flows. Redis is its default backend, but describing BullMQ as Redis-only is no longer accurate: its optional PostgreSQL backend exposes the Queue, Worker, QueueEvents, and FlowProducer APIs on PostgreSQL. The BullMQ documentation calls Redis the more battle-tested backend. Evaluate the backend you will actually use, including its connection and migration requirements.
BullMQ supports retry attempts with fixed or exponential backoff and optional jitter. A job that stalls can fail and be retried, so a payment or email worker still needs an application idempotency key. You control worker concurrency and should decide which job types merit separate queues and processes. For recurring work, current Job Schedulers create delayed jobs; their documented production rate can slip when workers are busy. That behavior matters if your requirement is an exact wall-clock completion time rather than a recurring enqueue target. Retry behavior and scheduler behavior should both be tested against your workload.
The operational trade-off is flexibility versus ownership. Plan for datastore persistence and backups, worker deployment, failed-job retention, and a monitoring interface. BullMQ exposes events and metrics, but deciding how to alert and investigate is still part of your system design. Choose it when queue control is the central need and you are comfortable integrating it with your application's models, permissions, and schedules.
3. pg-boss: PostgreSQL-centered job processing

Best fit: A Node.js application that already operates PostgreSQL and wants a job system built around it. The pg-boss project documentation lists delayed jobs, cron and RRULE scheduling, retries with exponential backoff, dead-letter queues and redrive, queue policies, and a separate dashboard package. It also describes enqueueing within an existing database transaction through supported adapters. That transaction boundary can matter when creating a business record and its follow-up work must succeed or fail together.
Workers still need running processes, suitable database connection limits, and a plan for long-running tasks and queue growth. PostgreSQL is familiar infrastructure, not an absence of infrastructure. Ask how job tables, migrations, retention, and vacuuming will fit your database operations; then test a worker termination while a job is active. The job API documents retry options and retrying failed jobs, which gives an operator a more concrete recovery path than simply checking logs.
The distinction from BullMQ’s optional PostgreSQL backend is product architecture, not a claim that only one can use SQL. pg-boss is designed around PostgreSQL from the outset; BullMQ offers the same queue-facing API across its documented Redis and PostgreSQL backends. Pick based on the behavior and integration you validate, not on a one-word storage label.
4. Trigger.dev: Hosted or self-hosted task execution

Best fit: TypeScript teams that want durable tasks, queues, schedules, retries, and a run interface without building their own queue-worker stack for a cloud deployment. Trigger.dev’s documentation describes SDK-defined tasks, configurable concurrency, cron schedules, recorded runs and logs, plus cloud and self-hosting options. A task can run in a separate execution environment, so deployment configuration and data access still deserve design work.
Retry settings can be configured for a project or an individual task; the error and retry guide also describes inspecting logs and retrying smaller tasks independently. That visibility is especially useful for long-running AI processing, where one opaque catch-all job makes diagnosis difficult. Verify which state you need to keep in your own product database for user-facing progress and authorization instead of assuming the platform dashboard is your customer UI.
The operational comparison changes with the hosting choice. On cloud, account for platform dependency and execution usage; with self-hosting, your team takes on the platform deployment as well. Evaluate the complete cost and execution limits for your workload rather than comparing only library installation effort. Choose Trigger.dev when run-level control and a TypeScript task platform outweigh the desire to keep coordination entirely inside your application runtime.
5. Inngest: Event-driven functions with durable steps

Best fit: An application where events or schedules initiate multi-step work and step-level retry matters more than direct queue-record control. Inngest Functions are triggered by events, cron schedules, or webhooks; function code runs on your compute while Inngest coordinates execution. Its documented step model stores completed steps so a later failure can retry from a checkpoint instead of rerunning every previous step.
This is a different mental model from pushing a named payload onto your own database queue. Define event identity, concurrency and throttling policy, then decide what application state needs to remain in your database. Inngest’s observability documentation describes function-run metrics, event logs, and step-level visibility. You still need to reason about a step that performs an irreversible external action before its completion is recorded. A durable checkpoint does not make that external system idempotent.
The main trade-off is adopting a coordination platform and its event/function model instead of operating a conventional queue directly. Compare that integration and service dependency with your own compute and monitoring costs. Choose Inngest when your real unit of work is an event-driven workflow whose steps need independent retry and inspection.
Test the failure path before you choose
A feature checklist can make every framework look capable. We would run the same small test in each finalist: accept a report request, persist a job identity, stop its worker after an external write but before completion, restart it, and show both the customer-facing outcome and the operator's recovery steps. That test exposes duplicate writes and ambiguous completion without requiring a synthetic throughput contest.
In the trial, record: enqueue-to-start wait time, attempts consumed, retry delay, terminal failure location, and the command or UI used to replay one job. Repeat with a malformed payload, a missing worker for a named queue, a rate-limited provider, and a scheduler restarted across a missed minute. If scheduled output must run once per business day, also test timezone and daylight-saving boundaries. Microsoft’s guidance recommends tracking completion and missed schedules, not merely task starts.
Consider deployability alongside correctness. Can a worker running yesterday's code safely read today's queued payload? Where are migrations applied? Can one noisy AI generation queue delay billing notifications? What alert fires if work remains pending even though no process throws an error? A framework is operationally complete for your application only when the answers are written down and tested. For db3.ai specifically, the queued-work walkthrough provides a starting point for retry, backpressure, and replay experiments rather than a guarantee for every provider or deployment.
Which approach fits your application?
| Situation | Start your evaluation with | Why |
|---|---|---|
| One TypeScript app needs jobs alongside its application services | db3.ai | Queue and scheduler share application concepts, with deployment still under your control. |
| You want a dedicated queue library and backend flexibility | BullMQ | Queue primitives with Redis by default and a documented PostgreSQL alternative. |
| PostgreSQL transactions are central to enqueueing | pg-boss | PostgreSQL-first jobs with transaction adapters and operational queue policies. |
| You want a task execution platform and run-level UI | Trigger.dev | Task-oriented retries, schedules, and observability, with cloud or self-hosting choices. |
| Events trigger workflows with independently retriable steps | Inngest | Event/function coordination and persisted step checkpoints. |
This is a shortlist, not a ranking by reliability. No framework substitutes for application-owned idempotency, access control, and a recovery procedure. Choose the one whose failure model you can explain to an on-call engineer and whose infrastructure your team can maintain.
Frequently asked questions
Is a scheduler the same as a job queue?
No. A scheduler determines when a task is due; a queue stores work for a worker to process. When a scheduled task enqueues a job, inspect both its occurrence history and the queued job's outcome. A successful enqueue is not a successful report. The db3.ai Scheduler guide makes this distinction explicit.
Does a durable queue mean each job executes exactly once?
No. A crash after an external side effect and before a completion record can lead to another attempt. Use a stable business key, make writes repeatable when possible, and reconcile ambiguous external outcomes before replaying. Microsoft’s background job guidance explains why duplicate execution must be part of the design.
If you are evaluating an integrated runtime, start with the db3.ai Queue guide, register one real job in a shared app bootstrap, and run its failure and replay paths against a disposable database. The goal is not to find the longest feature list. It is to know precisely what your users and operators will see when a job succeeds, waits, or fails.