FastAPI BackgroundTasks vs a Task Queue
FastAPI BackgroundTasks run inside your web process and disappear on restart. When that is fine, when you need a real task queue, and how to move over.
You add a document endpoint. The upload is small, but indexing it means fetching the file, splitting it, calling an embeddings API a few times, and writing vectors to a store. That is several seconds of work you do not want the caller waiting on. So you reach for BackgroundTasks, return a 202, and let the indexing happen after the response goes out. It works on the first try, which is exactly why you should understand what you just signed up for.
This post is for engineers building FastAPI backends who have used BackgroundTasks once and now wonder whether it is enough. I cover where the work actually runs, the failure mode that catches people, when the built-in tool is the right call, and what a real task queue buys you when it is not.
The running example is the kind of ingestion work I dealt with on Archi, the retrieval copilot I built for CMS computing operations at CERN. There, “index this document” has to survive a deploy.
Where BackgroundTasks actually runs
BackgroundTasks comes from Starlette, the ASGI toolkit FastAPI is built on (ASGI is the standard interface between async Python servers and web apps). When you add a task, Starlette holds it until your response has been sent, and then runs it. The important question is where: the task runs in the same process, on the same event loop, as the endpoint that scheduled it.
In the endpoint below, the handler only schedules index_document and returns right away:
from fastapi import BackgroundTasks, FastAPI
app = FastAPI()
@app.post("/documents", status_code=202)
async def ingest(doc_id: str, tasks: BackgroundTasks):
tasks.add_task(index_document, doc_id)
return {"status": "queued", "doc_id": doc_id}The caller gets a 202 in a few milliseconds. index_document runs afterward, inside the Uvicorn worker that handled the request:
- If the function is
async, it runs on the event loop. - If it is a plain
def, Starlette runs it in a thread pool so it does not block the loop. This is the same threadpool behaviour I wrote about in FastAPI event loop blocking.
Either way, there is no broker and no second process to deploy. That is the appeal, and it is also the whole problem.
The failure mode: the job dies with the worker
The job lives and dies with the worker process. Nothing records that it started, nothing watches for it to finish, and nothing will run it again if it does not.
Walk through a deploy. A request comes in, you return 202, and index_document is three embedding calls deep. Then your orchestrator sends the worker a SIGTERM (the “please shut down” signal) to roll out the new version. Uvicorn tries to shut down gracefully and waits for the in-flight request coroutine, which technically includes the background task. But the client already has its 202 and has moved on, and the grace period is finite.
On Kubernetes, the default terminationGracePeriodSeconds is 30. A job that runs longer than what is left of that window gets a SIGKILL (an immediate kill the process cannot intercept), and the half-finished work goes with it. A worker crash, an out-of-memory kill, or an unhandled exception in the task drops the job with no grace at all.
Meanwhile, the client saw 202 and believes the document is indexed. It is not. There is no error anywhere the client can see, because from the HTTP layer’s point of view the request succeeded minutes ago.
The same gap shows up in smaller ways:
- Exceptions vanish. An unhandled exception inside the task is logged and then dropped. There is no retry.
- Background jobs compete with requests. During a traffic spike, every worker is also running background jobs, which compete with request handling for the same event loop and CPU.
- There is no queue to inspect. “How many documents are waiting to index?” has no answer. There is no list, only whatever happens to be mid-flight in each process.
None of this is a bug in BackgroundTasks. It does exactly what the Starlette docs say. The trouble is that once the work matters, “run this after the response” quietly turns into “run this reliably, exactly once, even across restarts.” The tool never promised the second thing.
The diagram below puts the two designs side by side.
When BackgroundTasks is the right call
Reaching for a queue too early is its own mistake. A queue means a broker to run, a worker process to deploy, and a serialization boundary to think about. Plenty of work needs none of that.
BackgroundTasks is a good fit when losing the task is annoying but not wrong:
- Firing an analytics event or a webhook, where a dropped one does not corrupt anything.
- Sending a notification email that the user can re-trigger.
- Invalidating a cache key, where the next request repopulates it anyway.
- Cleaning up a temp file after streaming a response.
The test I use: if this task silently fails once a week, do I have to explain it to someone? If the honest answer is “no one would notice,” the built-in tool is fine, so keep it. If the answer is “I would have to find and re-run it by hand,” you have outgrown it.
What a task queue gives you
A task queue splits the work across a process boundary. Your web process does one cheap thing: it serializes the job’s name and arguments and hands them to a broker (a separate service that stores jobs until a worker takes them), usually Redis or RabbitMQ. A separate pool of worker processes pulls jobs off the broker and runs them. The web replica and the worker no longer share a fate.
That boundary buys you what BackgroundTasks cannot offer:
- Jobs survive deploys. Once a job is in the broker, a web deploy is irrelevant to it. A worker picks it up whenever one is free.
- Failed jobs retry. A job that raises can be retried with backoff. After a few failed attempts, it lands in a dead-letter list (a holding area for jobs that keep failing) instead of vanishing.
- The backlog is visible. The broker holds a real queue you can inspect, so “how far behind are we?” becomes a number you can measure and alert on.
- Web and workers scale separately. Ingestion is heavy and bursty, while request handling is light and steady. Separate workers let you scale each on its own signal, instead of overprovisioning web replicas to soak up background load. On Kubernetes, the worker pool is just another deployment you can scale on queue depth.
The cost is real, and worth stating. You now operate a broker. You deploy and monitor a worker process. And arguments cross a serialization boundary, so you pass a doc_id, not a live database handle, and the worker loads what it needs itself.
A minimal queue with Arq
For an async FastAPI app, I default to Arq, a small Redis-backed queue from the author of Pydantic. It speaks async/await natively, so the worker code looks like the rest of a FastAPI codebase instead of forcing a sync context. Celery is the older, heavier standard, with more features and more operational surface. RQ and Dramatiq sit in between. The shape of the code below is the same across all of them.
The worker file defines the job and its retry policy. Note max_tries, which caps the attempts, and job_timeout, which stops a stuck job:
# worker.py
from arq.connections import RedisSettings
async def index_document(ctx, doc_id: str):
doc = await load_document(doc_id)
chunks = split(doc)
vectors = await embed(chunks) # the slow, retry-worthy part
await upsert(doc_id, vectors)
class WorkerSettings:
functions = [index_document]
redis_settings = RedisSettings(host="redis")
max_tries = 4 # retry with backoff before giving up
job_timeout = 300 # seconds, so a wedged job cannot run foreverThe endpoint stops doing the work and only enqueues it. Create the Redis pool once at startup and reuse it, rather than opening a connection per request:
# app.py
from arq import create_pool
from arq.connections import RedisSettings
from fastapi import FastAPI
app = FastAPI()
@app.on_event("startup")
async def startup():
app.state.redis = await create_pool(RedisSettings(host="redis"))
@app.post("/documents", status_code=202)
async def ingest(doc_id: str):
await app.state.redis.enqueue_job("index_document", doc_id)
return {"status": "queued", "doc_id": doc_id}Run the worker as its own process, arq worker.WorkerSettings, in its own container or pod. The endpoint returns as soon as the job is in Redis, so the client still gets a fast 202. The difference is that the 202 now means something durable: the job is recorded, and if the first attempt fails, it will be tried again.
Tradeoffs, and where I draw the line
You can go wrong in two opposite directions. One is running every trivial fire-and-forget through Celery because a blog post said queues are the correct answer. The other is pushing money movement or document indexing through BackgroundTasks because adding a broker felt like too much on a Friday.
How I decide:
- Can the task be lost without anyone noticing or any data going wrong? Use
BackgroundTasks. - Does the task need to survive a deploy, retry on failure, or be countable when it backs up? Use a queue.
- Not sure yet? Start with
BackgroundTasks, and keep the task function pure: it takes an id and loads its own data. Moving it behindenqueue_joblater is then a small change, because the function already does not depend on request state.
One trap deserves a callout. BackgroundTasks lulls you because it works perfectly in development. You never restart mid-job on your laptop, so you never see the loss. It shows up in production on your first busy deploy, as a handful of documents that report success but never got indexed. If you cannot afford that kind of silent gap, the boundary between web and worker is what removes it. No amount of care inside a single process will.
Closing: what happens when the process goes away
The choice is not “which tool is better.” It is “what happens to this job when the process it started in goes away?” If the answer can be nothing, BackgroundTasks is the least machinery that does the job. If the answer has to be it still runs, you want the work on the far side of a broker.
On Archi, indexing new operations documents falls squarely in the second bucket. An operator asks Archi about an incident and expects the relevant logbook entries to be searchable; a document that silently failed to index is a wrong answer with no error attached. That is the line for me: the moment “it probably ran” is not good enough, the work belongs in a queue. You can see how the retrieval side of that system fits together in the Archi project write-up, and more of the backend patterns I lean on across my other projects.