{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-08-04-fastapi-background-tasks-vs-task-queue/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"5e10ccb3-9f3e-54f2-bed6-38b9b0617a1a","excerpt":"You add a document endpoint. The upload is small, but indexing it means fetching the file, splitting it, calling an embeddings API a few times, and writing…","html":"<p>You add a document endpoint. The upload is small, but indexing it means fetching the file, splitting it, calling an embeddings API a few times, and writing vectors to a store. That is several seconds of work you do not want the caller waiting on. So you reach for <a href=\"https://fastapi.tiangolo.com/tutorial/background-tasks/\"><code class=\"language-text\">BackgroundTasks</code></a>, return a <code class=\"language-text\">202</code>, and let the indexing happen after the response goes out. It works on the first try, which is exactly why you should understand what you just signed up for.</p>\n<p>This post is for engineers building <a href=\"https://fastapi.tiangolo.com/\">FastAPI</a> backends who have used <code class=\"language-text\">BackgroundTasks</code> once and now wonder whether it is enough. I cover where the work actually runs, the failure mode that catches people, when the built-in tool is the right call, and what a real task queue buys you when it is not.</p>\n<p>The running example is the kind of ingestion work I dealt with on <a href=\"/project/archi/\">Archi</a>, the retrieval copilot I built for CMS computing operations at CERN. There, “index this document” has to survive a deploy.</p>\n<h2>Where BackgroundTasks actually runs</h2>\n<p><code class=\"language-text\">BackgroundTasks</code> comes from <a href=\"https://www.starlette.io/background/\">Starlette</a>, the ASGI toolkit FastAPI is built on (ASGI is the standard interface between async Python servers and web apps). When you add a task, Starlette holds it until your response has been sent, and then runs it. The important question is <em>where</em>: the task runs in the same process, on the same event loop, as the endpoint that scheduled it.</p>\n<p>In the endpoint below, the handler only schedules <code class=\"language-text\">index_document</code> and returns right away:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> fastapi <span class=\"token keyword\">import</span> BackgroundTasks<span class=\"token punctuation\">,</span> FastAPI\n\napp <span class=\"token operator\">=</span> FastAPI<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token decorator annotation punctuation\">@app<span class=\"token punctuation\">.</span>post</span><span class=\"token punctuation\">(</span><span class=\"token string\">\"/documents\"</span><span class=\"token punctuation\">,</span> status_code<span class=\"token operator\">=</span><span class=\"token number\">202</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">ingest</span><span class=\"token punctuation\">(</span>doc_id<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">,</span> tasks<span class=\"token punctuation\">:</span> BackgroundTasks<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    tasks<span class=\"token punctuation\">.</span>add_task<span class=\"token punctuation\">(</span>index_document<span class=\"token punctuation\">,</span> doc_id<span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">return</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\"status\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"queued\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"doc_id\"</span><span class=\"token punctuation\">:</span> doc_id<span class=\"token punctuation\">}</span></code></pre></div>\n<p>The caller gets a <code class=\"language-text\">202</code> in a few milliseconds. <code class=\"language-text\">index_document</code> runs afterward, inside the Uvicorn worker that handled the request:</p>\n<ul>\n<li>If the function is <code class=\"language-text\">async</code>, it runs on the event loop.</li>\n<li>If it is a plain <code class=\"language-text\">def</code>, Starlette runs it in a thread pool so it does not block the loop. This is the same threadpool behaviour I wrote about in <a href=\"/blog/2026-07-10-fastapi-blocking-event-loop/\">FastAPI event loop blocking</a>.</li>\n</ul>\n<p>Either way, there is no broker and no second process to deploy. That is the appeal, and it is also the whole problem.</p>\n<h2>The failure mode: the job dies with the worker</h2>\n<p>The job lives and dies with the worker process. Nothing records that it started, nothing watches for it to finish, and nothing will run it again if it does not.</p>\n<p>Walk through a deploy. A request comes in, you return <code class=\"language-text\">202</code>, and <code class=\"language-text\">index_document</code> is three embedding calls deep. Then your orchestrator sends the worker a <code class=\"language-text\">SIGTERM</code> (the “please shut down” signal) to roll out the new version. Uvicorn tries to shut down gracefully and waits for the in-flight request coroutine, which technically includes the background task. But the client already has its <code class=\"language-text\">202</code> and has moved on, and the grace period is finite.</p>\n<p>On Kubernetes, the default <code class=\"language-text\">terminationGracePeriodSeconds</code> is 30. A job that runs longer than what is left of that window gets a <code class=\"language-text\">SIGKILL</code> (an immediate kill the process cannot intercept), and the half-finished work goes with it. A worker crash, an out-of-memory kill, or an unhandled exception in the task drops the job with no grace at all.</p>\n<p>Meanwhile, the client saw <code class=\"language-text\">202</code> and believes the document is indexed. It is not. There is no error anywhere the client can see, because from the HTTP layer’s point of view the request succeeded minutes ago.</p>\n<p>The same gap shows up in smaller ways:</p>\n<ul>\n<li><strong>Exceptions vanish.</strong> An unhandled exception inside the task is logged and then dropped. There is no retry.</li>\n<li><strong>Background jobs compete with requests.</strong> During a traffic spike, every worker is also running background jobs, which compete with request handling for the same event loop and CPU.</li>\n<li><strong>There is no queue to inspect.</strong> “How many documents are waiting to index?” has no answer. There is no list, only whatever happens to be mid-flight in each process.</li>\n</ul>\n<p>None of this is a bug in <code class=\"language-text\">BackgroundTasks</code>. It does exactly what the <a href=\"https://www.starlette.io/background/\">Starlette docs</a> say. The trouble is that once the work matters, “run this after the response” quietly turns into “run this reliably, exactly once, even across restarts.” The tool never promised the second thing.</p>\n<p>The diagram below puts the two designs side by side.</p>\n<p><img src=\"/88562bc888fa508de8c200cff5cb52d5/queue-architecture.svg\" alt=\"Left: BackgroundTasks run inside the Uvicorn worker after the response returns, so a restart or crash kills the job with no retry and no record. Right: a task queue has the web process enqueue the job to a Redis broker, which a separate worker process pulls from, so the job survives the request and the deploy, failed jobs retry, and the backlog is visible.\"></p>\n<h2>When BackgroundTasks is the right call</h2>\n<p>Reaching for a queue too early is its own mistake. A queue means a broker to run, a worker process to deploy, and a serialization boundary to think about. Plenty of work needs none of that.</p>\n<p><code class=\"language-text\">BackgroundTasks</code> is a good fit when losing the task is annoying but not wrong:</p>\n<ul>\n<li>Firing an analytics event or a webhook, where a dropped one does not corrupt anything.</li>\n<li>Sending a notification email that the user can re-trigger.</li>\n<li>Invalidating a cache key, where the next request repopulates it anyway.</li>\n<li>Cleaning up a temp file after streaming a response.</li>\n</ul>\n<p>The test I use: if this task silently fails once a week, do I have to explain it to someone? If the honest answer is “no one would notice,” the built-in tool is fine, so keep it. If the answer is “I would have to find and re-run it by hand,” you have outgrown it.</p>\n<h2>What a task queue gives you</h2>\n<p>A task queue splits the work across a process boundary. Your web process does one cheap thing: it serializes the job’s name and arguments and hands them to a broker (a separate service that stores jobs until a worker takes them), usually <a href=\"https://redis.io/\">Redis</a> or <a href=\"https://www.rabbitmq.com/\">RabbitMQ</a>. A separate pool of worker processes pulls jobs off the broker and runs them. The web replica and the worker no longer share a fate.</p>\n<p>That boundary buys you what <code class=\"language-text\">BackgroundTasks</code> cannot offer:</p>\n<ul>\n<li><strong>Jobs survive deploys.</strong> Once a job is in the broker, a web deploy is irrelevant to it. A worker picks it up whenever one is free.</li>\n<li><strong>Failed jobs retry.</strong> A job that raises can be retried with backoff. After a few failed attempts, it lands in a dead-letter list (a holding area for jobs that keep failing) instead of vanishing.</li>\n<li><strong>The backlog is visible.</strong> The broker holds a real queue you can inspect, so “how far behind are we?” becomes a number you can measure and alert on.</li>\n<li><strong>Web and workers scale separately.</strong> Ingestion is heavy and bursty, while request handling is light and steady. Separate workers let you scale each on its own signal, instead of overprovisioning web replicas to soak up background load. On <a href=\"/blog/2026-07-15-kubernetes-oomkilled-requests-vs-limits/\">Kubernetes</a>, the worker pool is just another deployment you can scale on queue depth.</li>\n</ul>\n<p>The cost is real, and worth stating. You now operate a broker. You deploy and monitor a worker process. And arguments cross a serialization boundary, so you pass a <code class=\"language-text\">doc_id</code>, not a live database handle, and the worker loads what it needs itself.</p>\n<h2>A minimal queue with Arq</h2>\n<p>For an async FastAPI app, I default to <a href=\"https://arq-docs.helpmanual.io/\">Arq</a>, a small Redis-backed queue from the author of Pydantic. It speaks <code class=\"language-text\">async</code>/<code class=\"language-text\">await</code> natively, so the worker code looks like the rest of a FastAPI codebase instead of forcing a sync context. <a href=\"https://docs.celeryq.dev/\">Celery</a> is the older, heavier standard, with more features and more operational surface. <a href=\"https://python-rq.org/\">RQ</a> and <a href=\"https://dramatiq.io/\">Dramatiq</a> sit in between. The shape of the code below is the same across all of them.</p>\n<p>The worker file defines the job and its retry policy. Note <code class=\"language-text\">max_tries</code>, which caps the attempts, and <code class=\"language-text\">job_timeout</code>, which stops a stuck job:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token comment\"># worker.py</span>\n<span class=\"token keyword\">from</span> arq<span class=\"token punctuation\">.</span>connections <span class=\"token keyword\">import</span> RedisSettings\n\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">index_document</span><span class=\"token punctuation\">(</span>ctx<span class=\"token punctuation\">,</span> doc_id<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    doc <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> load_document<span class=\"token punctuation\">(</span>doc_id<span class=\"token punctuation\">)</span>\n    chunks <span class=\"token operator\">=</span> split<span class=\"token punctuation\">(</span>doc<span class=\"token punctuation\">)</span>\n    vectors <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> embed<span class=\"token punctuation\">(</span>chunks<span class=\"token punctuation\">)</span>      <span class=\"token comment\"># the slow, retry-worthy part</span>\n    <span class=\"token keyword\">await</span> upsert<span class=\"token punctuation\">(</span>doc_id<span class=\"token punctuation\">,</span> vectors<span class=\"token punctuation\">)</span>\n\n<span class=\"token keyword\">class</span> <span class=\"token class-name\">WorkerSettings</span><span class=\"token punctuation\">:</span>\n    functions <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span>index_document<span class=\"token punctuation\">]</span>\n    redis_settings <span class=\"token operator\">=</span> RedisSettings<span class=\"token punctuation\">(</span>host<span class=\"token operator\">=</span><span class=\"token string\">\"redis\"</span><span class=\"token punctuation\">)</span>\n    max_tries <span class=\"token operator\">=</span> <span class=\"token number\">4</span>          <span class=\"token comment\"># retry with backoff before giving up</span>\n    job_timeout <span class=\"token operator\">=</span> <span class=\"token number\">300</span>      <span class=\"token comment\"># seconds, so a wedged job cannot run forever</span></code></pre></div>\n<p>The endpoint stops doing the work and only enqueues it. Create the Redis pool once at startup and reuse it, rather than opening a connection per request:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token comment\"># app.py</span>\n<span class=\"token keyword\">from</span> arq <span class=\"token keyword\">import</span> create_pool\n<span class=\"token keyword\">from</span> arq<span class=\"token punctuation\">.</span>connections <span class=\"token keyword\">import</span> RedisSettings\n<span class=\"token keyword\">from</span> fastapi <span class=\"token keyword\">import</span> FastAPI\n\napp <span class=\"token operator\">=</span> FastAPI<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token decorator annotation punctuation\">@app<span class=\"token punctuation\">.</span>on_event</span><span class=\"token punctuation\">(</span><span class=\"token string\">\"startup\"</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">startup</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    app<span class=\"token punctuation\">.</span>state<span class=\"token punctuation\">.</span>redis <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> create_pool<span class=\"token punctuation\">(</span>RedisSettings<span class=\"token punctuation\">(</span>host<span class=\"token operator\">=</span><span class=\"token string\">\"redis\"</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token decorator annotation punctuation\">@app<span class=\"token punctuation\">.</span>post</span><span class=\"token punctuation\">(</span><span class=\"token string\">\"/documents\"</span><span class=\"token punctuation\">,</span> status_code<span class=\"token operator\">=</span><span class=\"token number\">202</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">ingest</span><span class=\"token punctuation\">(</span>doc_id<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">await</span> app<span class=\"token punctuation\">.</span>state<span class=\"token punctuation\">.</span>redis<span class=\"token punctuation\">.</span>enqueue_job<span class=\"token punctuation\">(</span><span class=\"token string\">\"index_document\"</span><span class=\"token punctuation\">,</span> doc_id<span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">return</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\"status\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"queued\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"doc_id\"</span><span class=\"token punctuation\">:</span> doc_id<span class=\"token punctuation\">}</span></code></pre></div>\n<p>Run the worker as its own process, <code class=\"language-text\">arq worker.WorkerSettings</code>, in its own container or pod. The endpoint returns as soon as the job is in Redis, so the client still gets a fast <code class=\"language-text\">202</code>. The difference is that the <code class=\"language-text\">202</code> now means something durable: the job is recorded, and if the first attempt fails, it will be tried again.</p>\n<h2>Tradeoffs, and where I draw the line</h2>\n<p>You can go wrong in two opposite directions. One is running every trivial fire-and-forget through Celery because a blog post said queues are the correct answer. The other is pushing money movement or document indexing through <code class=\"language-text\">BackgroundTasks</code> because adding a broker felt like too much on a Friday.</p>\n<p>How I decide:</p>\n<ul>\n<li><strong>Can the task be lost without anyone noticing or any data going wrong?</strong> Use <code class=\"language-text\">BackgroundTasks</code>.</li>\n<li><strong>Does the task need to survive a deploy, retry on failure, or be countable when it backs up?</strong> Use a queue.</li>\n<li><strong>Not sure yet?</strong> Start with <code class=\"language-text\">BackgroundTasks</code>, and keep the task function pure: it takes an id and loads its own data. Moving it behind <code class=\"language-text\">enqueue_job</code> later is then a small change, because the function already does not depend on request state.</li>\n</ul>\n<p>One trap deserves a callout. <code class=\"language-text\">BackgroundTasks</code> lulls you because it works perfectly in development. You never restart mid-job on your laptop, so you never see the loss. It shows up in production on your first busy deploy, as a handful of documents that report success but never got indexed. If you cannot afford that kind of silent gap, the boundary between web and worker is what removes it. No amount of care inside a single process will.</p>\n<h2>Closing: what happens when the process goes away</h2>\n<p>The choice is not “which tool is better.” It is “what happens to this job when the process it started in goes away?” If the answer can be <em>nothing</em>, <code class=\"language-text\">BackgroundTasks</code> is the least machinery that does the job. If the answer has to be <em>it still runs</em>, you want the work on the far side of a broker.</p>\n<p>On Archi, indexing new operations documents falls squarely in the second bucket. An operator asks Archi about an incident and expects the relevant logbook entries to be searchable; a document that silently failed to index is a wrong answer with no error attached. That is the line for me: the moment “it probably ran” is not good enough, the work belongs in a queue. You can see how the retrieval side of that system fits together in the <a href=\"/project/archi/\">Archi project write-up</a>, and more of the backend patterns I lean on across my <a href=\"/\">other projects</a>.</p>","frontmatter":{"title":"FastAPI BackgroundTasks vs a Task Queue","date":"2026-08-04T00:00:00.000Z","description":"FastAPI BackgroundTasks run inside your web process and disappear on restart. When that is fine, when you need a real task queue, and how to move over.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAACBklEQVQoz2WRy27aUBCGXVUFbOO7fW62zzHY+AbYmEAugEJCqEqV0CbqqtuqfYJK3XdTVepDlF1eswPdtdKn0SzmP/PPfyTirygQQF165HKQvambp2G1hzqq3kEznryvmkegmX2I+htEr4i/BAkIJcX0ZYO1ulhzhUViPmjKZpWU58NmlVVXwGB4UdTLcrqq5msiSoskBuqBBIQgDi2c0HDIRImCTHNE1+ZdR8gma+sUhmQjhBmoshF0baE5PdUWqsVVS0iK4ds0jcerwXgTpWsaLVhvCWB+iYI5YhPklw6DhdlfPJojv4BGcyJJ0X3Ni51JNVt8e3x43t//fn17uFsftjeH7eb58uI7Ckc4OmP9ORJTwusmm2XDlZ/Mda8vyToDqx7LbTwwUQwYXmJ4sWKJiGUXyWiRjjHqKTbXwfPRtgDbuhupZiDJXWKTQX52I7IrImY0mgOYNyYqPDoUYR3xqUsrAxUmzk1cHEG5jlI4W5I1Ai85eRUkm7R4irM90E8fePw27O9EvIviHY/vCV8qLldhoRdpXt/EJYQndeCTbAHHuMcY0iM0s09AA+dgvzBJEvTWdfN5Mv1UT77ko4/h5FyzOYhR1wohT5vm/2DR3GUFpG3ihInF5vrn9vbH5vpXM/2KylqzwqP4BAbkU/0fyKWtulKrI73qvGi1X3YUVWOyhv8AdlFXAfSvsIkAAAAASUVORK5CYII=","aspectRatio":1.899441340782123,"src":"/static/3e4ce285c6ea6a6c3330336b6ea44ef3/40a76/hero.png","srcSet":"/static/3e4ce285c6ea6a6c3330336b6ea44ef3/c972b/hero.png 340w,\n/static/3e4ce285c6ea6a6c3330336b6ea44ef3/27625/hero.png 680w,\n/static/3e4ce285c6ea6a6c3330336b6ea44ef3/40a76/hero.png 1360w,\n/static/3e4ce285c6ea6a6c3330336b6ea44ef3/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-08-04-fastapi-background-tasks-vs-task-queue/","previous":"blog/2026-08-05-late-interaction-retrieval-colbert-rag/","next":"blog/2026-08-07-argocd-applicationsets-many-clusters/"}},"staticQueryHashes":["32046230"]}