FastAPI Event Loop Blocking: Sync vs Async

One blocking call in a FastAPI route stalls every other request, including live SSE streams. Here is how the event loop breaks, and how to keep it free.

The bug looks impossible the first time you see it. Your FastAPI service handles hundreds of requests a second in load tests. Then someone adds one endpoint that resizes an image or verifies a JWT with a slow library, and suddenly every request gets slow, not just that one. Health checks time out. A streaming chat that was landing tokens in 200 ms now hangs for seconds. Nothing changed in the other routes, but they all got worse together.

The cause is almost always the same: one blocking call sitting on the event loop. This post is for engineers running an async FastAPI backend who have hit that “why is the whole server slow” moment, or want to avoid it. It covers how the event loop works, how FastAPI decides where your route runs, a repro you can run on your laptop, and the right fix for each kind of slow work.

I hit this building the FastAPI backends behind CloudCanvasAI and Archi. There, a stalled event loop does not just slow one request down; it freezes the live SSE stream (Server-Sent Events) that every user is watching.

One thread runs every request

FastAPI is async, and async in Python means one event loop running on one thread per worker process. That loop is not doing many things at once. It does one thing at a time and switches between them very fast, at the points where your code says await.

That is the whole contract. When an async def route hits await client.get(url), it tells the loop “I am waiting on the network; go run someone else.” The loop parks that request, picks up another, and comes back when the network call is ready. A hundred requests can be in flight because, at any instant, ninety-nine of them are parked at an await and only one is actually executing.

A single event loop serves every request by interleaving them at await points; when one request makes a blocking call with no await, the loop cannot switch and every other request, including live SSE streams, stalls until it returns

The failure mode falls straight out of that design. If one request runs code that takes 500 ms and never yields, the loop cannot switch away from it, because there is no await to hand control back. For that half second the single thread is busy, and every other request, no matter how cheap, waits in line behind it. You did not slow down one endpoint. You paused the server.

The signature you write decides where code runs

FastAPI gives you a way out of this, and most people use it without knowing. It checks whether your path operation (the route function) is def or async def and schedules the two differently. Starlette, the framework FastAPI sits on, does the actual work. There are three cases:

Three ways a FastAPI route is scheduled: async def with await runs on the event loop and is safe; a plain def route is offloaded to a threadpool and is safe until the pool fills; an async def containing a blocking call runs on the loop and holds it, which is the trap

  • async def + real await. Runs directly on the event loop. This is correct and concurrent, as long as every slow thing inside it is actually awaited.
  • Plain def. FastAPI runs it in a threadpool instead of on the loop, using anyio under the hood. Your blocking requests.get() still blocks, but it blocks a worker thread, not the loop, so other requests keep flowing. The catch is that the pool is bounded: Starlette’s default limiter is 40 threads, and once they are all busy, new sync requests queue.
  • async def with a blocking call inside. This is the trap. You wrote async def, so FastAPI trusts you and runs it on the loop. Then you called something synchronous, like requests.get() or a CPU-heavy function, with no await. Now you have the worst of both: on the loop, and holding it.

The rule that falls out of this: async def is a promise that you will not block. If you cannot keep that promise for a given call, either change the route to plain def or offload the slow part explicitly.

A repro you can run

Here is the trap in about twenty lines. time.sleep stands in for any blocking work: a synchronous DB driver, a hashing round, a PIL resize, a requests call.

import time
from fastapi import FastAPI

app = FastAPI()

@app.get("/blocking")
async def blocking():
    time.sleep(1)          # blocks the event loop for a full second
    return {"ok": True}

@app.get("/health")
async def health():
    return {"status": "up"}

Run it with uvicorn main:app, then fire ten /blocking requests at once while timing /health:

# hammer the blocking route
for i in $(seq 10); do curl -s localhost:8000/blocking & done

# meanwhile, time a trivial health check
time curl -s localhost:8000/health

/health does nothing and should answer instantly. Instead it waits behind the queue of time.sleep calls, because all of them are stacked on the one loop thread. With ten blocking requests, /health can take several seconds. That is the exact shape of the production incident, reproduced on your laptop.

Now change one word on the blocking route, async def to def, and run it again. /health stays fast under the same load, because FastAPI pushed the def route onto the threadpool and left the loop free. Same blocking sleep, completely different blast radius.

Fixing it: match the tool to the kind of work

Dropping to def is the quick patch, but it is not always the right one. It also does not help if the blocking call lives deep inside an async def you cannot easily convert. The durable fix depends on what kind of slow work it is.

A decision guide: network or DB calls with an async client should stay on the loop and use await; a sync-only library that is mostly I/O should be pushed to the threadpool with asyncio.to_thread; CPU-bound work such as parsing or embeddings should go to a process pool via run_in_executor

I/O with an async client available

This is the clean case. Swap the synchronous library for its async counterpart and await it: httpx instead of requests, and asyncpg or SQLAlchemy’s async engine instead of a sync driver. The work stays on the loop, and the loop stays free. In this route, the await on client.get is where the loop gets control back:

import httpx

@app.get("/fetch")
async def fetch():
    async with httpx.AsyncClient() as client:
        r = await client.get("https://example.com/data")  # yields at the await
    return r.json()

A sync-only library doing I/O

Sometimes there is no async version, or rewriting is not worth it. Push that one call to the threadpool with asyncio.to_thread (Python 3.9+) or Starlette’s run_in_threadpool. This is safe for I/O. CPython’s GIL (global interpreter lock) lets only one thread run Python code at a time, but CPython releases it while a thread waits on a socket, so the loop keeps running. Here requests.get runs on a worker thread while the route awaits its result:

import asyncio
import requests

@app.get("/legacy")
async def legacy():
    # requests has no async API; run it off the loop
    data = await asyncio.to_thread(requests.get, "https://example.com/data")
    return data.json()

CPU-bound work

This is the one people get wrong: parsing a big file, hashing, resizing an image, running a local embedding model. Threads do not save you here, because CPU work holds the GIL and one busy thread starves the loop anyway. Send it to a separate process instead. Below, run_in_executor hands expensive_embed to a process pool and awaits the result:

import asyncio
from concurrent.futures import ProcessPoolExecutor

pool = ProcessPoolExecutor()

@app.post("/embed")
async def embed(text: str):
    loop = asyncio.get_running_loop()
    vector = await loop.run_in_executor(pool, expensive_embed, text)
    return {"dim": len(vector)}

Failure modes that hide the real cause

It passes every load test, then falls over in prod. A single-user test never has two requests competing for the loop, so a blocked loop looks fast; you only see the problem under concurrency. Test with at least a handful of parallel connections, and keep a trivial /health route in the mix as your canary.

The threadpool is a ceiling, not a fix. Moving everything to def routes feels like it solved the problem, but the anyio pool defaults to 40 threads. Under enough concurrent slow requests, request 41 waits for a free thread. For genuinely high-throughput blocking work you need real async, not a bigger pool.

More --workers masks it. Running uvicorn --workers 4 gives you four processes, so a blocked loop in one worker affects only a quarter of traffic. That raises your ceiling and hides the bug at low load, but the endpoint still blocks; you have just spread the pain. See the uvicorn deployment docs for what workers actually buy you.

Streaming makes it obvious and worse. A blocked loop cannot push SSE bytes for any open stream, so one bad request freezes every user’s live output at once. That is why I care about this more on LLM backends than on plain CRUD services. The SSE streaming setup I use in CloudCanvasAI stays smooth only because nothing on the hot path blocks the loop between token yields.

To catch these before prod, turn on asyncio’s debug mode with PYTHONASYNCIODEBUG=1 or asyncio.run(main(), debug=True). It logs a warning whenever a callback holds the loop longer than loop.slow_callback_duration (100 ms by default), which points a finger straight at the offending call.

What I would do differently

Early on, when the server got slow under load, I reached for --workers and a bigger threadpool, treating it as a capacity problem. It usually is not. It is one specific line that blocks the loop, and adding processes just buys headroom while the real fix is a one-line offload or an async client. Now, when latency spreads across unrelated endpoints, the first thing I check is whether some async def on the hot path is calling something synchronous. That single question resolves this class of bug far more often than any amount of horizontal scaling.

The second habit: decide the boundary early. Every async def that touches the outside world gets an async client or an explicit to_thread, and anything CPU-heavy goes to a process pool from the start, not after an incident. Keeping the loop clean is much cheaper than finding the one call that dirtied it at 2 a.m.

Keep the loop free

FastAPI’s speed comes from a single event loop juggling everything, and that same design is why one careless blocking call takes the whole thing down. The framework already hands you the escape hatches: def for sync routes, to_thread for stray blocking calls, and a process pool for CPU work. The discipline is knowing which kind of work you are looking at, and never quietly parking a slow synchronous call on the loop.

This is the quiet backbone under the LLM services I’ve built, from the streaming backend in CloudCanvasAI and Gemini Alchemy to the operational console for CMS workflow management at CERN. None of them feel fast because of clever code. They feel fast because the loop is never blocked.


Diagrams by M. Hassan Ahmed, released under CC0. Image credit: original work by the author.