How FastAPI Dependency Injection Actually Works

A practical guide to FastAPI dependency injection: how Depends resolves a graph, yield setup and teardown, per-request caching, and where it leaks.

Every endpoint in a real backend needs the same handful of things before it can do any work: a database session, the authenticated user, a settings object, maybe a rate limiter. You can build those inline at the top of each handler. It works, and it rots. The construction code gets copy-pasted, the teardown gets forgotten, and the day you want to run the handler in a test, you find there is no seam to swap the real database for a fake one.

FastAPI’s answer is Depends: you declare what a handler needs as a parameter, and the framework builds it before the handler runs. This post is for engineers who already use Depends and want to know what it does underneath. It covers how a dependency graph gets resolved, why a shared dependency runs only once, what yield actually guarantees about cleanup, and the places it leaks.

The examples follow the shape of the Archi backend I worked on for CMS computing operations at CERN. There, every request has to be authenticated against CERN SSO (single sign-on) and handed a scoped database session before it touches a thing.

Depends is a resolver, not a container

If your mental model of dependency injection comes from Spring or .NET, drop the part about a global registry. There is no container you register services into, and no separate file that configures lifetimes. A dependency in FastAPI is just a callable. Depends(get_db) on a parameter means one thing: before you call this function, call get_db, and pass what it returns as this argument.

In this example, the endpoint needs a user, and getting the user needs a token and settings:

from fastapi import Depends, FastAPI

app = FastAPI()

def get_settings() -> Settings:
    return Settings()  # env-driven config

async def get_current_user(
    token: str = Depends(get_token),
    settings: Settings = Depends(get_settings),
) -> User:
    return await verify(token, settings.jwt_key)

@app.post("/documents")
async def create_document(user: User = Depends(get_current_user)):
    return {"owner": user.id}

The wiring is the function signature. get_current_user needs a token and settings, so it declares them as Depends parameters, and FastAPI resolves those first. That nesting is the whole idea: dependencies can depend on other dependencies, and the result is a graph that FastAPI walks for you. There is no registration step and no lifetime annotation. Whatever a function needs, it asks for in its parameters.

One request, one pass through the graph

When a request arrives, FastAPI resolves the graph from the endpoint down. create_document needs a user; the user needs a token and settings; the token might need the raw request. FastAPI walks that tree and calls each node.

The part worth internalizing is caching. Within a single request, FastAPI calls each dependency once and reuses the result for every other dependency that asks for it. If get_current_user and get_db both depend on get_settings, get_settings runs a single time and both receive the same object. This is on by default; the cache key is the callable plus its arguments.

A request to POST /documents resolves a dependency graph: the endpoint depends on get_current_user and get_db, get_current_user depends on get_token and get_settings, and get_db also depends on get_settings. The shared get_settings node is called once and cached for the rest of the request, so the second branch receives the same result rather than triggering a second call.

The trap is assuming this cache spans requests. It does not. A dependency that runs “once” runs once per request, and the next request starts clean. So get_settings above builds a fresh Settings() on every single request, which is wasteful if reading config is expensive. If you want a genuine app-wide singleton, memoize it yourself:

from functools import lru_cache

@lru_cache
def get_settings() -> Settings:
    return Settings()

Now get_settings still appears as a normal dependency and can still be overridden in tests, but the object is built once for the life of the process. The lru_cache handles the app-scoped part; Depends handles the request-scoped wiring. Keeping those two responsibilities separate is most of what “getting DI right” means here.

yield dependencies: setup, hand off, tear down

A plain return dependency produces a value. A yield dependency produces a value and a cleanup step, which is exactly what a database session or any other acquired resource needs:

async def get_db():
    session = SessionLocal()
    try:
        yield session
    finally:
        await session.close()

The function has three parts:

  • Setup: everything before yield runs before your handler.
  • Hand-off: the value you yield (here, session) is what the handler receives.
  • Teardown: everything after yield runs at the end.

The timing of that teardown is the detail people get wrong. The exit code runs after the response has been sent to the client, not before. Your handler returns, the bytes go out, and only then does session.close() fire.

A timeline of one request. Before the response line, setup runs outside-in: open span, then open db session, then the handler runs and yields its value. A dashed marker shows the response being sent to the client. After that line, teardown unwinds in reverse order (LIFO): close db session, then close span. A highlighted strip warns that the exit code runs after the response, so an exception there can no longer reach the handler's exception handlers, and a background task that outlives the teardown may find the session already closed.

With more than one yield dependency, FastAPI unwinds them in reverse order of setup, the way a stack of context managers would: the last one opened is the first one closed. That ordering is a guarantee, not luck. A dependency that opened a transaction inside a dependency that opened a connection tears down in the safe order.

Two sharp edges come straight out of this timing. Both are documented behavior rather than bugs:

  • Exceptions in teardown arrive too late. An exception raised in the exit code, after yield, happens after your response is already out the door, so your route’s exception handlers have run and cannot catch it. Worse, if you wrap the yield in try/except and swallow the exception without re-raising, FastAPI cannot tell an error occurred at all. Catch, log, and re-raise unless you truly mean to bury it.
  • Background tasks can race the teardown. Because teardown fires around the time the response is sent, resources from a yield dependency and code running in a BackgroundTask can race. Do not assume a database session opened in a yield dependency is still open inside a background task. This ordering has shifted across FastAPI releases, so pin down the behavior of the version you actually run before you rely on it.

Auth is the dependency that earns its keep

The single best use of Depends is authentication, because it turns “is this caller allowed?” into one declared input that every protected route shares. In CloudCanvasAI, a Claude-powered document platform I built, the server verifies a Firebase ID token on every request. That verification is a dependency, not a line repeated in forty handlers:

async def require_user(token: str = Depends(get_token)) -> User:
    user = await verify_firebase_token(token)
    if user is None:
        raise HTTPException(status_code=401, detail="invalid token")
    return user

When a handler wants the user object, it takes user: User = Depends(require_user). When a route only needs the check and not the return value, attach the dependency at the router so it covers a whole group of routes at once:

router = APIRouter(dependencies=[Depends(require_user)])

For anything parameterized, like role checks, reach for a class with __call__. The instance holds the config (here, the required role), and the call does the work, so you get a reusable, testable guard:

class RequireRole:
    def __init__(self, role: str):
        self.role = role

    def __call__(self, user: User = Depends(require_user)) -> User:
        if self.role not in user.roles:
            raise HTTPException(status_code=403, detail="forbidden")
        return user

admin_only = RequireRole("admin")

Testing is the reason this pays off

Here is the seam that inline construction never gives you. FastAPI lets you override any dependency at test time without touching a single handler:

from myapp import app, get_db

def override_get_db():
    yield test_session

app.dependency_overrides[get_db] = override_get_db

Every route that depends on get_db, directly or three levels deep, now receives the test session instead of the real one. There is no monkeypatching of module globals and no conditional branch for a test mode. The override map is keyed by the original callable, so the same trick swaps out require_user for a fixture that returns a canned user, or get_settings for a config that points at a throwaway database.

This is why declaring dependencies beats constructing them inline. The declaration gives you a named place to reach in and substitute; inline construction gives you nothing to grab.

Where it breaks

Sync dependencies and the event loop. A dependency written as def, not async def, runs in an external threadpool, exactly like a sync path operation, which keeps it from blocking the loop. But the threadpool is finite, so a sync dependency doing slow blocking I/O on a hot path will consume workers and starve everything else, the same failure I wrote about for blocking the FastAPI event loop. If a dependency does real I/O, make it async and use an async client; if it must stay sync, keep it cheap.

Treating request scope as app scope. The most common performance bug I see is a dependency that constructs an expensive client (an HTTP session, a model handle, a connection) on every request, because it lives in a plain Depends. Per-request is the default, and per-request construction of a shared resource is pure overhead. App-lifetime objects belong in the lifespan handler and get read out of application state, not rebuilt per call.

Silent cache surprises. Per-request caching is usually what you want, but if a dependency has side effects and you expected it to run twice in one request, it will run only once. Set Depends(dep, use_cache=False) when you genuinely need a fresh call each time. This is rare, and reaching for it is often a sign the logic wants to be plain code inside the handler instead.

Over-broad dependencies. A Depends that quietly does three unrelated things (verifies auth, opens a transaction, and logs) is hard to override in a test, because you cannot replace one part without replacing all three. Keep each dependency to one job. Small dependencies compose; fat ones fight you.

Tradeoffs, and what I would keep in mind

FastAPI’s dependency system is a request-scoped resolver with per-request caching and stack-ordered cleanup. That is a smaller thing than a full inversion-of-control container (the kind of framework that creates and wires every object for you), and the mistake is trying to grow it into one. It has no notion of scoped, transient, and singleton lifetimes. There is request scope, and there is whatever you memoize yourself for the process.

Once you stop expecting the missing features, the model is clean:

  • parameters declare needs,
  • yield handles cleanup in the right order, and
  • dependency_overrides gives every layer a test seam.

For the backends I build, that is enough, and I would not reach for more. Archi authenticates every operator request and hands it a scoped session through exactly this mechanism; CloudCanvasAI verifies a token on every call the same way. The plumbing stays boring and out of the handlers, which is the point: let the framework build what each request needs, so the endpoint is left to do the one thing it exists to do.


Diagrams by M. Hassan Ahmed, created for this post and released under CC0 (public domain). No external image was used.