{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-08-10-async-connection-pooling-fastapi/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"11ede45d-5c6f-52a5-9d56-37d99accd26e","excerpt":"The first time a FastAPI service I ran fell over under load, nothing crashed. Requests just started hanging. A handful returned fast; the rest sat there for…","html":"<p>The first time a FastAPI service I ran fell over under load, nothing crashed. Requests just started hanging. A handful returned fast; the rest sat there for thirty seconds and then failed with <code class=\"language-text\">QueuePool limit of size 5 overflow 10 reached, connection timed out</code>. CPU was near idle, Postgres was fine, and the app was doing almost no work. It was waiting on itself.</p>\n<p>That error means the connection pool has run out of connections to lend. It is one of the most common ways an otherwise healthy async backend falls apart, and the fix is not “make the pool bigger.” The fix is to understand what the pool does, so you can size it against the one hard limit that matters: how many connections Postgres will accept.</p>\n<p>This post is for engineers building FastAPI backends on SQLAlchemy and Postgres who have hit a pool-exhaustion error, or want to avoid their first one. It covers the two settings that control the pool, how to size them against Postgres, stale connections, when to turn the pool off, and the failure modes I have hit in practice.</p>\n<p>I ran the FastAPI service behind the <a href=\"/project/cms-workflow-operations/\">CMS workflow operations console</a> at CERN and the backend for <a href=\"/project/archi/\">Archi</a>, the retrieval copilot. Both talk to databases on every request. Pool sizing never comes up in a demo, and then it decides whether your service survives its first busy afternoon.</p>\n<h2>Why pools exist: connections are expensive</h2>\n<p>Opening a Postgres connection is not free. It takes a TCP handshake, plus TLS if you use it, and on the server side Postgres forks a new backend process for each connection. Doing that per request would add tens of milliseconds and pile processes onto the database.</p>\n<p>So SQLAlchemy keeps a set of open connections and lends them out. A request borrows one, runs its queries, and returns it. The connection stays open for the next borrower.</p>\n<p>That borrowing is the whole game. With an async engine, SQLAlchemy manages it with an <a href=\"https://docs.sqlalchemy.org/en/20/core/pooling.html\"><code class=\"language-text\">AsyncAdaptedQueuePool</code></a>, which it selects automatically for <code class=\"language-text\">create_async_engine</code>. The plain <code class=\"language-text\">QueuePool</code> is not asyncio-safe, so you get the async-adapted one whether you ask for it or not.</p>\n<p><img src=\"/977bc3f3a5036452c283f3803c346700/pool-anatomy.svg\" alt=\"A connection pool sits between FastAPI request handlers and Postgres. Requests borrow a connection from the pool, use it, and return it. The pool keeps pool_size connections open, can open up to max_overflow more on demand, and Postgres caps the total at max_connections.\"></p>\n<p>As the diagram shows, two numbers control the pool. Both have defaults you will outgrow.</p>\n<h2>pool<em>size and max</em>overflow: the two numbers that set the ceiling</h2>\n<p><code class=\"language-text\">pool_size</code> is how many connections the pool keeps open and ready. It defaults to <strong>5</strong>. These connections are persistent: once opened, they stay in the pool between requests, so most checkouts (borrowing a connection from the pool) are instant.</p>\n<p><code class=\"language-text\">max_overflow</code> is how many <em>extra</em> connections the pool may open when all the persistent ones are busy. It defaults to <strong>10</strong>. Overflow connections are temporary. When a request returns one and the pool already holds <code class=\"language-text\">pool_size</code> connections, the pool closes it instead of keeping it. Overflow is a burst valve, not extra capacity you get to keep.</p>\n<p>So the real ceiling per pool is <code class=\"language-text\">pool_size + max_overflow</code>, which is <strong>15</strong> by default. Ask for a sixteenth concurrent connection and the pool has nothing to give. The request waits for up to <code class=\"language-text\">pool_timeout</code>, which defaults to <strong>30 seconds</strong>. If no connection frees up by then, you get the timeout error I opened with.</p>\n<p>Here is the engine and session setup I actually use. The comment on each line says what the setting is for:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> sqlalchemy<span class=\"token punctuation\">.</span>ext<span class=\"token punctuation\">.</span>asyncio <span class=\"token keyword\">import</span> create_async_engine<span class=\"token punctuation\">,</span> async_sessionmaker\n\nengine <span class=\"token operator\">=</span> create_async_engine<span class=\"token punctuation\">(</span>\n    <span class=\"token string\">\"postgresql+asyncpg://ops:***@db:5432/workflows\"</span><span class=\"token punctuation\">,</span>\n    pool_size<span class=\"token operator\">=</span><span class=\"token number\">10</span><span class=\"token punctuation\">,</span>        <span class=\"token comment\"># persistent connections</span>\n    max_overflow<span class=\"token operator\">=</span><span class=\"token number\">20</span><span class=\"token punctuation\">,</span>     <span class=\"token comment\"># burst room; ceiling is 30 per process</span>\n    pool_timeout<span class=\"token operator\">=</span><span class=\"token number\">10</span><span class=\"token punctuation\">,</span>     <span class=\"token comment\"># fail fast instead of hanging for 30s</span>\n    pool_recycle<span class=\"token operator\">=</span><span class=\"token number\">1800</span><span class=\"token punctuation\">,</span>   <span class=\"token comment\"># recycle connections older than 30 min</span>\n    pool_pre_ping<span class=\"token operator\">=</span><span class=\"token boolean\">True</span><span class=\"token punctuation\">,</span>  <span class=\"token comment\"># check liveness before handing one out</span>\n<span class=\"token punctuation\">)</span>\n\nAsyncSessionLocal <span class=\"token operator\">=</span> async_sessionmaker<span class=\"token punctuation\">(</span>engine<span class=\"token punctuation\">,</span> expire_on_commit<span class=\"token operator\">=</span><span class=\"token boolean\">False</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>I lowered <code class=\"language-text\">pool_timeout</code> to 10 seconds on purpose. A request that cannot get a connection in ten seconds is not going to have a good day at thirty. Failing fast turns a slow hang into a clean error the client can retry, and it keeps one stuck dependency from tying up every worker. <code class=\"language-text\">pool_recycle</code> and <code class=\"language-text\">pool_pre_ping</code> deal with stale connections, which get their own section below.</p>\n<h2>Size the pool against Postgres, not against your app</h2>\n<p>The trap is treating the pool size as a per-app setting. It is per <em>process</em>. Each Uvicorn or Gunicorn worker imports your module, calls <code class=\"language-text\">create_async_engine</code>, and gets its own independent pool. Run four workers and you have four pools, each able to open <code class=\"language-text\">pool_size + max_overflow</code> connections.</p>\n<p>Postgres does not care how many pools you have. It enforces a single cluster-wide <code class=\"language-text\">max_connections</code>, which <a href=\"https://www.postgresql.org/docs/current/runtime-config-connection.html\">defaults to 100</a>. About 3 of those are reserved for superusers, so roughly 97 are available to applications. That budget is shared across every process, every service, and every human running <code class=\"language-text\">psql</code>.</p>\n<p><img src=\"/9fda87128e0b2547f5bfb5b0d22022bb/connection-math.svg\" alt=\"A bar chart comparing total connections requested against the 97-slot Postgres ceiling. One worker asks for 15 connections and is safe, four workers ask for 60 and are safe, six workers ask for 90 and are tight, eight workers ask for 120 and exceed the cap.\"></p>\n<p>The arithmetic is unforgiving. With the config above (a ceiling of 30 per process) and 4 workers, you are asking for up to 120 connections from a database that will give you 97. That is before any background job, migration, or second service takes its share. The bar chart shows the same idea with the default settings. It stays green while the worker count is low, and then one deploy that bumps replicas tips you over into <code class=\"language-text\">FATAL: sorry, too many clients already</code>.</p>\n<p>So size from the database backward. The total across all processes has to fit under the Postgres limit, with headroom left for everything else:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">worker_processes  x  (pool_size + max_overflow)  &lt;  max_connections - headroom</code></pre></div>\n<p>Two examples of that rule in practice:</p>\n<ul>\n<li><strong>A service pinned at 4 workers against default Postgres:</strong> <code class=\"language-text\">pool_size=10, max_overflow=5</code> gives a ceiling of 60, which leaves room for everything else.</li>\n<li><strong>A service that scales on Kubernetes,</strong> as I do for the operations console: replicas multiply too. 3 pods times 4 workers is 12 processes, and the pool numbers have to be small enough that twelve of them still fit.</li>\n</ul>\n<p>This is the same “your local number multiplied by the fleet” mistake that shows up in <a href=\"/blog/2026-07-15-kubernetes-oomkilled-requests-vs-limits/\">Kubernetes memory limits</a>, just with a different resource.</p>\n<h2>Stale connections: pool<em>recycle and pool</em>pre_ping</h2>\n<p>A pooled connection can go bad while it sits idle:</p>\n<ul>\n<li>A firewall drops long-lived TCP sessions.</li>\n<li>Postgres restarts.</li>\n<li>A managed database fails over to a replica.</li>\n<li><code class=\"language-text\">idle_in_transaction_session_timeout</code> closes the connection on the server side.</li>\n</ul>\n<p>The pool does not know any of this happened. It still thinks it holds a live connection, so it hands it to a request, and the query dies with a connection-reset error. The error looks random because it depends on how long the connection sat unused.</p>\n<p>Two settings handle this:</p>\n<ul>\n<li><strong><code class=\"language-text\">pool_recycle</code></strong> sets a maximum connection age. A connection older than that is quietly replaced on its next checkout. The default is <code class=\"language-text\">-1</code>, meaning never, so on any real deployment set it below the shortest idle timeout in your network path. 1800 seconds is a safe start.</li>\n<li><strong><code class=\"language-text\">pool_pre_ping</code></strong> goes further. It runs a cheap liveness check before each checkout and swaps in a fresh connection if the old one is dead. It costs a tiny round trip per checkout, in exchange for never serving a broken connection to a user.</li>\n</ul>\n<p>On managed Postgres, where failovers happen without warning, I keep both on.</p>\n<h2>When the pool is the wrong tool: PgBouncer and NullPool</h2>\n<p><a href=\"https://www.pgbouncer.org/\">PgBouncer</a> and cloud poolers like RDS Proxy sit in front of Postgres and pool connections themselves. If you run one, you have two poolers stacked, and the outer one is where connection reuse should happen.</p>\n<p>Running SQLAlchemy’s pool on top of PgBouncer in transaction mode causes a specific, confusing bug. asyncpg caches prepared statements per connection. In transaction mode, PgBouncer hands a different physical connection to each transaction, so the cached statement is not there. Under load, you get <code class=\"language-text\">DuplicatePreparedStatementError</code> or <code class=\"language-text\">InvalidSQLStatementNameError</code>.</p>\n<p>The fix has two parts, and you need both:</p>\n<ol>\n<li>Turn off SQLAlchemy’s pool with <code class=\"language-text\">NullPool</code>, so it opens and closes a connection per checkout and lets PgBouncer do the pooling.</li>\n<li>Disable asyncpg’s statement caches, <a href=\"https://docs.sqlalchemy.org/en/20/dialects/postgresql.html#prepared-statement-cache\">both of them</a>. The second is an LRU cache that keeps the error rate nonzero if you only clear the first.</li>\n</ol>\n<p>In the code below, <code class=\"language-text\">poolclass=NullPool</code> handles the first part, and the two <code class=\"language-text\">connect_args</code> entries set both caches to zero:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> sqlalchemy<span class=\"token punctuation\">.</span>pool <span class=\"token keyword\">import</span> NullPool\n\nengine <span class=\"token operator\">=</span> create_async_engine<span class=\"token punctuation\">(</span>\n    <span class=\"token string\">\"postgresql+asyncpg://ops:***@pgbouncer:6432/workflows\"</span><span class=\"token punctuation\">,</span>\n    poolclass<span class=\"token operator\">=</span>NullPool<span class=\"token punctuation\">,</span>\n    connect_args<span class=\"token operator\">=</span><span class=\"token punctuation\">{</span>\n        <span class=\"token string\">\"statement_cache_size\"</span><span class=\"token punctuation\">:</span> <span class=\"token number\">0</span><span class=\"token punctuation\">,</span>\n        <span class=\"token string\">\"prepared_statement_cache_size\"</span><span class=\"token punctuation\">:</span> <span class=\"token number\">0</span><span class=\"token punctuation\">,</span>\n    <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n<span class=\"token punctuation\">)</span></code></pre></div>\n<p>If you are not behind an external pooler, do not reach for <code class=\"language-text\">NullPool</code>. You would pay the full connection cost on every request, which is the exact expense the pool exists to remove.</p>\n<h2>Failure modes I have actually hit</h2>\n<p><strong>Sessions that outlive the request.</strong> The most common cause of exhaustion is not too many requests. It is a session that never gets returned. If you open a session and forget to close it (an early <code class=\"language-text\">return</code>, or an exception that skips your cleanup), that connection stays checked out forever. Use a dependency that closes it no matter what:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">get_session</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">async</span> <span class=\"token keyword\">with</span> AsyncSessionLocal<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span> <span class=\"token keyword\">as</span> session<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">yield</span> session</code></pre></div>\n<p>The <code class=\"language-text\">async with</code> returns the connection on both the success and the error path. Leaking one connection per failing request drains the pool in minutes, and the symptom (timeouts) points at the pool, not at the leak.</p>\n<p><strong>Blocking the event loop while holding a connection.</strong> Suppose a handler checks out a connection and then does something synchronous and slow, such as a <code class=\"language-text\">requests</code> call or a heavy CPU loop. It holds that connection the entire time and starves every other request of it. This is the connection-pool version of a problem I wrote about in <a href=\"/blog/2026-07-10-fastapi-blocking-event-loop/\">blocking the FastAPI event loop</a>: the pool is only as available as your slowest handler lets it be.</p>\n<p><strong>Long transactions.</strong> A connection stays checked out for the life of a transaction, not the life of a query. Open a transaction, call an external API for two seconds inside it, and that connection is unavailable for two seconds. Keep transactions short, and do slow I/O outside them.</p>\n<p><strong>Migrations and shells eat from the same 97.</strong> During a deploy, your migration tool and any open <code class=\"language-text\">psql</code> session count against <code class=\"language-text\">max_connections</code> too. Leave headroom so a routine migration during traffic does not push you over the edge.</p>\n<h2>What I would do differently</h2>\n<p>I would set <code class=\"language-text\">pool_recycle</code> and <code class=\"language-text\">pool_pre_ping</code> on day one, instead of adding them after the first mysterious connection reset in production. They cost almost nothing, and they remove a class of bug that is miserable to reproduce because it depends on idle timing.</p>\n<p>I would also write down the connection budget somewhere visible, ideally as a comment next to the engine. The number that is safe today quietly becomes unsafe the moment someone adds a second service to the same database or bumps the replica count. Pool sizing is not a one-time decision. It is a shared budget against a fixed ceiling, and the failures come from forgetting that the ceiling is shared.</p>\n<p>None of this is exotic. It is the boring plumbing under every FastAPI backend I’ve run, including the ones behind <a href=\"/project/archi/\">Archi</a> and the CMS operations tooling. Getting it right is the difference between a service that degrades gracefully under load and one that hangs for thirty seconds and then gives up.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, created for this post and released under CC0 (public domain). No external image was used.</em></p>","frontmatter":{"title":"Async Connection Pooling in FastAPI","date":"2026-08-10T00:00:00.000Z","description":"Your async FastAPI app stalls under load, then times out. Here is how SQLAlchemy pool_size and max_overflow really work, and how to size them for Postgres.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAACMElEQVQozz3Oy2/TMAAG8BzG2jiJEzupY8dxXk2yvh9sbKxdu62HCTgg9hLlgDhMGgN1PAYnEEzaJAQIBOIA/wESiAP8h7grQvopspTvsz+lmvu1C/KQl+nKUmXYb6+tNkfri4Nea3PQHfbaa1eba6utYb+z1M0WUvY/r8zDsAAj+VVRAknFT5ezxiCp9fPmMMxXFtobWXOQNoZ5a73SGXnJEiRV1YpnLUXHQsfBFBKGHalmoNlpKbiMWKMkOpi35Y2qJTQkgDVNysy/PBaKJn9cMJ0YkxCX/O0RPT0oHe24k33y4p5zeIt4ngdwGXt1N+yyqFsSbYtWdRwqwPQlFXJMU+aXCeUnY/T7rPBpon19An6+Ln58qEcBU3GG/WaSr97Y3JbjEW/qaFrmM3K25QSm7Y+3nDeHxvGe9WxsnR3AyR7iHlOtyOYNniyK9IoTdBCrybEKgJ6kGgy7ZS4S12WP99GvU/XDA+3Lsfb9JXh3ZESCXtI8YDLPTVbieubGBZ1q0FNUyKYMJl9Gjm9ifmfLfntff7RvPh+b5wfG8a7JPTqnMQf7vSC/nlR7Ip83KIBMUQ0qFXUXlWLPjyh1T26bf84Lnyfg21Pw41Xx/ZEWCqKolGC+ndV383pfZHOaCwwqy65U1AkiMRcxY2S0bB/eNHc2rLvXzMYCA1Zo2z5GPHKjBi+3/DRn8aylyNqMCuUSuYIAg2iQgAsGEtAOoR3IqGlx3fJ0i0PENcMt6OQvpL1eAoequOIAAAAASUVORK5CYII=","aspectRatio":1.899441340782123,"src":"/static/6be4c650659d2ce67014e656d7abb978/40a76/hero.png","srcSet":"/static/6be4c650659d2ce67014e656d7abb978/c972b/hero.png 340w,\n/static/6be4c650659d2ce67014e656d7abb978/27625/hero.png 680w,\n/static/6be4c650659d2ce67014e656d7abb978/40a76/hero.png 1360w,\n/static/6be4c650659d2ce67014e656d7abb978/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-08-10-async-connection-pooling-fastapi/","previous":"blog/2026-08-06-fastapi-dependency-injection-explained/","next":"blog/2026-08-09-byte-pair-encoding-llm-tokenization/"}},"staticQueryHashes":["32046230"]}