{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-05-sandboxing-llm-generated-code/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"d6031681-282c-5d5f-8d8c-67462fea5b86","excerpt":"There is a moment in building any code-writing agent when the model hands you a string of Python, and the question stops being “is this a good answer?” and…","html":"<p>There is a moment in building any code-writing agent when the model hands you a string of Python, and the question stops being “is this a good answer?” and becomes “where do I run this?” In <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a> that moment happens constantly. A user asks for a quarterly report, and Claude decides the cleanest way to produce a <code class=\"language-text\">.docx</code> is to write a short script against <code class=\"language-text\">python-docx</code>. Now there is a block of code that has to actually execute for the user to get a file back. The model wrote it, nobody reviewed it, and the next request comes from a different user on the same server.</p>\n<p>This post is for engineers who have an <a href=\"/blog/2026-06-30-llm-agent-tool-loop/\">agent loop</a> working and have reached the execution step. Running the code is easy. The part that needs design is running code you did not write, on behalf of users who do not trust each other, without one session’s mistake becoming everyone’s outage. I walk through why the obvious approaches fail, what each isolation option actually buys you, and how the per-user sandbox in CloudCanvasAI is put together with <a href=\"https://e2b.dev/docs\">E2B</a>.</p>\n<h2>The model writes code; your server runs it</h2>\n<p>Start by being precise about the mechanics, because they decide where the risk sits. A language model produces text. When it “runs code,” what actually happens is that your loop parses a tool call, pulls out the code string, and <em>your</em> process executes it. The <a href=\"https://docs.claude.com/en/api/agent-sdk/overview\">Claude Agent SDK</a> makes this ergonomic, but it does not change the shape: the model proposes, and your infrastructure disposes.</p>\n<p>That means the security boundary is entirely yours to place. The model has no opinion about whether the code touches your database credentials, because from its side it only ever emitted a string. What matters is what that string can reach once it runs.</p>\n<p>\n  <a\n    class=\"gatsby-resp-image-link\"\n    href=\"/static/df3d69535603299c30c74c5f454daf86/60356/trust-boundary.png\"\n    style=\"display: block\"\n    target=\"_blank\"\n    rel=\"noopener\"\n  >\n  \n  <span\n    class=\"gatsby-resp-image-wrapper\"\n    style=\"position: relative; display: block; margin: 7vw 0; max-width: 1360px; margin-left: auto; margin-right: auto;\"\n  >\n    <span\n      class=\"gatsby-resp-image-background-image\"\n      style=\"padding-bottom: 56.1764705882353%; position: relative; bottom: 0; left: 0; background-image: url('data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAABj0lEQVQoz3WS147dMAxE9/8/MMAGCyS3ualQ1aokQ+cmj2scGKSg0dAjfxDxRO6DvK/OnYfygjHBQgBwDrwG2K3ZrDkcQEqQUx5tMBLTRx8MkXeFt8d4rbjsdH9OCJwrsxxMlHhsM36G7TNtP+P+wy9f6dCYG82P1vmsDDZ5n+vgmMZxWZwinr331o4SJ6FysGi1Gi3vPkaiXmhczj5e+pDIOASPqXA6KWSOIQqqSTdiq65kdyZfzzK7o1JFLOOXyuKhddh22DYIMY+J1vdWGs7pqO6nfwWje7JY1uTWCFI3nn+dE4PvGopx1foqhYvTuC62IUTZp2pYo1Utmpn37AQpLvGYbDzbwMtRH1u+L2m3Xdp4Ms4hiLOmc2v+FvTNq0c00irMXQLr81IqwNden9vFqroJLN9Mc9I1dpFsX8X9jvqezCNbNdO/tGVsG/mw9Fznm+VAmeUSIwpA5cD869RfWb3ZRlRvsdylZCb+4z9SXysdO0CzpuGQYLtcHOMbkckKyk/C3z5EY2Dv32/gP1lYeG3y0pQPAAAAAElFTkSuQmCC'); background-size: cover; display: block;\"\n    >\n      <picture>\n        <source\n          srcset=\"/static/df3d69535603299c30c74c5f454daf86/51a8e/trust-boundary.webp 340w,\n/static/df3d69535603299c30c74c5f454daf86/713b7/trust-boundary.webp 680w,\n/static/df3d69535603299c30c74c5f454daf86/54376/trust-boundary.webp 1360w\"\n          sizes=\"(max-width: 1360px) 100vw, 1360px\"\n          type=\"image/webp\"\n        />\n        <source\n          srcset=\"/static/df3d69535603299c30c74c5f454daf86/ad208/trust-boundary.png 340w,\n/static/df3d69535603299c30c74c5f454daf86/a5a26/trust-boundary.png 680w,\n/static/df3d69535603299c30c74c5f454daf86/60356/trust-boundary.png 1360w\"\n          sizes=\"(max-width: 1360px) 100vw, 1360px\"\n          type=\"image/png\"\n        />\n        <img\n          class=\"gatsby-resp-image-image\"\n          style=\"width: 100%; height: 100%; margin: 0; vertical-align: middle; position: absolute; top: 0; left: 0; box-shadow: inset 0px 0px 0px 400px white;\"\n          src=\"/static/df3d69535603299c30c74c5f454daf86/60356/trust-boundary.png\"\n          alt=\"The model proposes code as text on the trusted side of a boundary; the code executes inside a per-user microVM on the untrusted side, so a crash or a rogue script only takes down one ephemeral sandbox instead of the app server\"\n          title=\"\"\n          src=\"/static/df3d69535603299c30c74c5f454daf86/60356/trust-boundary.png\"\n        />\n      </picture>\n      </span>\n  </span>\n  \n  </a>\n    </p>\n<h2>Why “just run it” is a trap</h2>\n<p>The first version everyone writes is some flavor of in-process execution:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token comment\"># Do not do this in production.</span>\n<span class=\"token keyword\">exec</span><span class=\"token punctuation\">(</span>generated_code<span class=\"token punctuation\">,</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\"__builtins__\"</span><span class=\"token punctuation\">:</span> __builtins__<span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>This is a sandbox in name only. The code runs inside your interpreter, so it can read every global you have loaded, including the database handle and the API keys sitting in module state. It can open any file the server process can open. An infinite loop pins a worker: <code class=\"language-text\">while True: pass</code>, and the request never returns.</p>\n<p>People have a long history of trying to make <code class=\"language-text\">exec</code> safe by stripping builtins, and there is a matching history of <a href=\"https://nedbatchelder.com/blog/201206/eval_really_is_dangerous.html\">sandbox escapes</a>. These walk back up the object graph from an innocent-looking value to <code class=\"language-text\">__subclasses__</code> and out. Treat CPython’s own docs as authoritative here: there is <a href=\"https://docs.python.org/3/library/functions.html#eval\">no supported way</a> to run untrusted Python in-process safely.</p>\n<p>Moving to a <code class=\"language-text\">subprocess</code> is a real improvement, and it is often where people stop. A separate OS process keeps a crash contained, and you can wrap it in a timeout, which handles the runaway loop. But the process shares your kernel, your disk, and your network. Nothing stops the code from reading <code class=\"language-text\">/etc/passwd</code>, writing into a directory another request will read, or making an outbound call to somewhere it should not. Subprocess isolation stops accidents. It does not stop a hostile input, and once the code is machine-generated on behalf of arbitrary users, “hostile” is just a matter of who is typing.</p>\n<h2>The isolation ladder: from shared process to hardware boundary</h2>\n<p>A useful way to think about the options is as a ladder. At the bottom, the code shares your process; at the top, it runs behind a hardware boundary. Each rung costs more to operate and walls off more.</p>\n<p><img src=\"/7c73823ea15c8acdff7d61f28fb0c2de/isolation-layers.svg\" alt=\"Four ways to run model-written code, weakest to strongest: exec() in-process which is not isolation, a subprocess which stops crashes but shares disk and network, a container or gVisor which namespaces the filesystem and network but shares one Linux kernel, and a Firecracker microVM which gives each user their own guest kernel behind a hardware boundary\"></p>\n<p>A <strong>container</strong> is the first rung that isolates by construction rather than by convention. Linux namespaces give the code its own view of the filesystem, network, and process table, and cgroups (control groups) cap its CPU and memory. For a lot of internal tooling, this is enough.</p>\n<p>The caveat is the shared kernel. Every container on the host talks to the same Linux kernel, so a kernel-level vulnerability is a shared-fate surface across all of them. <a href=\"https://gvisor.dev/\">gVisor</a> narrows that by putting a user-space kernel in front of the real one and intercepting system calls, which is the tradeoff Google made for running untrusted workloads at scale.</p>\n<p>The strongest common rung is a <strong>microVM</strong>. Instead of sharing the host kernel, each sandbox boots its own guest kernel inside a lightweight virtual machine. <a href=\"https://firecracker-microvm.github.io/\">Firecracker</a>, the VMM (virtual machine monitor) that AWS built for Lambda and Fargate, is the one most of this ecosystem runs on. The <a href=\"https://www.usenix.org/conference/nsdi20/presentation/agache\">NSDI 2020 paper</a> is the readable account of how it boots a VM in around 125 ms with a deliberately tiny device model. You get a real VM boundary with startup fast enough to put in a request path, which is exactly what running per-user code needs.</p>\n<p>That top rung is where CloudCanvasAI sits. Rather than operate Firecracker directly, I use E2B, which is that microVM model as a service.</p>\n<h2>Wiring up E2B</h2>\n<p><a href=\"https://github.com/e2b-dev/code-interpreter\">E2B</a> is open source and gives you a sandbox that is a Firecracker microVM, with a Python or JavaScript SDK on top. The core loop is small: create a sandbox, run code in it, read the output, and kill it.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> e2b_code_interpreter <span class=\"token keyword\">import</span> Sandbox\n\n<span class=\"token comment\"># A fresh microVM. The default wall-clock timeout is 300 seconds.</span>\nsbx <span class=\"token operator\">=</span> Sandbox<span class=\"token punctuation\">(</span>timeout<span class=\"token operator\">=</span><span class=\"token number\">300</span><span class=\"token punctuation\">)</span>\n\nexecution <span class=\"token operator\">=</span> sbx<span class=\"token punctuation\">.</span>run_code<span class=\"token punctuation\">(</span><span class=\"token string\">\"print(sum(range(1000)))\"</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">print</span><span class=\"token punctuation\">(</span>execution<span class=\"token punctuation\">.</span>logs<span class=\"token punctuation\">.</span>stdout<span class=\"token punctuation\">)</span>   <span class=\"token comment\"># ['499500\\n']</span>\n\nsbx<span class=\"token punctuation\">.</span>kill<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>   <span class=\"token comment\"># tear it down; the filesystem goes with it</span></code></pre></div>\n<p>For document generation, the interesting part is the filesystem. The generated code writes a file inside the sandbox, and you need to get the bytes back out across the boundary. E2B exposes that directly, along with shell commands for installing whatever the generated script imports:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">with</span> Sandbox<span class=\"token punctuation\">(</span>timeout<span class=\"token operator\">=</span><span class=\"token number\">300</span><span class=\"token punctuation\">)</span> <span class=\"token keyword\">as</span> sbx<span class=\"token punctuation\">:</span>\n    sbx<span class=\"token punctuation\">.</span>commands<span class=\"token punctuation\">.</span>run<span class=\"token punctuation\">(</span><span class=\"token string\">\"pip install python-docx\"</span><span class=\"token punctuation\">)</span>     <span class=\"token comment\"># deps for the generated script</span>\n    sbx<span class=\"token punctuation\">.</span>run_code<span class=\"token punctuation\">(</span>generated_docx_script<span class=\"token punctuation\">)</span>             <span class=\"token comment\"># writes /home/user/report.docx</span>\n    file_bytes <span class=\"token operator\">=</span> sbx<span class=\"token punctuation\">.</span>files<span class=\"token punctuation\">.</span>read<span class=\"token punctuation\">(</span><span class=\"token string\">\"/home/user/report.docx\"</span><span class=\"token punctuation\">)</span>\n    <span class=\"token comment\"># ... hand file_bytes to your file-serving endpoint</span>\n<span class=\"token comment\"># leaving the block kills the sandbox; nothing persists</span></code></pre></div>\n<p>The context manager matters more than it looks. When the block exits, the sandbox is destroyed, and its entire filesystem with it. There is no cleanup step to forget and no directory that a later request could stumble into, because the disk that request wrote to no longer exists. Being ephemeral by default is the property that makes multi-tenancy safe, not an afterthought you bolt on.</p>\n<h3>One sandbox per user, not per app</h3>\n<p>The rule I hold to is that a sandbox belongs to a session, never to the service. Two users never share one. This is the difference between “the code is isolated from the host” and “the code is isolated from <em>other users</em>,” and only the second one is a real multi-tenant guarantee. If everyone ran in a shared sandbox, one user’s script could read the file another user’s script just wrote. You would have rebuilt the cross-tenant leak you spun up microVMs to avoid.</p>\n<p>The lifecycle then looks like this:</p>\n<ul>\n<li><strong>Session starts:</strong> create a sandbox and store its ID against that session.</li>\n<li><strong>Each turn:</strong> run the generated code in that same sandbox, so state carries across the conversation.</li>\n<li><strong>Session ends or times out:</strong> kill the sandbox.</li>\n</ul>\n<p>E2B keeps a sandbox alive up to a ceiling (an hour on the base tier, longer on paid), and you can push the window out with <code class=\"language-text\">set_timeout</code> while a user is active. The sandbox ID is the handle you reconnect to between requests.</p>\n<h2>The failure modes nobody warns you about</h2>\n<p>Getting the happy path working takes an afternoon. The failures are what the second week is about.</p>\n<p><strong>Orphaned sandboxes are a bill.</strong> A microVM you forgot to kill keeps running until its timeout. Under load, or after a deploy that drops your in-memory map of sessions to sandboxes, you leak live VMs and pay for them. Two things save you. The wall-clock timeout is a backstop that guarantees nothing runs forever. And you want an out-of-band reaper, a separate job that lists live sandboxes and kills any whose session is gone. Do not rely only on the happy-path <code class=\"language-text\">kill()</code>; the process that was supposed to call it is exactly the one that crashed.</p>\n<p><strong>The timeout is a feature, and it will bite the legitimate case.</strong> 300 seconds is generous for a report and short for a script that decided to fetch and process a large file. When a real task hits the wall, you want to surface “this run exceeded its time budget,” not a bare stack trace. Set the timeout to the slowest reasonable task rather than the fastest, and extend it deliberately with <code class=\"language-text\">set_timeout</code> for the operations you know are long.</p>\n<p><strong>Package installs are slow and sometimes offline.</strong> Generated code imports libraries the base image does not have, so a <code class=\"language-text\">pip install</code> sneaks onto the critical path and adds seconds. For genuinely untrusted code, you may also want to cut network egress (outbound traffic) from the sandbox. If you do, <code class=\"language-text\">pip install</code> fails outright, and you need a pre-baked image with the common libraries already present. Decide the network posture early, because it changes how you build the sandbox image.</p>\n<p><strong>Getting output back is its own transport problem.</strong> <code class=\"language-text\">stdout</code> comes back fine. But a generated matplotlib chart or a <code class=\"language-text\">.docx</code> is bytes on the sandbox filesystem, and streaming a long-running run’s progress to the user is a second channel entirely. In CloudCanvasAI the results ride back over <a href=\"/blog/2026-06-30-fastapi-sse-streaming-llm/\">SSE from FastAPI</a>, so the user watches the document take shape instead of waiting on a spinner. The sandbox produces; the streaming layer delivers. Keep them separate in your head, or you will conflate “the code finished” with “the user has the file.”</p>\n<p><strong>Cold starts are real but small.</strong> A microVM boot is on the order of a couple hundred milliseconds. That is invisible next to model latency, but it is not zero. If you create a sandbox per turn instead of per session, you pay it every turn for no benefit, which is one more reason the sandbox should outlive a single request.</p>\n<h2>Tradeoffs, and what I would do differently</h2>\n<p>The honest tradeoff of the microVM approach is that you have taken on a dependency, whether that is E2B as a service or Firecracker to operate yourself. If your code executor only ever runs your own trusted code, that is overkill, and a plain container is the right call. The microVM earns its keep precisely when the code is untrusted and multi-tenant, which is the AI-agent case. So the decision is really “is this input adversarial?”, and for anything user-facing, the answer is yes.</p>\n<p>The thing I would tell my earlier self is to build the reaper on day one, not after the first surprise invoice. The happy-path teardown is the easy 90 percent. The leaked-VM cleanup is the 10 percent that only shows up under exactly the conditions where your normal cleanup did not run: a crash or a bad deploy. The shift that made the system predictable was treating sandbox lifecycle as something with its own out-of-band garbage collector, rather than a <code class=\"language-text\">try/finally</code> you trust.</p>\n<p>The second thing is to decide the network posture before writing much code, because it is load-bearing in ways that ripple outward:</p>\n<ul>\n<li><strong>Full egress</strong> makes <code class=\"language-text\">pip install</code> and any legitimate API call easy, and it makes exfiltration easy too.</li>\n<li><strong>No egress</strong> is safest, and it forces a pre-baked image and a story for anything the code genuinely needs to reach.</li>\n</ul>\n<p>There is no default that is right for everyone. There is a wrong move, though: not choosing, and discovering your posture by accident.</p>\n<h2>Where this runs</h2>\n<p>This is the execution layer under <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a>, where Claude drives a set of per-format skills, one each for <code class=\"language-text\">docx</code>, <code class=\"language-text\">pptx</code>, <code class=\"language-text\">pdf</code>, and <code class=\"language-text\">xlsx</code>. Every user session gets its own E2B sandbox, with an isolated filesystem that is thrown away when the session ends. The <a href=\"https://docs.claude.com/en/api/agent-sdk/overview\">Agent SDK</a> decides <em>what</em> code to write, the sandbox decides <em>where</em> it is allowed to run, and the <a href=\"/blog/2026-06-30-fastapi-sse-streaming-llm/\">streaming layer</a> decides how the result gets back. Keeping those three concerns separate is most of what turns an agent that writes and runs code into something you can put in front of real users, instead of a demo you run on your own laptop.</p>\n<p><em>Diagrams by M. Hassan Ahmed, made for this post. Image credit: M. Hassan Ahmed, released under CC0.</em></p>","frontmatter":{"title":"Running LLM-Generated Code in a Sandbox","date":"2026-07-05T00:00:00.000Z","description":"An LLM that writes and runs code needs real isolation, not a try/except. How to sandbox AI-generated code with E2B microVMs, and the failure modes.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB4klEQVQoz12R627cIBCF/a+xvQzY+I4BAwbfN46aKNkoyW6q9v1fqWxWUaVKn0YHNGc4GgIybsm4kWFFdqr3J/f+OV5+69cPX9XLm3p516f38fxr+fyTrTu4mYzrtf/LFcTAI9QmeZ+VY83XYT6t+4ebTtt+9trNp2F5PT5cjvulYktSWKA6PLQR4t4YxIkIcXugHeQKZSokwte0HqDow4SHiYhSGRJ+53tyhXMFVEIqD6mMUxFkzaTH54KvlC2Q91cyg4g8UFnez+U6gO7yZWT7DKazwjyr8ciUzK8Tg6Qahv1ij+dc3Cf1QuoZKotFj5hKV0NGgSRLJkM3jTRTTL613QsTPGt93gAXttaPjXkq5ANlW9qsmA20H7CyYCUSDXYaegWGw9D1XH1wc+6syuUdsICUrlaPrH8s5E7qJW02SHUUVxFuUudSZ0FrYno6WDCmkEb2jitDSxmiJoDC0tY/OCX1RMqRVJNfmHdGlBWby0cNsqWzrTaHu9anyI8OWxE3PET1NXbOt8b8bPSVrF3Bm1EdJQ2eFZ47UD6wFxK6Bqwiq8K2RU0bHqoAUYUL58Pf8Np/Q4iqKKn9nkFy7AyoDkSLrQatQHBsOtSKMC6Cu0Ph+RHnN27Hb8ov/tP/Lv8CCRJIesEq3/EAAAAASUVORK5CYII=","aspectRatio":1.899441340782123,"src":"/static/276b514ff6dac30102034af9dd35cc27/40a76/hero.png","srcSet":"/static/276b514ff6dac30102034af9dd35cc27/c972b/hero.png 340w,\n/static/276b514ff6dac30102034af9dd35cc27/27625/hero.png 680w,\n/static/276b514ff6dac30102034af9dd35cc27/40a76/hero.png 1360w,\n/static/276b514ff6dac30102034af9dd35cc27/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-05-sandboxing-llm-generated-code/","previous":"blog/2026-07-01-opensearch-workflow-monitoring/","next":"blog/2026-07-04-measuring-rag-retrieval-quality/"}},"staticQueryHashes":["32046230"]}