{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-06-30-llm-agent-tool-loop/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"3b4f72a0-8f23-5bd7-8a40-c769e8d9b0a9","excerpt":"The first time you wire a tool into a language model, the model seems to do something it cannot do. You ask a question, and partway through answering, the model…","html":"<p>The first time you wire a tool into a language model, the model seems to do something it cannot do. You ask a question, and partway through answering, the model “runs a search” or “queries a database” and comes back with a fact it had no way of knowing. It looks as if the model reached out and grabbed the data itself.</p>\n<p>It did not. A loop sits between you and the model, and that loop is the entire trick. Once you have written it, the mystery is gone. What is left is a control-flow pattern that you fully control.</p>\n<p>This post is for engineers who have made a single LLM call and now want the model to use tools: search a corpus, hit an API, run a query, then keep going. I build the loop from scratch in Python with <a href=\"https://www.claude.com/\">Claude</a>, because the raw version makes every agent framework you meet afterward easy to read. The same loop runs under <a href=\"/project/archi/\">Archi</a>, the retrieval copilot I worked on for CMS computing operations at CERN. There, “search the logbook” is a tool the model calls when it decides a question needs history it does not already have.</p>\n<h2>A tool is a description, not a function</h2>\n<p>The first misconception is that the model executes code. It cannot. The model only ever produces text and structured requests. So what you hand it is not a function but a <em>description</em> of one: a name, a sentence about when to use it, and a <a href=\"https://json-schema.org/\">JSON Schema</a> that describes the arguments. The list below defines a single tool, <code class=\"language-text\">search_logbook</code>, that takes one required string argument:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\">tools <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span>\n    <span class=\"token punctuation\">{</span>\n        <span class=\"token string\">\"name\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"search_logbook\"</span><span class=\"token punctuation\">,</span>\n        <span class=\"token string\">\"description\"</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">(</span>\n            <span class=\"token string\">\"Search the operations logbook for past entries. Use this when the \"</span>\n            <span class=\"token string\">\"user asks about a previous incident, an error message, or how \"</span>\n            <span class=\"token string\">\"something was handled before.\"</span>\n        <span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n        <span class=\"token string\">\"input_schema\"</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">{</span>\n            <span class=\"token string\">\"type\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"object\"</span><span class=\"token punctuation\">,</span>\n            <span class=\"token string\">\"properties\"</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">{</span>\n                <span class=\"token string\">\"query\"</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\"type\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"string\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"description\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"Free-text search query\"</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n            <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n            <span class=\"token string\">\"required\"</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">[</span><span class=\"token string\">\"query\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span>\n        <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n    <span class=\"token punctuation\">}</span>\n<span class=\"token punctuation\">]</span></code></pre></div>\n<p>The model reads that list the same way it reads the conversation. When it decides <code class=\"language-text\">search_logbook</code> fits, it still does not call anything. It emits a structured block that says “I want to call <code class=\"language-text\">search_logbook</code> with <code class=\"language-text\">{&quot;query&quot;: &quot;T1_US_FNAL transfer error&quot;}</code>,” and then it stops and waits. Running the search is your job. The model has handed control back to you mid-thought.</p>\n<p>That handoff is the whole mechanism, and one field in the response records it: <a href=\"https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview\"><code class=\"language-text\">stop_reason</code></a>. When the model wants a tool, <code class=\"language-text\">stop_reason</code> is <code class=\"language-text\">tool_use</code>. When it has finished and written its answer, <code class=\"language-text\">stop_reason</code> is <code class=\"language-text\">end_turn</code>. The whole loop is built around that one field.</p>\n<h2>The loop, in full</h2>\n<p>Here is the complete agent. There is less to it than the word “agent” suggests. It calls the model, checks <code class=\"language-text\">stop_reason</code>, runs any requested tools, appends the results to the conversation, and repeats:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> anthropic\n\nclient <span class=\"token operator\">=</span> anthropic<span class=\"token punctuation\">.</span>Anthropic<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">run_tool</span><span class=\"token punctuation\">(</span>name<span class=\"token punctuation\">,</span> tool_input<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">if</span> name <span class=\"token operator\">==</span> <span class=\"token string\">\"search_logbook\"</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">return</span> search_logbook<span class=\"token punctuation\">(</span>tool_input<span class=\"token punctuation\">[</span><span class=\"token string\">\"query\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>   <span class=\"token comment\"># your real implementation</span>\n    <span class=\"token keyword\">return</span> <span class=\"token string-interpolation\"><span class=\"token string\">f\"Unknown tool: </span><span class=\"token interpolation\"><span class=\"token punctuation\">{</span>name<span class=\"token punctuation\">}</span></span><span class=\"token string\">\"</span></span>\n\nmessages <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span>\n    <span class=\"token punctuation\">{</span><span class=\"token string\">\"role\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"user\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"Has the T1_US_FNAL transfer error happened before?\"</span><span class=\"token punctuation\">}</span>\n<span class=\"token punctuation\">]</span>\n\n<span class=\"token keyword\">while</span> <span class=\"token boolean\">True</span><span class=\"token punctuation\">:</span>\n    response <span class=\"token operator\">=</span> client<span class=\"token punctuation\">.</span>messages<span class=\"token punctuation\">.</span>create<span class=\"token punctuation\">(</span>\n        model<span class=\"token operator\">=</span><span class=\"token string\">\"claude-sonnet-4-6\"</span><span class=\"token punctuation\">,</span>\n        max_tokens<span class=\"token operator\">=</span><span class=\"token number\">1024</span><span class=\"token punctuation\">,</span>\n        tools<span class=\"token operator\">=</span>tools<span class=\"token punctuation\">,</span>\n        messages<span class=\"token operator\">=</span>messages<span class=\"token punctuation\">,</span>\n    <span class=\"token punctuation\">)</span>\n\n    <span class=\"token comment\"># The model wrote its final answer. Stop.</span>\n    <span class=\"token keyword\">if</span> response<span class=\"token punctuation\">.</span>stop_reason <span class=\"token operator\">!=</span> <span class=\"token string\">\"tool_use\"</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">break</span>\n\n    <span class=\"token comment\"># Record what the model said, tool requests and all.</span>\n    messages<span class=\"token punctuation\">.</span>append<span class=\"token punctuation\">(</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"role\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"assistant\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> response<span class=\"token punctuation\">.</span>content<span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span>\n\n    <span class=\"token comment\"># Run every tool the model asked for, collect the results.</span>\n    tool_results <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span><span class=\"token punctuation\">]</span>\n    <span class=\"token keyword\">for</span> block <span class=\"token keyword\">in</span> response<span class=\"token punctuation\">.</span>content<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">if</span> block<span class=\"token punctuation\">.</span><span class=\"token builtin\">type</span> <span class=\"token operator\">==</span> <span class=\"token string\">\"tool_use\"</span><span class=\"token punctuation\">:</span>\n            result <span class=\"token operator\">=</span> run_tool<span class=\"token punctuation\">(</span>block<span class=\"token punctuation\">.</span>name<span class=\"token punctuation\">,</span> block<span class=\"token punctuation\">.</span><span class=\"token builtin\">input</span><span class=\"token punctuation\">)</span>\n            tool_results<span class=\"token punctuation\">.</span>append<span class=\"token punctuation\">(</span><span class=\"token punctuation\">{</span>\n                <span class=\"token string\">\"type\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"tool_result\"</span><span class=\"token punctuation\">,</span>\n                <span class=\"token string\">\"tool_use_id\"</span><span class=\"token punctuation\">:</span> block<span class=\"token punctuation\">.</span><span class=\"token builtin\">id</span><span class=\"token punctuation\">,</span>\n                <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> result<span class=\"token punctuation\">,</span>\n            <span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span>\n\n    <span class=\"token comment\"># Feed the results back as the next user turn, and loop.</span>\n    messages<span class=\"token punctuation\">.</span>append<span class=\"token punctuation\">(</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"role\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"user\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> tool_results<span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span>\n\nanswer <span class=\"token operator\">=</span> <span class=\"token string\">\"\"</span><span class=\"token punctuation\">.</span>join<span class=\"token punctuation\">(</span>b<span class=\"token punctuation\">.</span>text <span class=\"token keyword\">for</span> b <span class=\"token keyword\">in</span> response<span class=\"token punctuation\">.</span>content <span class=\"token keyword\">if</span> b<span class=\"token punctuation\">.</span><span class=\"token builtin\">type</span> <span class=\"token operator\">==</span> <span class=\"token string\">\"text\"</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">print</span><span class=\"token punctuation\">(</span>answer<span class=\"token punctuation\">)</span></code></pre></div>\n<p>Notice what is <em>not</em> there: no memory, no state machine, no planner. The only thing that persists across iterations is <code class=\"language-text\">messages</code>. That list grows by two entries per tool call:</p>\n<ul>\n<li>the assistant turn that requested the tool, and</li>\n<li>the user turn that carries the result.</li>\n</ul>\n<p>The model rereads the whole list on every pass. That is how it “remembers” what it already searched for.</p>\n<p>The diagram below draws out the same loop’s control flow.</p>\n<p><img src=\"/abb9514c332940b1673a4c7897b764f9/agent-loop.svg\" alt=\"An LLM agent loop: the user turn goes to the model, the model sets a stop_reason of tool_use or end_turn, your code runs the requested tool and appends the result, and the loop repeats until the model returns final text\"></p>\n<h3>Two rules you cannot skip</h3>\n<p>Two rules in that code are mandatory. Get either one wrong and you get a confusing error instead of a clean one.</p>\n<ol>\n<li><strong>Append the model’s entire <code class=\"language-text\">response.content</code> back into <code class=\"language-text\">messages</code>, not just the text.</strong> Each <code class=\"language-text\">tool_use</code> block carries an <code class=\"language-text\">id</code>. The API rejects the next turn if it cannot match every <code class=\"language-text\">tool_result</code> to the <code class=\"language-text\">tool_use</code> that asked for it.</li>\n<li><strong>Set each result’s <code class=\"language-text\">tool_use_id</code> to that exact id.</strong></li>\n</ol>\n<p>Together, these two rules let the model line up “I asked to search for X” with “here is what the search found.”</p>\n<h2>The model can ask for several tools at once</h2>\n<p>A single assistant turn can contain more than one <code class=\"language-text\">tool_use</code> block. The model might decide it needs to search two logbooks, or read three files, before it can answer. The loop above already handles this. It iterates over every block in <code class=\"language-text\">response.content</code> and gathers all the results before sending them back.</p>\n<p>The detail that bites people is the return shape. All of those results go back in <strong>one</strong> user message, as a list of <code class=\"language-text\">tool_result</code> blocks. If you split them across several messages, you quietly teach the model to stop making parallel calls, because the conversation no longer looks the way it did when the calls succeeded. The <a href=\"https://platform.claude.com/docs/en/agents-and-tools/tool-use/overview\">Anthropic tool-use docs</a> are explicit about this, so follow it from the first version.</p>\n<h2>Where the loop breaks</h2>\n<p>The happy path is short. Real systems spend their time in the failures, and most of those are not exotic.</p>\n<p><strong>A tool that throws.</strong> Your search hits a timeout or a bad query. Do not let that exception escape the loop and kill the turn. Catch it and report it back to the model as a result with <code class=\"language-text\">is_error</code> set, so the model can recover instead of the whole request dying. The <code class=\"language-text\">try</code>/<code class=\"language-text\">except</code> below wraps the tool call from the main loop:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">try</span><span class=\"token punctuation\">:</span>\n    result <span class=\"token operator\">=</span> run_tool<span class=\"token punctuation\">(</span>block<span class=\"token punctuation\">.</span>name<span class=\"token punctuation\">,</span> block<span class=\"token punctuation\">.</span><span class=\"token builtin\">input</span><span class=\"token punctuation\">)</span>\n    tool_results<span class=\"token punctuation\">.</span>append<span class=\"token punctuation\">(</span><span class=\"token punctuation\">{</span>\n        <span class=\"token string\">\"type\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"tool_result\"</span><span class=\"token punctuation\">,</span>\n        <span class=\"token string\">\"tool_use_id\"</span><span class=\"token punctuation\">:</span> block<span class=\"token punctuation\">.</span><span class=\"token builtin\">id</span><span class=\"token punctuation\">,</span>\n        <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> result<span class=\"token punctuation\">,</span>\n    <span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">except</span> Exception <span class=\"token keyword\">as</span> exc<span class=\"token punctuation\">:</span>\n    tool_results<span class=\"token punctuation\">.</span>append<span class=\"token punctuation\">(</span><span class=\"token punctuation\">{</span>\n        <span class=\"token string\">\"type\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"tool_result\"</span><span class=\"token punctuation\">,</span>\n        <span class=\"token string\">\"tool_use_id\"</span><span class=\"token punctuation\">:</span> block<span class=\"token punctuation\">.</span><span class=\"token builtin\">id</span><span class=\"token punctuation\">,</span>\n        <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> <span class=\"token string-interpolation\"><span class=\"token string\">f\"Tool failed: </span><span class=\"token interpolation\"><span class=\"token punctuation\">{</span>exc<span class=\"token punctuation\">}</span></span><span class=\"token string\">\"</span></span><span class=\"token punctuation\">,</span>\n        <span class=\"token string\">\"is_error\"</span><span class=\"token punctuation\">:</span> <span class=\"token boolean\">True</span><span class=\"token punctuation\">,</span>\n    <span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>Given an error this way, the model usually reads it, adjusts the query, and tries again, or tells the user it could not find the answer. Given an uncaught exception, your process just crashes.</p>\n<p><strong>A loop that never ends.</strong> Nothing in <code class=\"language-text\">while True</code> guarantees the model ever returns <code class=\"language-text\">end_turn</code>. A poorly described tool, or a task the model cannot actually complete, can keep it calling tools in circles. Always cap the iterations. A counter that breaks after, say, eight passes turns an expensive runaway into a bounded, debuggable one.</p>\n<p><strong>Context that only grows.</strong> Every tool result is appended and never removed, so a long agent run sends more and more tokens through the model on each pass. For a handful of calls this is fine. For a long-running session you eventually need to trim or summarize old tool results, because you pay to resend the entire history on every iteration.</p>\n<p><strong>A tool the model ignores, or overuses.</strong> The model decides whether to call a tool entirely from its <code class=\"language-text\">description</code>. A vague description means the tool fires at the wrong times or never fires at all. The fix is almost always in the prose, not the code: say plainly <em>when</em> to use the tool, not just what it does. Anthropic’s recent guidance is that a description stating the trigger condition (“call this when the user asks about a past incident”) gives a real lift over one that only names the capability.</p>\n<h2>Tool runner versus the manual loop</h2>\n<p>Every major SDK now ships a helper that hides this loop. Anthropic’s <a href=\"https://github.com/anthropics/anthropic-sdk-python\">Python SDK</a> has a tool runner that calls your functions and feeds the results back until the model is done. You write the tool bodies and skip the plumbing. Once the pattern is familiar, the runner is the right default.</p>\n<p>The manual loop earns its place when you need a seat in the middle of it, for example:</p>\n<ul>\n<li>a human approval step before a tool with side effects runs;</li>\n<li>custom logging on every call;</li>\n<li>a tool you only allow under certain conditions;</li>\n<li>a hard budget on cost.</li>\n</ul>\n<p>In <a href=\"/project/archi/\">Archi</a>, retrieval is read-only and safe to run automatically. Anything that could change state, though, is exactly the kind of action you want to gate, and gating is much easier when you own the loop. My advice: start manual to learn it, switch to the runner when the loop becomes boilerplate, and drop back to manual the moment you need to step in between the model’s request and the tool’s execution.</p>\n<h2>The limits of a simple loop</h2>\n<p>The loop is simple, and that simplicity is also its ceiling. It is reactive: the model takes one step, sees the result, and decides the next step. That works well for retrieval and lookups, which is most of what a tool-using assistant does. It is also the shape behind the <a href=\"https://arxiv.org/abs/2210.03629\">ReAct</a> pattern (interleaved reasoning and acting), which a lot of agent work descends from.</p>\n<p>It fits less well when a task genuinely needs a plan laid out before any action, or when many independent subtasks should run at once. For those cases you build structure <em>on top of</em> this loop; you do not replace it. Anthropic’s <a href=\"https://www.anthropic.com/engineering/building-effective-agents\">Building Effective Agents</a> is a good map of when to add that structure and when a single loop is enough. The honest answer is that a plain loop covers more ground than its simplicity suggests.</p>\n<h2>What I would do differently</h2>\n<p>Early on, I underrated tool descriptions. I treated them as documentation and spent my time on the code. In fact they are part of the prompt and deserve the same iteration. A model that calls the wrong tool, or refuses to call any, is usually telling you the description is unclear, not that the model is broken. Now I write the description first, test which queries trigger the tool, and tune the wording before I touch the implementation.</p>\n<p>I would also add the iteration cap and the per-tool error handling on day one, not after the first runaway. Both are a few lines. Together they separate an agent that fails loudly and recoverably from one that crashes or quietly burns tokens in a circle.</p>\n<p>If you want to see this loop inside real products:</p>\n<ul>\n<li><a href=\"/project/archi/\">Archi</a> wraps it around CERN operations data.</li>\n<li><a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a> uses tool calls to drive live document edits next to a chat panel.</li>\n<li><a href=\"/project/llm-dev-mate/\">LLM DevMate</a> is the smaller, editor-side cousin of the same idea.</li>\n</ul>\n<p>The surfaces differ, but the same handful of lines decides when to call out and when to answer.</p>\n<hr>\n<p><em>Image credit: agent-loop diagrams by M. Hassan Ahmed, created for this post, released under CC0 (public domain).</em></p>","frontmatter":{"title":"How an LLM Agent Tool-Calling Loop Works","date":"2026-06-30T00:00:00.000Z","description":"A practical look at the agent loop behind LLM tools: how the model asks to call a tool, your code runs it, and the result feeds back until the answer is done.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB0klEQVQoz02QWY/aMBSF81QgXuPETkK8ZCFJSTIzUAgMDGol2gppXkrb//9fepl2UKVPlu1zju7i5e50x7kXa45Df9nvfo3b63bzY7u57rY/d+ONrvtu9OF/vzdl8ztMlUiaJO+LbrsYdlU/lsNYP+yXq1PzeEjznqmKx+Xd7+HA4sC9YVlU9c2rUg/Ojk190noThR0Xtc8tAgJHo4KEBRYOCwt4iJu/+FyTwC3zi0v2hXkGSn2w8RiGLUiI6ZuHaZ/+AzHj3d7vzOg8lC0LK8JzFtTgFvIj5s7nGQqtTzMUaJI6pkumC6xu4eydOeY2iBomykX2udXnWD4xWYXJ0peGFg2OC9rWYnD5l3X23JGF9nw6RyxDb2Ea5FTkMuyX7uu6fXXJgaWLYN4gCTNrKIsjg2ODlEaxxlJ7M5pOiJqSZILVjKRUFIQ7LXcm2smow9xgZj4gOcExMMUxeO54sAmlhkg2cbKUcQt3LkrMNQ0cSLCCNHmKVAsqEEYNdDojCZSE0wvDelVdF/a46i5DczZqLUQNeVgy1AxE9an5XbuXx/bb0J5LfaTMTbGcEWgh9qATTDWhFhNDqKHMTmEKrHySTnAEl7uKqSHM+iSBMPwDfwCNBEUA8WM2TAAAAABJRU5ErkJggg==","aspectRatio":1.899441340782123,"src":"/static/7250ff100f79cc48a94fdb61a4cb4e3e/40a76/hero.png","srcSet":"/static/7250ff100f79cc48a94fdb61a4cb4e3e/c972b/hero.png 340w,\n/static/7250ff100f79cc48a94fdb61a4cb4e3e/27625/hero.png 680w,\n/static/7250ff100f79cc48a94fdb61a4cb4e3e/40a76/hero.png 1360w,\n/static/7250ff100f79cc48a94fdb61a4cb4e3e/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-06-30-llm-agent-tool-loop/","previous":null,"next":"blog/2026-06-30-fastapi-sse-streaming-llm/"}},"staticQueryHashes":["32046230"]}