{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-29-llm-sampling-temperature-top-p-top-k/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"8df25a11-b866-5893-bcaa-49ba9200f5cd","excerpt":"The first time my RAG copilot returned a subtly wrong factual answer, I blamed retrieval. But the right document was in context. The model had the source in…","html":"<p>The first time my RAG copilot returned a subtly wrong factual answer, I blamed retrieval. But the right document was in context. The model had the source in front of it and still drifted into a plausible-sounding detail that the source did not contain. Retrieval was fine. The problem was the generation config: <code class=\"language-text\">temperature=0.8</code>, copied from an example tuned for chat. That one number did more damage than any missing chunk.</p>\n<p>Sampling parameters are the part of the stack most people set once and never revisit. They sit in a config file next to the model name, get copied between projects, and quietly decide how adventurous every token is.</p>\n<p>This post is for engineers who call an LLM as one component of a system and want to know what <code class=\"language-text\">temperature</code>, <code class=\"language-text\">top_p</code>, and <code class=\"language-text\">top_k</code> actually do to its output, instead of treating them as three dials to jiggle when something looks off. I start with how a model picks a token, show where each knob acts, and then spend time on how the knobs interact and the failure modes that bite in production.</p>\n<h2>Where the probabilities come from: logits and softmax</h2>\n<p>A language model does not emit words. At each step it produces a vector of <strong>logits</strong>: one raw score per token in its vocabulary, which means tens of thousands of scores. These scores are unbounded, so they are not probabilities. To get something you can sample from, you pass them through <a href=\"https://en.wikipedia.org/wiki/Softmax_function\">softmax</a>. Softmax exponentiates each score and then normalizes, so the results sum to one:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">p_i = exp(z_i) / Σ_j exp(z_j)</code></pre></div>\n<p>Here <code class=\"language-text\">z_i</code> is the logit for token <code class=\"language-text\">i</code> and <code class=\"language-text\">p_i</code> is its probability. The result is a probability distribution over the whole vocabulary, and every decoding strategy is a rule for picking one token from it:</p>\n<ul>\n<li><strong>Greedy decoding</strong> takes the highest-probability token every time.</li>\n<li><strong>Sampling</strong> draws at random according to the probabilities, so a token with 12% probability gets picked roughly 12% of the time.</li>\n</ul>\n<p>Everything in the rest of this post either reshapes that distribution or restricts which part of it you may draw from, before the draw happens.</p>\n<h2>Temperature: making the distribution sharper or flatter</h2>\n<p>Temperature is a single number, <code class=\"language-text\">T</code>, that divides the logits before softmax runs:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">p_i = exp(z_i / T) / Σ_j exp(z_j / T)</code></pre></div>\n<p>That division is the entire mechanism. With <code class=\"language-text\">T</code> below 1, the logits get larger, which widens the gaps between them, and softmax piles even more probability onto the tokens that were already likely. With <code class=\"language-text\">T</code> above 1, the gaps shrink and probability spreads out toward the tail (the long list of unlikely tokens). The diagram below shows the same six logits at three temperatures.</p>\n<p><img src=\"/f6c66733bd513b5e176710fe3ac4354d/temperature-softmax.svg\" alt=\"The same six candidate tokens after softmax at three temperatures. At T=0.5 the top token holds 83% of the mass and the rest is negligible. At T=1.0, the model&#x27;s raw distribution, the top token is 54%. At T=1.7 the distribution is flat: the top token is down to 37% and the tail carries real weight.\"></p>\n<p>Two points stand out in that picture.</p>\n<p>First, temperature adds no information. It cannot make the model consider a token it wasn’t already considering. It only moves probability among the candidates the logits already ranked. High-temperature “creativity” is really the model being allowed to pick further down its own list.</p>\n<p>Second, the endpoints are special. As <code class=\"language-text\">T</code> approaches 0, softmax collapses onto the single top token, which is exactly greedy decoding. Most APIs special-case <code class=\"language-text\">temperature=0</code> to mean greedy instead of dividing by zero.</p>\n<p>The mental model that serves me well: temperature is a confidence knob, not a quality knob. Low temperature makes the model commit to what it already thinks is most likely. You want that for extraction, classification, and the grounded factual answers in a retrieval system. It’s the same instinct behind <a href=\"/blog/2026-07-14-reliable-json-from-llms/\">getting reliable structured output out of a model</a>. High temperature buys variety, but the price is the model wandering into its own long tail, which is where confident nonsense lives.</p>\n<h2>Top-k and top-p: cutting off the tail before you sample</h2>\n<p>Temperature reshapes the whole distribution but never rules a token out. Even at <code class=\"language-text\">T=1</code>, a token with 3% probability at the bottom of the list is still a legal draw. Over thousands of tokens, those occasional draws from the far tail are where text goes off the rails. <strong>Truncation sampling</strong> fixes this by discarding the tail entirely and sampling only from what remains. The two common rules differ in where they cut.</p>\n<h3>Top-k: keep a fixed number of tokens</h3>\n<p><strong>Top-k</strong> is the blunt version. Keep the <code class=\"language-text\">k</code> highest-probability tokens, discard the rest, renormalize the survivors so they sum to one, and sample. <code class=\"language-text\">top_k=40</code> means “only ever consider the 40 most likely next tokens.”</p>\n<p>It is simple and cheap. Its weakness is that <code class=\"language-text\">k</code> is a fixed count that ignores the shape of the distribution:</p>\n<ul>\n<li>When the model is certain (a peaked distribution, where one token holds most of the probability), a <code class=\"language-text\">k</code> of 40 drags in 39 tokens the model barely wanted.</li>\n<li>When the model is genuinely uncertain across many reasonable options, the same <code class=\"language-text\">k</code> may cut off good candidates.</li>\n</ul>\n<h3>Top-p (nucleus sampling): keep a fixed share of the probability</h3>\n<p><strong>Top-p</strong>, also called <a href=\"https://arxiv.org/abs/1904.09751\">nucleus sampling</a> (Holtzman et al., 2019), adapts to the shape instead. Sort the tokens by probability and walk down the list, adding up probability as you go. Keep the smallest set whose cumulative probability reaches <code class=\"language-text\">p</code>, then renormalize and sample.</p>\n<p><code class=\"language-text\">top_p=0.9</code> means “keep however many tokens it takes to cover 90% of the probability, and no more.” On a peaked distribution that might be two tokens; on a flat one it might be fifty. The cutoff moves with the model’s confidence, which is exactly the property top-k lacks.</p>\n<p><img src=\"/0f78cf97c86004949ddec464eef0a79b/topk-vs-topp.svg\" alt=\"Two truncation rules on the same distribution. Top-k = 3 keeps exactly the three highest-probability tokens regardless of their mass. Top-p = 0.9 keeps the four tokens whose cumulative probability first reaches 90 percent, so the count follows the shape of the distribution.\"></p>\n<p>The paper that introduced nucleus sampling is worth reading for the full motivation. It showed that pure greedy decoding and pure temperature sampling both degrade into repetitive or incoherent text, and that truncating the unreliable tail is what keeps generation both varied and sane. That is why nearly every production stack defaults to top-p.</p>\n<h2>How the three knobs combine, and why order matters</h2>\n<p>You rarely set just one knob. In most implementations all three are active at once, and the order in which they apply matters. The common pipeline, which is also the default in <a href=\"https://huggingface.co/docs/transformers/en/main_classes/text_generation\">Hugging Face’s <code class=\"language-text\">transformers</code></a>, runs like this:</p>\n<ol>\n<li>Apply temperature to the logits.</li>\n<li>Filter with top-k.</li>\n<li>Filter with top-p.</li>\n<li>Renormalize whatever survived and draw a token.</li>\n</ol>\n<p>Here is that sampling step as pseudocode, without the tensor plumbing. The numbered comments match the order above.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token comment\"># The shape of what a sampling step does, minus the tensor plumbing.</span>\nlogits <span class=\"token operator\">=</span> model_step<span class=\"token punctuation\">(</span>context<span class=\"token punctuation\">)</span>          <span class=\"token comment\"># one score per vocab token</span>\nlogits <span class=\"token operator\">=</span> logits <span class=\"token operator\">/</span> temperature         <span class=\"token comment\"># 1. reshape</span>\nlogits <span class=\"token operator\">=</span> keep_top_k<span class=\"token punctuation\">(</span>logits<span class=\"token punctuation\">,</span> k<span class=\"token operator\">=</span><span class=\"token number\">40</span><span class=\"token punctuation\">)</span>     <span class=\"token comment\"># 2. hard cap on count</span>\nlogits <span class=\"token operator\">=</span> keep_top_p<span class=\"token punctuation\">(</span>logits<span class=\"token punctuation\">,</span> p<span class=\"token operator\">=</span><span class=\"token number\">0.9</span><span class=\"token punctuation\">)</span>    <span class=\"token comment\"># 3. adaptive cap on mass</span>\nprobs  <span class=\"token operator\">=</span> softmax<span class=\"token punctuation\">(</span>logits<span class=\"token punctuation\">)</span>              <span class=\"token comment\"># 4. renormalize survivors</span>\ntoken  <span class=\"token operator\">=</span> sample<span class=\"token punctuation\">(</span>probs<span class=\"token punctuation\">)</span>                <span class=\"token comment\"># 5. draw</span></code></pre></div>\n<p>Reading it as a pipeline explains two things.</p>\n<p>First, combinations that look redundant are not. <code class=\"language-text\">top_k=40</code> with <code class=\"language-text\">top_p=0.9</code> is not two ways of saying the same thing. Top-k sets a hard ceiling on how many tokens can ever be considered, and top-p tightens that further whenever the model is confident.</p>\n<p>Second, it exposes a trap. If you set <code class=\"language-text\">temperature=0</code> for deterministic output, top-p and top-k are irrelevant, because greedy decoding already ignores everything except the top token. Leaving them in the config is harmless, but it misleads the next person who reads it.</p>\n<h2>Failure modes I’ve actually hit</h2>\n<p><strong>Temperature 0 does not guarantee identical output.</strong> Greedy decoding is deterministic in theory, yet the same prompt sent twice to a hosted model can return different text. The logits aren’t bit-identical from run to run because of floating-point non-associativity (the order of additions changes the rounding), changing batch sizes, and routing differences in mixture-of-experts models. When the top two tokens are close, a tiny perturbation flips which one wins. If your tests assert exact string equality on model output, they will flake. Assert on the parsed result instead.</p>\n<p><strong>Low temperature plus aggressive truncation produces loops.</strong> Set temperature low <em>and</em> set a tight <code class=\"language-text\">top_k</code> or <code class=\"language-text\">top_p</code>, and you can starve the distribution down to one or two tokens at every step. The model gets stuck repeating a phrase because you removed every alternative it might have escaped through. If output degenerates into repetition, the fix is usually to loosen truncation, not tighten it.</p>\n<p><strong>Inherited defaults were tuned for a different job.</strong> An SDK example might ship <code class=\"language-text\">temperature=0.7</code>. That is reasonable for open-ended chat and wrong for a classifier whose label your code parses. When the model is one component feeding another, “close enough” is a bug, and the temperature that felt fine in a demo is why your extractor occasionally invents a field. That was exactly my RAG bug: a chat-tuned temperature on a task that needed the model to stay pinned to its source.</p>\n<p><strong>Some frontier models no longer expose the knobs.</strong> This matters when you pick a provider. The newest Anthropic models drop <code class=\"language-text\">temperature</code>, <code class=\"language-text\">top_p</code>, and <code class=\"language-text\">top_k</code> entirely and reject requests that set them, steering you toward prompting and other controls instead (see the <a href=\"https://platform.claude.com/docs/en/about-claude/models/migration-guide\">Claude Opus 4.7 migration notes</a>). If your pipeline depends on a specific temperature, that is a real portability cost. Don’t assume every model exposes the same dials.</p>\n<h2>What I’d reach for</h2>\n<p><strong>When code consumes the output</strong>, I default to <code class=\"language-text\">temperature=0</code> and stop there. That covers extraction, classification, tool-argument generation, and the grounded answer step in a RAG pipeline. Determinism is worth more than variety here, and the truncation parameters don’t matter under greedy decoding anyway. This is the setting I used for the factual paths in <a href=\"/project/archi/\">Archi</a>, the retrieval copilot I worked on for CMS computing operations at CERN. In ops answers, a fabricated detail is worse than a blunt one.</p>\n<p><strong>When variety is the point</strong> (drafting, brainstorming, generating diverse candidates to rank later), I start from a moderate temperature around 0.7 with <code class=\"language-text\">top_p=0.9</code> and adjust from there. I treat temperature as the coarse knob and top-p as the safety rail that keeps the tail out. Newer alternatives such as <a href=\"https://arxiv.org/abs/2407.01082\">min-p sampling</a> are worth watching for high-temperature settings, but top-p remains the dependable default.</p>\n<p>Mostly, though, the advice is smaller than any specific number: read the sampling config as carefully as you read the prompt. It is part of the model’s behavior, not boilerplate. The difference between a grounded answer and a confident hallucination is sometimes a single division by <code class=\"language-text\">T</code>. If you’re wiring a model into a larger system, the same care applies one layer up, in <a href=\"/blog/2026-06-30-llm-agent-tool-loop/\">the agent tool loop</a>, where a stray token can send a whole chain of tool calls the wrong way.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.</em></p>","frontmatter":{"title":"LLM Sampling: Temperature, Top-p, and Top-k","date":"2026-07-29T00:00:00.000Z","description":"Temperature, top-p, and top-k are the three knobs that shape how an LLM picks each token. How each one works, when to reach for it, and how they interact.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAABzklEQVQoz21QyY6bQBTkaKChVxbTzTbYbrMZMs4otsGeUXKKMusll/xA/v8L8tqWIkWKVGq9qq63WlRPAF7P2/P79vKhz+/V8Xl1egG6Oj3r89tmfl1PLyDWjx/r6XUzvwEVzZnqk+X4ietLx1tWej/un7phboZp2x0B7TD346XZTe1u0u2hNfoB9Gpzj7CERMtl2RW5jdUCpzbJXJo7Bpl5r9QmKdDra/SFryCAFMsFCSsSlLLoSihJUxcrgIMlYlmU1h41BrCBiILKW9Y47b2kcWlmOcan/GDNZM/lzmOFy0sUrqE2i3SQdCSuEcmcWzleEtnRbPSDFVCzsycqKkeqRiIHmvSuGTgN0kFUR7Y6iuIzi2sbToMlK/Z8deTVAQLECgtO5YeaLLv49FO0X0nceLx0cELjhuuL+v6bryeS9NAD0ZyqXfjpx/LxFy0ecKSthRcjfkflwIoHlt0bH1GOH4ukY9mebebg7gt0BhvMybJR1E9B+43n0Dm3bC8GwA7QE8c1zIyp4gKWTCBHyB6H2sXJ1RbBOUnSmp1FZaMIkiMAFjkN7wAeVTcFUcXCkkUVDcu/IgQkKACY50CtBQoBtz/AAvCv8j/RDAv0D2etR5Bw7PryAAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/c7670afd8e01157c75428c788f6b297f/40a76/hero.png","srcSet":"/static/c7670afd8e01157c75428c788f6b297f/c972b/hero.png 340w,\n/static/c7670afd8e01157c75428c788f6b297f/27625/hero.png 680w,\n/static/c7670afd8e01157c75428c788f6b297f/40a76/hero.png 1360w,\n/static/c7670afd8e01157c75428c788f6b297f/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-29-llm-sampling-temperature-top-p-top-k/","previous":"blog/2026-07-24-speculative-decoding-llm-inference/","next":"blog/2026-07-28-choosing-embedding-model-rag/"}},"staticQueryHashes":["32046230"]}