#Transformers(2)

September 2026
#LLM #Transformers #AI #GPU #Inference

RoPE: How LLMs Encode Position and Extend Context

How rotary position embeddings encode token order by rotating query and key vectors, and how position interpolation and YaRN stretch a model's context window.

Read more →
August 2026
#AI #LLM #GPU #Inference #Transformers

How Grouped-Query Attention Shrinks the KV Cache

Grouped-query attention shares key/value heads across query heads to cut the KV cache. How GQA sits between MHA and MQA, with the memory math and code.

Read more →