Splitting an LLM Across GPUs: Tensor vs Pipeline

A model too big for one GPU has to be split. How tensor and pipeline parallelism divide an LLM, what each costs, and when to reach for which.

At some point the model does not fit. You try to load a 70B-parameter checkpoint onto a single 80GB card. The weights alone want around 140GB in fp16, and the process dies before it serves a single token. The weights are only part of it: you also need room for activations and for the KV cache that grows with every request in flight. One GPU is not enough, so the model has to live on several. The question is how you cut it, and the two standard answers pull in opposite directions.

This post is for engineers past the toy stage of LLM serving who are now staring at a model that won’t fit on one card, or one that fits but runs too slowly. I’ll walk through tensor parallelism and pipeline parallelism: what communication each one forces onto the hot path (the work every token has to wait for), where each fails, and how to choose between them without guessing.

Both techniques come out of the training world. Tensor parallelism traces to the Megatron-LM paper and pipeline parallelism to GPipe, but the same tradeoffs decide how you serve.

Two ways to cut a model

You can slice a model along two axes, and they are genuinely different cuts.

Tensor parallelism cuts inside each layer. A transformer layer is mostly big matrix multiplies. You can split a weight matrix by its columns, hand each GPU a slice, and have every GPU compute a partial result on the same input at the same time. The catch is that the partial results have to be summed back together before the layer’s output is correct, and that reduction happens on every layer.

Pipeline parallelism cuts between layers. GPU 0 gets layers 1 through 8, GPU 1 gets layers 9 through 16, and a request walks through them in order, like parts down an assembly line. Each GPU holds fewer layers, so each holds less of the model. The GPUs only talk when a request crosses a stage boundary.

The distinction sounds academic until you look at where the network traffic lands. That is what decides your latency and your hardware bill.

Tensor parallelism: fast, chatty, wants NVLink

Here is the shape of tensor parallelism for one layer split across two GPUs:

Tensor parallelism splitting one transformer layer across two GPUs. A hidden state feeds into both GPU 0, holding the first half of the weight matrix columns and the first half of the attention heads, and GPU 1, holding the second half. Both compute partial outputs in parallel, then an all-reduce step sums the partial outputs before the result flows into the next layer. A note says the all-reduce runs on every layer in the forward pass, on the critical path of each token, which is why tensor parallelism wants a fast intra-node link like NVLink rather than the network.

Every GPU stores a slice of every weight matrix and runs the full hidden state (the vector that represents each token as it moves through the model) through its slice. Because the slices are column-partitioned, each GPU produces a partial output covering part of the feature dimension. An all-reduce, a collective operation that combines values from every GPU and hands each one the result, collects those partials and sums them. The layer’s output then matches what a single GPU would have produced.

Megatron’s arrangement is clever about this. It pairs a column-split matrix with a row-split one, so a full layer needs two reductions in the forward pass rather than one per matrix multiply. The principle holds either way: you communicate once or twice per layer.

That communication is the whole story. For a model with dozens of layers, decoding a single token triggers dozens of all-reduces, and each one blocks until every GPU in the group has checked in. This is why tensor parallelism belongs inside a node, across GPUs wired together with NVLink, NVIDIA’s direct GPU-to-GPU link, where the interconnect is measured in hundreds of GB/s. Run the same all-reduce over ordinary Ethernet between machines, and the reduction, not the math, becomes what you are waiting on. The compute is embarrassingly parallel; the synchronization is not.

What you buy for that cost is real. Tensor parallelism cuts both the per-GPU weight memory and the per-GPU activation memory. And because all GPUs work on the same token simultaneously, it lowers the latency of that token. When a single request needs to come back fast, this is the axis that helps.

Pipeline parallelism: memory-cheap, but it bubbles

Pipeline parallelism gives up that low latency in exchange for far less communication.

Pipeline parallelism splitting layers into stages across two GPUs, drawn as a schedule over time. GPU 0 owns stage 1 with layers 1 to 8, GPU 1 owns stage 2 with layers 9 to 16. Four microbatches flow through GPU 0 first, then through GPU 1 one slot later. GPU 1 sits idle at the start until GPU 0 hands it the first microbatch, and GPU 0 sits idle at the end after passing its last microbatch on. Those idle gaps are labelled the pipeline bubble. A note says more microbatches shrink the bubble's share of total time, and that the GPUs talk only at stage boundaries so the link between them can be slower.

A GPU only sends data to the next stage when a request finishes its block of layers. That is one transfer per stage boundary, not one reduction per layer, so pipeline parallelism tolerates a slower link. It is the axis you reach for across nodes, where you have no NVLink between machines.

The pipeline bubble

The cost is the gap you can see in the diagram: the pipeline bubble. Stage 2 cannot start until stage 1 hands it something, so GPU 1 idles while the pipeline fills, and GPU 0 idles while it drains. If you push one request through at a time, half your GPUs are always waiting.

The fix is to break the batch into microbatches (smaller slices of the batch) and keep several in flight. While stage 2 works on microbatch 1, stage 1 is already busy on microbatch 2.

The Megatron-LM scaling paper gives the bubble’s share of total time for the simple schedule as (p - 1) / (m + p - 1), where p is the number of pipeline stages and m is the number of microbatches. The lesson is in the arithmetic:

  • With 4 stages and 4 microbatches, you waste roughly 43% of your compute to the bubble.
  • With 4 stages and 32 microbatches, it drops under 9%.

More microbatches, smaller bubble.

Pipeline parallelism is the memory-friendliest cut, since each GPU holds only its fraction of the layers and nothing has to be summed across the group. But it does nothing for the latency of a single token, and it needs steady traffic to stay efficient.

Choosing, and combining

The two are not rivals so much as tools for different constraints, and large deployments use both at once.

The rule I start from: tensor-parallel within a node, pipeline-parallel across nodes.

  • Put tensor parallelism where the interconnect is fast enough to absorb an all-reduce per layer. Today that means the GPUs sharing an NVLink domain inside one box.
  • Reach for pipeline parallelism to span boxes. It only pays a transfer at each stage boundary, so it survives a slower link between machines.

A model that spans, say, sixteen GPUs across two nodes is often tensor-parallel 8-way inside each node and pipeline-parallel 2-way between them. Serving stacks expose exactly these two knobs. In vLLM they are two arguments, tensor_parallel_size and pipeline_parallel_size. The first call below splits every layer 8 ways inside one node; the second adds a two-stage pipeline across two nodes:

from vllm import LLM

# 8 GPUs in one node, NVLink between them: split each layer 8 ways.
llm = LLM(
    model="meta-llama/Llama-3.1-70B-Instruct",
    tensor_parallel_size=8,
)

# 16 GPUs across two nodes: tensor-parallel inside each box,
# pipeline-parallel across the slow link between them.
llm = LLM(
    model="meta-llama/Llama-3.1-70B-Instruct",
    tensor_parallel_size=8,
    pipeline_parallel_size=2,
)

If the model already fits on one GPU, do neither. A single card serving a model that fits, with continuous batching to keep it busy, beats any split, because every form of parallelism adds communication that a lone GPU never pays. Reach for a split only when you have run out of memory or out of speed on one card.

Where it bites

The failure modes are consistent enough to list.

The tensor-parallel size has to divide the model evenly. You split attention heads across GPUs, so the head count has to divide by the tensor-parallel size. Ask a model with 64 heads to run tensor-parallel 6 ways and it will refuse or pad awkwardly. Powers of two that divide both the head count and the hidden size are the safe choices.

A slow link quietly erases the win. The most common disappointment is running tensor parallelism across nodes, over Ethernet, and finding it slower than a smaller model on one GPU. The all-reduce is now bottlenecked on the network, and the GPUs spend their time blocked. If you must cross a node boundary, make that boundary a pipeline stage, not a tensor split.

Pipeline parallelism with a small batch is mostly bubble. In interactive serving, where requests trickle in one at a time, there may not be enough microbatches to fill the pipeline, so you pay the bubble on every request. Pipeline parallelism rewards steady, batched load and punishes sparse, latency-sensitive traffic.

The KV cache does not shrink the way weights do. Splitting the model spreads the weights across GPUs, but each request still carries a KV cache proportional to its context length. That cache lives on whichever GPUs hold the layers it passes through: under tensor parallelism it is sharded with the layer, and under pipeline parallelism each stage holds the cache for its own layers. Either way, how you manage that cache still decides how many concurrent requests you can hold, and no amount of model splitting rescues you from a cache that outgrows the memory left after the weights.

What I would reach for first

Start by asking whether you need to split at all. If the model fits on one GPU, keep it there and spend your effort on batching. If it does not fit, the next question is which resource you ran out of:

  • Out of memory, but latency is fine: pipeline parallelism is the cheapest way to spread the weights, as long as you feed it enough microbatches.
  • Latency is the problem, and you have a fast intra-node link: tensor parallelism pulls the per-token time down.
  • Out of both, across more GPUs than one node holds: combine them, tensor inside the box and pipeline across boxes.

I spend my days on the HPC and workflow tooling that keeps the CMS experiment at CERN running. There, the reflex to check what actually saturates before adding machinery is what keeps a cluster from quietly burning compute. The same discipline applies to a model split across GPUs on Archi, the RAG copilot I built for CMS operations. Parallelism is not free throughput. It is a trade: you pay in communication for the memory or the latency you needed. Know which one you were short on, put the chatty cut where the wires are fast, and the split earns its keep.


Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.