#HPC(1)

September 2026
#LLM #GPU #Inference #HPC #DevOps

Splitting an LLM Across GPUs: Tensor vs Pipeline

A model too big for one GPU has to be split. How tensor and pipeline parallelism divide an LLM, what each costs, and when to reach for which.

Read more →