Chengshu@skadai · 2026.10.03
RSS
2,733 字 · 2,730 词 · 约 12 分钟

如何扩展你的模型(0):引言

Google DeepMind 的开源书《How To Scale Your Model》全书中英对照第 0 篇:为什么「让模型跑得快」和「让模型变强」同等重要,强扩展(strong scaling)、roofline、通信与显存这些词到底在说什么,以及全书 12 个部分如何组织。左栏英文原文,右栏中文译文,原文配图全部保留。

本篇属于系列 如何扩展你的模型(How To Scale Your Model) · 第 0 篇

原文 · English中文译文

原文:How To Scale Your Model(Google DeepMind,作者 Jacob Austin、Sholto Douglas、Roy Frostig、Anselm Levskaya、Charlie Chen、Sharad Vikram、Federico Lebron、Peter Choy、Vinay Ramasesh、Albert Webson、Reiner Pope)。

本站是中英对照排版:左栏是英文原文,右栏是对应的中文译文。不好直译的术语(roofline、strong scaling、ICI、MXU、KV cache、FSDP、TP 等)保留英文写法;原文配图全部保留。

本文是系列《如何扩展你的模型(How To Scale Your Model)》的第 0 篇,共 13 篇。系列目录。

原文以 MIT 许可证发布,版权归 Google LLC;本译文仅作学习交流之用,如有错漏以原文为准。

Much of deep learning still boils down to a kind of black magic, but optimizing the performance of your models doesn’t have to — even at huge scale! Relatively simple principles apply everywhere — from dealing with a single accelerator to tens of thousands — and understanding them lets you do many useful things:

深度学习在很大程度上仍像一门玄学,但优化模型的性能不必如此——即使是在超大规模下也一样!相对简单的原则到处都适用——从单颗加速器到上万颗——理解这些原则能让你做很多有用的事:

  • Ballpark how close parts of your model are to their theoretical optimum.
  • Make informed choices about different parallelism schemes at different scales (how you split the computation across multiple devices).
  • Estimate the cost and time required to train and run large Transformer models.
  • Design algorithms that take advantage of specific hardware affordances.
  • Design hardware driven by an explicit understanding of what limits current algorithm performance.
  • 粗略估计模型的各个部分距离它们的理论最优还有多远。
  • 在不同规模下为不同的并行方案(如何把计算切分到多台设备上)做出有依据的选择。
  • 估算训练和运行大型 Transformer 模型所需的成本与时间。
  • 设计能利用特定硬件特性的算法。
  • 基于「当前算法的性能瓶颈究竟在哪」这一明确认识来设计硬件。

Expected background: We’re going to assume you have a basic understanding of LLMs and the Transformer architecture but not necessarily how they operate at scale. You should know the basics of LLM training and ideally have some basic familiarity with JAX. Some useful background reading might include this blog post on the Transformer architecture and the original Transformer paper. Also check out this list for more useful concurrent and future reading.

前置背景: 我们假定你对 LLM 和 Transformer 架构有基本了解,但不要求你了解它们在大规模下如何运行。你应当知道 LLM 训练的基础知识,最好对 JAX 有初步熟悉。一些有用的背景阅读包括这篇讲 Transformer 架构的博客文章和原始的 Transformer 论文。也可以看看这个清单,里面有更多当下和未来的阅读材料。

Goals & Feedback: By the end, you should feel comfortable estimating the best parallelism scheme for a Transformer model on a given hardware platform, and roughly how long training and inference should take. If you don’t, email us or leave a comment! We’d love to know how we could make this clearer.

目标与反馈: 读完本书,你应当能从容地为给定硬件平台上的 Transformer 模型估算出最佳的并行方案,并大致判断训练和推理需要多久。如果做不到,写信给我们或留言!我们很想知道怎样能讲得更清楚。

You might also enjoy reading the new Section 12 on NVIDIA GPUs!

你可能也会喜欢新写的第 12 节:关于 NVIDIA GPU。

Why should you care?

为什么你该关心这些?

Three or four years ago, I don’t think most ML researchers would have needed to understand any of the content in this book. But today even “small” models run so close to hardware limits that doing novel research requires you to think about efficiency at scale. (Historically, ML research has followed something of a tick-tock cycle between systems innovations and software improvements. Alex Krizhevsky had to write unholy CUDA code to make CNNs fast but within a couple of years, libraries like Theano and TensorFlow meant you didn’t have to. Maybe that will happen here too and everything in this book will be abstracted away in a few years. But scaling laws have pushed our models perpetually to the very frontier of our hardware, and it seems likely that, for the foreseeable future, doing cutting-edge research will be inextricably tied to an understanding of how to efficiently scale models to large hardware topologies.) A 20% win on benchmarks is irrelevant if it comes at a 20% cost to roofline efficiency. Promising model architectures routinely fail either because they can’t run efficiently at scale or because no one puts in the work to make them do so.

三四年前,我认为大多数 ML 研究者都不需要理解本书里的任何内容。但今天,就连「小」模型也跑得如此贴近硬件极限,以至于做新颖的研究也需要你从规模化效率的角度去思考。(从历史上看,ML 研究一直像是在系统创新和软件改进之间来回摆动的滴答钟摆。Alex Krizhevsky 不得不写「不圣洁」的 CUDA 代码才能让 CNN 变快,但没过几年,Theano、TensorFlow 这样的库就让大家不必再这么做了。也许这里也会发生同样的事,几年后本书里的一切都会被抽象掉。但 scaling laws 把我们的模型不断推到硬件的最前沿,而且在可预见的未来,做前沿研究似乎都必然与「如何把模型高效扩展到大规模硬件拓扑上」的理解绑定在一起。)如果一项 20% 的基准提升要以 20% 的 roofline 效率损失为代价,那它毫无意义。 很有前景的模型架构常常失败,要么因为它们_无法_在大规模下高效运行,要么因为没人投入精力让它们做到。

The goal of “model scaling” is to be able to increase the number of chips used for training or inference while achieving a proportional, linear increase in throughput. This is known as “strong scaling”. Although adding additional chips (“parallelism”) usually decreases the computation time, it also comes at the cost of added communication between chips. When communication takes longer than computation we become “communication bound” and cannot scale strongly. (As your computation time decreases, you also typically face bottlenecks at the level of a single chip. Your shiny new TPU or GPU may be rated to perform 500 trillion operations per second, but if you aren’t careful it can just as easily do a tenth of that if it’s bogged down moving parameters around in memory. The interplay of per-chip computation, memory bandwidth, and total memory is critical to the scaling story.) If we understand our hardware well enough to anticipate where these bottlenecks will arise, we can design or reconfigure our models to avoid them. (Hardware designers face the inverse problem: building hardware that provides just enough compute, bandwidth, and memory for our algorithms while minimizing cost. You can imagine how stressful this “co-design” problem is: you have to bet on what algorithms will look like when the first chips actually become available, often 2 to 3 years down the road. The story of the TPU is a resounding success in this game. Matrix multiplication is a unique algorithm in the sense that it uses far more FLOPs per byte of memory than almost any other (N FLOPs per byte), and early TPUs and their systolic array architecture achieved far better perf / $ than GPUs did at the time they were built. TPUs were designed for ML workloads, and GPUs with their Tensor Cores are rapidly changing to fill this niche as well. But you can imagine how costly it would have been if neural networks had not taken off, or had changed in some fundamental way that TPUs (which are inherently less flexible than GPUs) could not handle.)

「模型扩展(model scaling)」的目标是:在增加用于训练或推理的芯片数量的同时,让吞吐量获得成比例的线性增长。 这被称为「强扩展(strong scaling)」。虽然增加芯片(「并行」)通常会减少计算时间,但代价是芯片之间新增的通信。当通信耗时长于计算时,我们就变成「通信受限」,无法做强扩展。(随着计算时间缩短,你通常还会在单颗芯片的层面遇到瓶颈。你崭新的 TPU 或 GPU 标称每秒能执行 500 万亿次运算,但如果不小心,它也很容易只做到十分之一——如果它把时间都耗在内存里搬参数上的话。单芯片计算、内存带宽和总内存三者之间的相互作用,是扩展这件事的关键。)如果我们足够了解硬件,能预判这些瓶颈会在哪里出现,就能设计或重新配置模型来避开它们。(硬件设计者面临的是反问题:在最小化成本的前提下,造出恰好提供足够计算、带宽和内存的硬件来支撑我们的算法。你可以想象这个「协同设计」问题有多让人焦虑:你必须押注第一批芯片真正可用时算法会是什么样子——而那往往是 2 到 3 年之后。TPU 的故事是这场博弈中一个响亮的成功。矩阵乘法是一种独特的算法:它每字节内存所需的 FLOPs 远超几乎任何其他算法(每字节 N 次 FLOPs),而早期 TPU 及其脉动阵列架构在诞生时所达到的性价比(perf / $)远好于当时的 GPU。TPU 是为 ML 负载设计的,而带 Tensor Core 的 GPU 也正在迅速填补这个生态位。但你可以想象,如果神经网络没有火起来,或者以某种根本性的方式变了形、以至于 TPU(天生没有 GPU 灵活)无法应对,那代价会有多大。)

Our goal in this book is to explain how TPU (and GPU) hardware works and how the Transformer architecture has evolved to perform well on current hardware. We hope this will be useful both for researchers designing new architectures and for engineers working to make the current generation of LLMs run fast.

本书的目标是解释 TPU(以及 GPU)硬件如何工作,以及 Transformer 架构如何演化成能在当前硬件上表现出色。我们希望这对两类人都有用:设计新架构的研究者,以及努力让当前这一代 LLM 跑得快的工程师。

High-Level Outline

全书结构

The overall structure of this book is as follows:

本书的整体结构如下:

Section 1 explains roofline analysis and what factors can limit our ability to scale (communication, computation, and memory). Section 2 and Section 3 talk in detail about how TPUs work, both as individual chips and — of critical importance — as an interconnected system with inter-chip links of limited bandwidth and latency. We’ll answer questions like:

第 1 节 讲解 roofline 分析,以及哪些因素会限制我们扩展的能力(通信、计算和内存)。第 2 节 和 第 3 节 详细讲 TPU 如何工作——既作为单颗芯片,也作为(极其重要的)一个由带宽和延迟都有限的片间链路连接起来的互联系统。我们会回答诸如这样的问题:

  • How long should a matrix multiply of a certain size take? At what point is it bound by compute or by memory or communication bandwidth?
  • How are TPUs wired together to form training clusters? How much bandwidth does each part of the system have?
  • How long does it take to gather, scatter, or re-distribute arrays across multiple TPUs?
  • How do we efficiently multiply matrices that are distributed differently across devices?
  • 给定尺寸的矩阵乘法应该花多长时间?它在什么点上会受限于计算、内存还是通信带宽?
  • TPU 是如何连起来组成训练集群的?系统各部分的带宽各是多少?
  • 在多个 TPU 之间 gather、scatter 或重新分布数组需要多久?
  • 如何高效地相乘那些在不同设备上以不同方式分布的矩阵?
Figure: a diagram from Section 2 showing how a TPU performs an elementwise product. Depending on the size of our arrays and the bandwidth of various links, we can find ourselves compute-bound (using the full hardware compute capacity) or memory-bound (bottlenecked by memory loading).

Five years ago ML had a colorful landscape of architectures — ConvNets, LSTMs, MLPs, Transformers — but now we mostly just have the Transformer[transformers]. We strongly believe it’s worth understanding every piece of the Transformer architecture: the exact sizes of every matrix, where normalization occurs, how many parameters and FLOPs (FLoating point OPs, basically the total number of adds and multiplies required. While many sources take FLOPs to mean “operations per second”, we use FLOPs/s to indicate that explicitly.) are in each part. Section 4 goes through this “Transformer math” carefully, showing how to count the parameters and FLOPs for both training and inference. This tells us how much memory our model will use, how much time we’ll spend on compute or comms, and when attention will become important relative to the feed-forward blocks.

五年前的 ML 还是架构百花齐放——ConvNet、LSTM、MLP、Transformer——但如今我们基本只剩下 Transformer[transformers]。我们坚信值得把 Transformer 架构的每一块都搞懂:每个矩阵的确切尺寸、归一化发生在哪里、每一部分有多少参数和 FLOPs(「FLoating point OPs」,基本上就是所需的加法与乘法总次数。注意很多资料把 FLOPs 当作「每秒运算次数」,我们用 FLOPs/s 来明确表示后者)。第 4 节 会仔细过一遍这套「Transformer 数学」,演示如何为训练和推理分别统计参数量和 FLOPs。这能告诉我们模型会占用多少内存、会把多少时间花在计算或通信上,以及注意力相对于前馈块会在什么时候变得重要。

Figure: a standard Transformer layer with each matrix multiplication (matmul) shown as a dot inside a circle. All parameters (excluding norms) are shown in purple. Section 4 walks through this diagram in more detail.

Section 5: Training and Section 7: Inference are the core of this book, where we discuss the fundamental question: given a model of some size and some number of chips, how do I parallelize my model to stay in the “strong scaling” regime? This is a simple question with a surprisingly complicated answer. At a high level, there are four primary parallelism techniques used to split models over multiple chips (data, tensor, pipeline, and expert), and a number of other techniques to reduce the memory requirements (rematerialization, optimizer/model sharding (aka ZeRO), host offload, gradient accumulation). We discuss many of these here.

第 5 节:训练 和 第 7 节:推理 是本书的核心,我们在这里讨论那个根本问题:给定某个大小的模型和一定数量的芯片,我该如何并行化我的模型,才能保持在「强扩展」区间?这个问题看似简单,答案却出人意料地复杂。在高层面,把模型切分到多颗芯片上的主要技术有四种(数据并行、张量并行、流水线并行和专家并行),另有一些用来降低内存需求的技术(重算 / rematerialization、优化器/模型分片(即 ZeRO)、主机卸载 / host offload、梯度累积)。我们会在这里讨论其中很多。

We hope by the end of these sections you should be able to choose among them yourself for new architectures or settings. Section 6 and Section 8 are practical tutorials that apply these concepts to LLaMA 3, a popular open-source model.

我们希望读完这些章节后,你能自己为新的架构或场景在这些技术之间做选择。第 6 节 和 第 8 节 是把这些概念套用到 LLaMA 3(一个流行的开源模型)上的实战教程。

Finally, Section 9 and Section 10 look at how to implement some of these ideas in JAX and how to profile and debug your code when things go wrong. Section 12 is a new section that dives into GPUs as well.

最后,第 9 节 和 第 10 节 讨论如何在 JAX 中实现其中一些想法,以及当出问题时如何做性能剖析和调试。第 12 节 是新增的一节,也深入讲 GPU。

Throughout we try to give you problems to work for yourself. Please feel no pressure to read all the sections or read them in order. And please leave feedback. For the time being, this is a draft and will continue to be revised. Thank you!

我们通篇都会给你留一些自己动手做的习题。请不必有压力去读完所有章节或按顺序阅读。也请留下反馈。目前这还是草稿,会继续修订。谢谢!

We’d like to acknowledge James Bradbury and Blake Hechtman who derived many of the ideas in this book.

我们要感谢 James Bradbury 和 Blake Hechtman,本书中的许多想法源自他们。

Without further ado, here is Section 1 about TPU rooflines.

废话少说,这是第 1 节,讲 TPU roofline。

各节链接

This series is probably longer than it needs to be, but we hope that won’t deter you. The first three chapters are preliminaries and can be skipped if you’re already familiar with the material, although they introduce notation used later. The final three parts might be the most practically useful, since they explain how to work with real models.

这个系列可能比它需要的更长,但希望这不会吓退你。前三章是预备知识,如果你已经熟悉这些内容可以跳过,不过它们引入了后面会用到的记号。最后三个部分可能最实用,因为它们讲的是如何对付真实模型。

Part 1: Preliminaries

第 1 部分:预备知识

Part 2: Transformers

第 2 部分:Transformer

  • Chapter 7: All About Transformer Inference. Once we’ve trained a model, we have to serve it. Inference adds a new consideration — latency — and changes up the memory landscape. We’ll talk about how disaggregated serving works and how to think about KV caches.
  • 第 7 章:关于 Transformer 推理的一切。模型训练完之后,我们还得部署它。推理引入了一个新的考量——延迟——并且改变了显存版图。我们会讲分离式服务(disaggregated serving)如何工作,以及如何看待 KV cache。

Part 3: Practical Tutorials

第 3 部分:实战教程

  • Chapter 9: How to Profile TPU Code. Real LLMs are never as simple as the theory above. Here we explain the JAX + XLA stack and how to use the JAX/TensorBoard profiler to debug and fix real issues.
  • Chapter 10: Programming TPUs in JAX. JAX provides a bunch of magical APIs for parallelizing computation, but you need to know how to use them. Fun examples and worked problems.

Part 4: Conclusions and Bonus Content

第 4 部分:结论与附加内容

  • 第 12 章:如何理解 GPU。关于 GPU 的附加章节:它们如何工作、如何组网、以及它们的 roofline 与 TPU 有何不同。

讨论

用 GitHub 账号留言;评论保存在公开仓库chengshu-blog-discussions的 Discussions 里。也可通过 RSS 订阅后续文章。