PJFP.com

Pursuit of Joy, Fulfillment, and Purpose

,

Amin Vahdat of Google: Why Power, Not Chips, Limits the AI Buildout

Amin Vahdat, who runs AI infrastructure at Google, explains how the company builds AI data centers, why it measures them by “goodput” and not chip speed, and why electricity is the one constraint it does not know how to solve. In a 64-minute interview with Sonya Huang of Sequoia Capital, in a year when Google is expected to spend more than $200 billion on capital projects, Vahdat covers the TPU program, co-design with DeepMind, what AI agents do to data center design, optical networking, chip lifetimes and data centers in orbit.

TLDW

An AI data center is a building designed around its hardware, not a general-purpose shell. Google judges it by useful work delivered per watt, because at 100,000 accelerators something fails several times a day. Most efficiency gains have come from models and software, with hardware adding a dependable doubling each year. The rise of long-running agents is pushing up demand for ordinary CPUs, networking and storage as well as accelerators. Power is the binding long-term limit, and Google prefers to work with utilities over generating its own. Seven and eight year old TPUs are still fully used.

Thoughts

Agents are changing what a data center has to be, and this is the least discussed point in the interview. When a person types a prompt, the person is the brake: reading the answer takes seconds, and the machines wait. An agent reads the answer in milliseconds and sends the next request. Vahdat’s second observation matters more. The thinking an agent does between model calls, parsing a response, fetching context from memory or disk, deciding what to ask next, runs on ordinary CPUs. So demand for the unglamorous parts of computing, processors, storage and networking, is “going through the roof” alongside demand for accelerators. That complicates the tidy story that AI is a GPU business. It also undoes some of the efficiency of purpose-built AI buildings, since CPU racks and accelerator racks want different power, density and cabling. The agent boom has a hardware bill that does not show up in chip counts.

On power, the interesting part is how unheroic Google’s approach is. Asked whether it is building its own turbines, Vahdat says the preferred model is to be connected to the grid and to plan with utilities years ahead. The reasoning is arithmetic. A site that must be up 99.99% of the time on its own power needs roughly twice the generation it uses. Share a grid and that redundancy is spread across everyone. Google keeps some local generation to cover gaps and can hand power back during the hottest weeks of the year. Vahdat also says Google pays for the transmission lines and substations built for it, so that other customers’ rates do not rise. That is the commitment critics of the buildout say they want. It deserves to be checked against utility filings, not taken from a podcast, but it is the right standard.

The description of co-design is the clearest account I have seen of what vertical integration buys. DeepMind researchers can ask for a change to a chip that is weeks from manufacture, engineers from both sides sit in a room for a week or two, and Google may delay the tape-out because the gain is worth it. Vahdat is careful to say this is hard across company lines, not impossible. Each layer of the stack might give 10% or 20% or more, and those multiply. The cost is stated just as plainly: design for one stack and you cannot pick up and run anywhere. It is worth holding that next to the later praise of open standards. Google wants both a tightly integrated system and customers who do not feel locked in. Supporting PyTorch on TPUs is real evidence for the second, but the commercial pull is toward the first.

Goodput is a useful idea outside data centers. Throughput is how much work a system does. Goodput is how much of that work counted toward the answer. If one chip among 100,000 fails during a tightly coordinated job, the whole job may stop, roll back to a checkpoint and redo hours of computation. All of that is work, and none of it is progress. So Google holds itself to delivered results under real failure conditions, not to a chip’s rated speed, which Vahdat treats as a number that is true only in theory. Any team that reports activity in place of outcomes could borrow the distinction. Hours worked, tickets closed and lines written are throughput.

One sentence near the end bears on the question investors keep asking about this spending. Vahdat says Google’s seven and eight year old TPUs are still at 100% utilization, and seems surprised the remark drew attention when it was first made. The bear case on AI capital spending rests partly on chips becoming worthless within a few years. If old accelerators still find paying work, the assets last longer than the pessimists assume. Two cautions apply. Full utilization inside Google during a shortage says little about resale value or about a world where supply catches up, and Vahdat notes that Google does retire hardware around the six-year depreciation mark because newer chips use less power for the same work. On orbital data centers, the tone is measured: more sunlight and no batteries, harder cooling and repairs, and “no fundamental showstoppers.”

Key Takeaways

  • A traditional data center is a 25 to 30 year building that hosts many generations of hardware. An AI data center is co-designed with the hardware going into it, down to cooling and power distribution.
  • A storage rack draws 10 to 40 kilowatts. An accelerator rack draws hundreds of kilowatts today, and megawatt racks are being discussed.
  • At 100,000 accelerators, something fails several times a day and possibly several times an hour. There is no dominant cause, only a long tail of hardware, network and software faults.
  • Google’s target is to double effective serving capacity, meaning the ability to generate tokens, about every six months. As much or more of that comes from software as from hardware.
  • Most gains in intelligence per watt have come from the model side. Hardware contributes a doubling or better each year that everything above it can rely on.
  • The TPU began in 2013 as a contrarian bet, when conventional wisdom said custom chips could not beat general-purpose ones riding Moore’s law. The first chip did inference for translation and voice recognition.
  • This year Google shipped two TPUs, one tuned for inference and one for training, because inference had grown large enough to justify its own chip. Each can still do the other’s job.
  • Whether to specialize a chip depends on how big a workload is, how big it will be in three or four years, and whether it lasts. A workload that fades in three months cannot be designed for.
  • Google plans hardware two to five years ahead and has five or six chip generations in progress at once, from concept to production.
  • Gemini is being used to help design hardware for future Gemini models, and Vahdat says hardware engineers now use AI as heavily as software engineers.
  • Optical circuit switches steer light with tiny mirrors, so a failed rack can be replaced by a spare in milliseconds without anyone moving a cable.
  • Training clusters want to be as large and concentrated as possible. Serving wants to be spread around the world, mixed with storage and CPUs, so inference capacity has to be built separately.
  • In a sun-synchronous orbit, solar panels get about 1.4 times the power and near-constant sunlight, against roughly 30% of the day on the ground.

Chapters

01:15 What Makes an AI Data Center Different

The ingredients are the same as any data center: concrete, electrical and mechanical yards, cooling, networking and storage. The difference is specialization. Making a building flexible enough for anything over 30 years means overbuilding it, so AI sites are designed for specific racks. Huang mentions a cluster Google delivered to a Sequoia portfolio company and the photographs of its fiber.

05:01 Goodput: Measuring What Was Delivered

Chip metrics describe a maximum under ideal conditions. Real performance depends on thousands of chips, the CPUs feeding them and the network between them. Vahdat explains goodput with a problem worked on paper: if you must go back to step one, you are still working, but not getting closer. Finding the failed part is a constant search for a needle in a haystack. The section ends with the six-month doubling target.

15:08 The TPU Bet and How Far to Specialize

People inside Google doubted the TPU in 2013. A second chip added training, transformers arrived around then, and recommendation systems and ads followed. On why not build a chip for the transformer alone, Vahdat says the linear algebra is already baked in, and the next step is specializing to a particular model, which some companies are exploring.

22:43 TPUs, GPUs and Co-Design With DeepMind

GPUs are more general, Google sells and uses plenty of them, and customers choose by workload. Vahdat lays out the case for and against co-design and describes daily work with DeepMind, including simulations that predict how future models will run on candidate chips. The TPU’s basic architecture has barely changed since the first version, much as a CPU’s instruction set outlives the software written for it.

34:03 What Long-Horizon Agents Do to the Data Center

Two changes: no human pacing the requests, and far more orchestration on CPUs with data pulled from memory and disk. Putting CPU racks beside accelerator racks spoils the uniform design. Putting them in the next building adds networking cost, reliability risk and latency.

37:49 Optical Circuit Switching

About 15 years ago Google brought wavelength division multiplexing into its data centers, along with switches that route light without converting it to electrical signals. They were first used to create shortcuts between clusters and to grow the network without rewiring. Asked why not drop fiber entirely, Vahdat cites signal loss and the difficulty of aiming beams across a large building.

42:50 Power: The One Constraint Without a Solution

Every constraint is hard and they shift, Vahdat says, but the others are solvable with time. Power is not, short of abundant clean energy such as nuclear at scale. A gigawatt cannot be ordered for tomorrow. If a utility can supply 700 megawatts by the date needed, Google can wait or fill the gap with solar and batteries.

47:51 How Big to Build, and How Long Chips Last

Google once debated putting everything in a single gigawatt site. A single point of failure and the impossibility of getting all its power in one place ruled that out. Sites now range from tens of megawatts at the network edge to near a gigawatt for training. Hardware is replaced by the pod, about 9,600 chips across some 150 racks, and a new pod rarely fits the old one’s footprint.

54:08 Open Standards and AI Inside the Team

Vahdat compares open interfaces to the internet protocol, which beat rivals because anything could run above it and below it. Google favors its JAX framework but supports unmodified PyTorch models on TPUs. Inside the team, time from design start to tape-out is shrinking, and AI now gathers the information for decisions such as campus size, without replacing the people who make them.

59:14 Orbital Data Centers and the Supercomputer of 2036

Google is pursuing computing in orbit as a moonshot. Cooling is harder in space than people expect, repairs are harder, and links would be lasers through open space. Ten years out, Vahdat expects racks built centrally, far more integrated, possibly drawing several megawatts each, and installed by connecting water, power and a small bundle of fiber.

Notable Quotes

“It’s delivered goodput, not theoretical benchmark throughput or goodness in theory.”

Amin Vahdat, on how Google holds its infrastructure accountable

“We’re living in a world right now where 2x or more year-over-year performance improvements is absolutely possible.”

Amin Vahdat, on what hardware contributes each year

“We’re using Gemini to design hardware for future Geminis as well.”

Amin Vahdat, on AI in chip design

“There’s no human in the loop that is going to naturally rate limit how quickly requests are going to go to the model.”

Amin Vahdat, on how agents change the load on a data center

“Power is the single most fundamental constraint that we face.”

Amin Vahdat, when asked to name the biggest limit on the buildout

“Our seven and 8 year old TPUs are still at 100% utilization.”

Amin Vahdat, on the useful life of AI chips

“There are no showstoppers here. No fundamental showstoppers.”

Amin Vahdat, on data centers in orbit

Vahdat explains the engineering more clearly than most executives explain a strategy. Watch the full interview here.

Related Reading