PJFP.com

Pursuit of Joy, Fulfillment, and Purpose

Tag: Nvidia Rubin

  • Why the Markets Are Pricing AI Wrong: Gavin Baker on the July 2026 Selloff, GPU Spot Prices, Memory LTAs, and Nvidia’s Credit Wrapper

    Gavin Baker of Atreides Management returned to Invest Like the Best with Patrick O’Shaughnessy days after one of the strangest months the AI trade has ever produced. AI and semiconductor names fell 40 to 60 percent in a straight line while, by Baker’s account, not a single quantitative metric on the ground deteriorated. He spent the week in Silicon Valley hunting for a bearish data point and came back with almost nothing except credit. This conversation is the result: a detailed argument that the market has mispriced the gap between contracted compute and spot compute, that open source is growing the infrastructure pie rather than shrinking it, and that the one risk actually worth fearing is political rather than financial.

    TLDW

    Gavin Baker describes July 2026 as “2022 packed into a single month,” a violent AI and semiconductor drawdown that happened while hyperscaler operating cash flow accelerated from roughly 28 percent growth to 32 percent, or closer to 35 percent adjusting for unusual legal charges. His core claim is that the installed base of GPU compute is locked into long-term contracts priced far below the current spot market, so as those contracts roll off, compute reprices higher, operating cash flow accelerates, and the buildout can be funded internally rather than with the debt that widening credit default swap spreads and a poorly received Meta bond have made look expensive. He walks through each catalyst of the selloff: Meta renting out compute (misread as a capex cut), the open source capability leap from GLM 5.2 and Kimi K3 (misread as deflationary when a token is a token and costs the same flops, watts, and memory to produce), China acquiring a domestic deep ultraviolet lithography machine (real but 25 years behind), and rising real yields (the only genuine negative). He covers the game theory of breaking a memory long-term agreement in a world where market share is set by supply allocations, Nvidia’s new credit wrapper plus revenue share model and why it is misunderstood, the router and fine-tuning stack from Fireworks and Baseten that turns “ChatGPT wrappers” into defensible AI natives, continual learning as the one technical development that could disrupt training demand, SRAM accelerators for disaggregated inference, SpaceX as an underappreciated compute company with orbital ambitions, and his view that regulation, not fundamentals, is the biggest risk because the industry has done a terrible job telling its own story. He also makes an unusual observation about market structure: everyone now feeds news into Claude, and Claude has become a kind of Walter Cronkite for the stock market, collapsing the diversity of interpretation that normally keeps markets stable.

    Thoughts

    The load-bearing claim in this episode is the spread between contracted and spot compute, and to Baker’s credit it is falsifiable in a way most bull cases are not. He is not arguing that AI will be transformative or that demand feels strong. He is arguing something narrow and checkable: hyperscalers and neoclouds signed multi-year GPU contracts in 2024 and 2025 at prices that assumed a gentle decline, prices instead went vertical, and the installed base is therefore systematically under-earning. A startup rented several thousand B200s in the mid two dollars per GPU hour range and expects to pay just under four dollars for an identical cluster seven months later. If that repricing is real and broad, hyperscaler operating cash flow mechanically accelerates and roughly 700 billion dollars of projected credit demand evaporates. If GPU rental prices roll over and stay down for two consecutive quarters, the thesis is dead. That is the number to watch rather than any earnings headline. The caveat he steps past quickly is that the open source mix shift he describes as bullish does not eliminate margin, it relocates it, out of the frontier labs and down into the infrastructure layer. Excellent if you sell GPUs, power, and memory. Considerably more awkward for the labs whose projected cash flows are the reason anyone believes the compute gets paid for at all.

    The Claude as Walter Cronkite observation deserves more attention than it got, where it passed as a joke. Baker is describing a genuine change in market microstructure. Every institutional and retail participant now feeds the same news into roughly the same models, and while those models are probabilistic, they are not producing meaningfully diverse readings of the same headline. He connects this to Michael Mauboussin’s argument that a breakdown in diversity, not leverage alone, is what produces bubbles and crashes. If that is what happened in July, then the Japanese capacitor stock chart he cites, an entire three-year cycle compressed into six weeks before the fundamentals had even arrived, is not a curiosity. It is the signature of a market where thousands of participants share one interpretive engine. That makes drawdowns faster and deeper without making them more informative, which argues for holding through machine-generated narrative cascades rather than trading them.

    The middle of the conversation contains the most consequential business idea in it, and it is one that got almost no coverage during the selloff: memory long-term agreements and Nvidia’s credit wrapper are the same move executed at two different layers of the stack. Both trade near-term upside for durability. The memory companies stopped maximizing spot price and started signing prepaid agreements with floors and ceilings, and the reason those agreements will hold is that the penalty for breaking one has changed category. Apple could renege on memory pricing for years because its volume was overwhelming and it had no equivalent competitor. In a world with four buyers that matter and where AI market share is set by supply allocation rather than product quality, a supplier can answer a broken price agreement by breaking the volume commitment and handing your allocation to a rival, in an industry where oversupply is always followed by undersupply. Nvidia is running the same play one layer up. The credit wrapper with a revenue share above a price floor converts a cyclical one-time chip sale into a royalty on recurring compute revenue, financed on someone else’s balance sheet, which is a materially better business than selling hardware. It also widens the moat, because a startup accelerator pays more at the foundry, pays more for high bandwidth memory, and cannot finance its chips at Nvidia’s rate. Baker is right that this is misunderstood, and it is a strange thing for a stock at a ten-year-low forward multiple to be quietly doing.

    The technical material in the back half reveals an asymmetry worth naming. Baker treats two efficiency developments very differently. Continual learning and sample efficient learning, which several labs believe are close, would collapse the token budget required to produce a capable model, and he handles this by asserting that training asymptotes to a small but nonzero share of compute and that the outcome would be wonderful for the world anyway. SRAM-based accelerators for disaggregated inference, running prefill on one chip, attention on a high-memory chip, and the feed forward network on SRAM, he embraces enthusiastically as a return-on-investment improvement across the installed base. Both are efficiency gains. One is treated as neutral, the other as clearly positive, and Jevons paradox is doing all the work in both directions. That is probably correct given everything we have observed so far, but it is an assumption rather than a finding, and it is the assumption on which the entire “cheaper compute is bullish for compute” framework rests. Worth noting too that the SRAM disaggregation point is genuinely underdiscussed: those chips sit on older nodes and do not compete for leading-edge capacity, so they are additive supply rather than substitute supply.

    The final twenty minutes hold both the largest unpriced upside and the largest unpriced risk, and neither is in consensus estimates. On the upside, only the hyperscalers, CoreWeave, Crusoe, and SpaceX have ever brought more than 500 megawatts online in a single year, and SpaceX has done it fastest and cheapest. When it dumped a large block of compute into the market, the market absorbed it without a blip, which tells you more about demand than any survey. Baker’s sanity check on orbital compute is the sharpest reasoning move in the episode: Benchmark, from entirely outside the Elon ecosystem and without the benefit of internal launch costs, funded StarCloud at a real valuation, so the set of people who would all have to be wrong keeps growing. On the downside, regulation is the risk he names first and it is the one his own framework cannot arbitrage. New York’s data center moratorium is not a fundamentals problem, and no amount of operating cash flow acceleration fixes a permitting ban. His diagnosis is that the industry finds the benefits so obvious that it never learned to explain them, which is how a water usage figure overstated by four orders of magnitude became conventional wisdom. Proposing a foundation that buys World Series ad time is a tell about how far behind he thinks the industry is. Every other risk in this conversation is priced somewhere. That one is not.

    Key Takeaways

    • Baker characterizes July 2026 as “2022 in a month,” with AI names down 40 to 60 percent from their highs in a straight line while underlying fundamentals improved.
    • He spent the week in Silicon Valley explicitly hunting for a negative quantitative metric and found essentially one: third-party data suggesting Anthropic’s growth curve came slightly off trajectory, a data point Anthropic shareholders reportedly dispute.
    • Nvidia was trading at its lowest forward price to earnings multiple in ten years at the time of recording. The only cheaper moments were the DeepSeek shock and Liberation Day, both of which proved to be V-bottoms.
    • A low forward multiple means the market believes these companies are significantly over-earning. Baker’s counter is that they are under-earning because their installed compute is contracted below spot.
    • Combined operating cash flow at Microsoft, Meta, and Amazon accelerated from roughly 28 percent to 32 percent growth, or to about 35 percent after adjusting for an unusual quarter of legal and regulatory charges.
    • Nobody in 2024 or 2025 modeled old GPU prices going vertical in 2026. The bull case assumed a slow decline in rental rates and the bear case assumed a steep one.
    • A concrete example: a well-known startup rented several thousand Blackwell B200s in the mid two dollars per GPU hour range and expects to pay just under four dollars for an identical cluster seven months later, a 50 to 60 percent increase.
    • One inference cloud stated publicly that it plans to pay roughly 100 percent more for Blackwells when its current contract expires.
    • Neoclouds were often forced into below-market long-term contracts because they needed an offtake agreement to finance the GPUs in the first place.
    • Consensus models hyperscalers monetizing Blackwell and Rubin at roughly Ampere rates, two generations behind, producing about 1.3 to 1.4 trillion dollars of hyperscale operating cash flow. Assuming monetization merely at a discount to current Blackwell rates pushes that closer to two trillion and removes roughly 700 billion dollars of credit demand.
    • The credit concerns are real and undeniable: real yields are up, spreads have widened, credit default swap levels for the large buyers have blown out, and a recent Meta bond did not price where a Meta bond should price.
    • Baker’s response is that debt-fueled buildouts demand immediate repayment and unwind violently, which is what happened in the internet buildout, but this buildout is still overwhelmingly funded from operating cash flow.
    • If credit is not available, he argues the existing flops simply become more valuable, which is self-correcting rather than catastrophic.
    • The Meta selloff catalyst was a misread. Meta renting out compute was interpreted as excess capacity and a capex cut. Meta did not cut capex, and the actual motivation appears to have been demonstrating strong internal rates of return on a small slice of capacity ahead of a capital raise.
    • The open source panic was also a misread. Open source taking token share moves margin dollars out of the frontier model layer, but a token still requires the same flops, memory, and watts to produce, so infrastructure demand rises rather than falls.
    • Frontier tokens carry gross margins somewhere in the 80 to 95 percent range. Open source tokens might carry 30 percent. The customer’s savings come almost entirely out of that margin, not out of compute consumption.
    • Baker calls open source “dark matter to the public markets,” growing rapidly through GLM 5.2, Kimi K3, and Nvidia’s Nemotron, but nearly impossible for public investors to measure since it runs through private inference clouds.
    • Jensen Huang being the world’s loudest supporter of open source is itself evidence that open source is good for Nvidia’s business.
    • Enterprises that blow through their AI budget in three months set up a router, which cuts their spend but often increases total GPU hours consumed by shifting volume to cheaper open source tokens.
    • Adoption is happening in staggered waves: AI natives are all in and hiring very few humans, coastal public companies are optimizing, East Coast and non-coastal companies have barely adopted, and Europe is trying to regulate AI before using it.
    • Roughly 500,000 people worldwide use agentic AI, and perhaps half that number use it seriously, yet the world is already in an acute compute shortage. The relevant question is what happens at 100 million or 500 million users.
    • Token spend at the most AI-forward companies now runs 20 to 25 percent of total compensation spend, with individual examples at 30 percent and reports as high as 50 percent, against a roughly 25 trillion dollar global knowledge work market.
    • Founder-controlled companies are not conducting large-scale layoffs, which suggests the cash flow to pay for AI is expected to come from growth rather than from labor substitution.
    • Memory is the dominant variable in token economics. More memory per unit of compute yields more tokens out, which lowers cost per token, which is why demand has shown no negative elasticity to memory pricing.
    • Memory suppliers have shifted from maximizing near-term price to signing long-term agreements with prepayments, floors, and ceilings, trading short-term upside for durability.
    • Breaking a memory long-term agreement is now potentially fatal. With four buyers that matter at scale and market share determined by supply allocation, a supplier can respond by breaking the volume commitment and handing your allocation to a competitor.
    • This is structurally different from the Apple era, when a single dominant buyer could break pricing agreements without consequence.
    • Nvidia’s new model is best described as a credit wrapper with a revenue share triggered when GPU prices exceed a floor. It is not vendor financing, since a third party lends the money, and it could produce a very large cloud-scale royalty business quickly.
    • Baker thinks this model is badly misunderstood, meaningfully increases Nvidia’s revenue per gigawatt, and strengthens its competitive position against startup accelerators that pay more at the foundry, pay more for high bandwidth memory, and cannot finance their chips as cheaply.
    • Nvidia has taken equity stakes across the ecosystem, and Baker’s read is that every time they have not taken a stake it has proven to be a mistake.
    • The scenario that would genuinely frighten him: hyperscaler operating cash flow stops accelerating, forcing the buildout onto debt, or a sustained sharp contraction in GPU rental prices. Nobody he has spoken to says they have too many GPUs.
    • Continual learning and sample efficient learning are the technical developments most likely to disrupt training demand, and several new labs including Safe Superintelligence are focused on them. Baker still thinks training asymptotes to a small share of compute rather than to zero, and that the change would be enormously good for the world regardless.
    • Fireworks launched a product called Nexus that plugs into Claude Code, OpenAI Codex, or Grok in roughly three lines of code, ingests a customer’s data, applies reinforcement learning to a model, and routes queries appropriately.
    • This stack is what converts an alleged “ChatGPT wrapper” into a defensible company. Shifting 30 to 60 percent of token consumption to a customized open model on top of frontier orchestration produces better outcomes at roughly half the cost.
    • Cheap, capable open source models may actually inflate the value of the very best frontier model, since a 160 IQ orchestrator becomes more valuable when it has an army of cheap 120 IQ models to direct.
    • The inference clouds are growing almost as fast as the frontier labs did in their early days while burning very little cash, which is extraordinary by any conventional software metric.
    • China obtaining a domestic deep ultraviolet lithography machine is a genuine phase transition and should not be dismissed, but the technology is roughly 25 years behind extreme ultraviolet, and lithography progress is learning by doing that cannot be teleported through.
    • Baker considers regulation the biggest single risk to AI, citing New York’s data center moratorium as the first of many and describing the current environment as post-factual and post-logical.
    • The public narrative that data centers raise power bills, drain water, and destroy jobs is largely wrong. Behind the meter deals typically lower local electricity prices, and modern community agreements include hospitals, schools, police and fire stations.
    • The widely cited data center water figure originated in a published error overstating usage by roughly 10,000 times, since acknowledged by the author, which Baker likens to the decimal point error that created the myth that spinach is exceptionally high in iron.
    • He argues data centers are among the best things to happen to blue collar wages in his lifetime, with ongoing rather than one-time employment from maintenance, replacement, and upgrade cycles.
    • SRAM-based accelerators built on older nodes and free of high bandwidth memory constraints could substantially improve return on investment by allowing disaggregated inference: prefill on one chip, attention on a high-memory chip, and the feed forward network on SRAM.
    • SpaceX has improved fundamentally since going public, and Baker believes the market does not yet understand it as a compute company. Only the hyperscalers, CoreWeave, Crusoe, and SpaceX have ever brought on more than 500 megawatts of power in a single year, and SpaceX has done it fastest and cheapest.
    • A widely circulated report claims SpaceX intends to bring on eight gigawatts of compute in 18 months. Baker doubts the number but notes that at roughly 50 billion dollars of monetization per gigawatt, even a fraction of it dwarfs the current consensus estimate.
    • When SpaceX dumped a large block of compute into the market, it was absorbed without a blip, which Baker reads as one of the more bullish demand signals of the year.
    • Orbital compute feels more real every day. Benchmark funding StarCloud, from outside the Elon ecosystem and without access to internal launch costs, functions as a useful sanity check on the idea.
    • Dark horse names Baker flags for the next phase: Lip-Bu Tan, Lin Qiao at Fireworks, and Scott Wu at Cognition.

    Detailed Summary

    A Selloff That Contradicted Every Fundamental

    Baker opens by describing July 2026 as 2022 compressed into a single month. AI names fell 40 to 60 percent from their highs in a nearly straight line. What made the month unusual was not the magnitude but the absence of a legible cause. In 2022 the market feared recession, rising rates, and inflation. During the DeepSeek shock and Liberation Day you knew exactly what the market was reacting to. This time the fundamentals moved in the opposite direction from the tape. GPU availability tightened, GPU rental pricing rose, DRAM spot prices rose, and token growth accelerated. Baker asked Patrick, who had also spent the summer in Silicon Valley, whether he had heard a single negative quantitative metric or a single instance of deceleration. The answer was nothing.

    Part of the problem is visibility. Public markets cannot see Anthropic or OpenAI directly, and they cannot see the American open source inference clouds like Fireworks, Baseten, Modal, and Together that monetize inference. Everyone stares at the same chart of semiconductor cash flow rising while hyperscaler free cash flow falls, and that chart omits the private companies entirely. It also omits the repricing dynamic Baker considers the most important fact in the market.

    The Spot Versus Contract Gap

    In 2024 and 2025 every serious forecast assumed GPU rental prices would decline, with the only debate being how fast. Neoclouds locked in long-term contracts partly out of prudence and partly because they needed offtake agreements to finance the hardware at all. The result is a large installed base of contracted compute trading at a steep discount to today’s spot market. Baker’s argument is that as those contracts roll off, compute reprices higher even if spot itself declines from current levels, and that repricing flows directly into hyperscaler operating cash flow.

    The anecdotes are stark. A prominent startup rented several thousand B200s in the mid two dollar per GPU hour range and expects to pay just under four dollars for an identical cluster seven months later. One inference cloud said publicly it plans to pay roughly double for Blackwells at contract renewal. Baker’s read is that hyperscalers are therefore under-earning across the board, which is the exact opposite of what a ten-year-low forward multiple implies the market believes.

    Financing the Buildout and the Credit Question

    Credit is the one bearish input Baker concedes is real. Real yields have risen, spreads have widened, credit default swap levels have blown out across the large buyers, and a recent Meta bond did not price the way a Meta bond should. Sophisticated private capital investors told him this is just banks hedging commitments, but he acknowledges the optics are bad and the facts are undeniable. His concern is the classic capital cycle: debt-financed buildouts demand immediate repayment, so when supply and demand slip out of alignment the unwind is fast and brutal, exactly as it was in the internet buildout.

    The math he ran is the counterweight. Consensus effectively models hyperscalers monetizing Blackwell and Rubin at Ampere rates, two generations behind, producing 1.3 to 1.4 trillion dollars of operating cash flow. Assume instead that they monetize merely at a modest discount to current Blackwell rates and the figure approaches two trillion, taking about 700 billion dollars of credit demand off the table. Better cash flow also improves the credit ratios, which makes debt cheaper if they choose to use it. And if credit disappears entirely, the flops already installed simply become more valuable. Microsoft brought on a large slug of capacity in June that did not even appear in second quarter results.

    How the Month Actually Unfolded

    Baker walks the sequence of catalysts. First, Meta announced it would rent out compute, which the market read as excess capacity and an imminent capex cut. Meta did not cut capex. What Meta appears to have seen was SpaceX selling trading-optimized clusters into the market at an enormous premium to contracted rates, and the plan was likely to demonstrate strong returns on a small slice of capacity before raising equity capital and increasing capex. Shortly afterward Meta released its best model in a long time, overshadowed by a competing release but a clear signal it was not easing off.

    Next came the open source freakout. Kimi K3 arrived, the widely watched token index dipped and flattened, and the two were connected: the index captures mix, and a shift from expensive frontier tokens toward open source tokens looks like weakness even when total compute consumption is rising. Then China’s deep ultraviolet lithography news triggered a broad selloff in semicap equipment. Finally, rising real yields and widening spreads gave the market a genuine reason to worry. Baker’s summary is that with the sole exception of credit, every one of these narratives was factually wrong, and a friend at Fidelity described the winning strategy of the past three years as doing the dumbest, most superficial thing as fast as possible and cycling between them.

    Open Source as Dark Matter

    The most important conceptual argument in the episode is that a token is a token. Regardless of which model produces it, a token consumes the same flops, the same memory, and the same watts. Open source taking share therefore does not reduce compute demand. It transfers margin from the frontier model layer, where gross margins might be 90 percent, to open weights inference at perhaps 30 percent, and the resulting price decline drives elasticity in token volume. Since frontier labs and open source models both run on the same underlying cloud infrastructure at the same compute cost, the effect is to push margin dollars down into the infrastructure layer.

    Baker calls open source dark matter to public markets. It is real, it is accelerating on the back of capability leaps from GLM 5.2 and Kimi K3, Nvidia continues to push Nemotron closer to the frontier, and yet none of it appears in audited financials that public investors can underwrite. He also notes the tell that should have settled the debate: Jensen Huang is the world’s most vocal supporter of open source, which would be an odd position for the largest beneficiary of frontier concentration to hold if open source actually threatened the business. Baker adds a normative point, that a world with only one or two dominant frontier models charging 90 percent margins is not good for humanity, and that many models is the better outcome.

    Routers, Fine-Tuning, and the End of the Wrapper Insult

    The practical mechanism behind the open source surge is the router plus fine-tuning stack. Inference clouds have become genuinely good at supervised fine-tuning and reinforcement learning, so a company can take its proprietary data, customize an open weights model, put it behind a router, and have the router send most queries to that model while escalating to a frontier model for verification or harder work. The result is often slightly better outcomes at half the cost. Fireworks shipped a product called Nexus that connects to Claude Code, OpenAI Codex, or Grok in roughly three lines of code and handles ingestion, reinforcement learning, and routing.

    This changes the durability question for AI natives. Two years ago the criticism was that these companies were thin wrappers with no defensibility. Now a company with domain-specific proprietary data can train on it, own the model serving 30 to 60 percent of its tokens, and get off the frontier lab treadmill it previously had no choice but to accept. Baker points to Cursor, Harvey, and others leaning hard into this. He also raises the counterargument fairly: some believe that once a frontier model achieves recursive self-improvement it will serve every intelligence level more cheaply through distillation, leaving no room for open source. He does not dismiss it, but he thinks the proprietary data held by AI natives and the orchestration value of the single smartest model make the multi-model future more likely. Cheap 120 IQ models arguably make a 160 IQ orchestrator more valuable, not less.

    Where the Money Comes From

    The pushback Baker gets on X is fair: even if hyperscalers are under-earning, where does the customer revenue ultimately come from? Definitionally it must come from faster economic growth through productivity or from labor substitution. He sees labor substitution happening at AI natives, though not through firing. They simply never hire the humans, and gross profit dollars per full-time employee at these companies is vertical compared with prior startup generations. Token spend now runs 20 to 25 percent of total compensation spend at the most aggressive companies, with individual examples at 30 percent and reports as high as 50 percent, against a roughly 25 trillion dollar global knowledge work market.

    The encouraging signal is that founder-controlled companies, the ones most likely to move fast on efficiency, are not conducting large-scale layoffs once you adjust for pandemic-era overhiring. That suggests they see continued opportunity for people plus large token budgets rather than a straight substitution. Data from Cognition, Ramp, and Stripe indicates that companies spending the most on AI are growing meaningfully faster, though Baker acknowledges the skeptics’ point that these datasets do not control for industry.

    The Memory Supply War and LTA Game Theory

    Everything is currently in shortage, and Baker argues the constraint is energizing gigawatts rather than manufacturing. Turbine makers and diesel generator makers are ramping, old aircraft turbines are being stripped and reconditioned for data center power, and regulatory policy is moving favorably. The transition he says he got wrong is the shift, especially in memory, from maximizing short-term pricing to signing long-term agreements with customer prepayments, price floors, and price ceilings.

    The reason those agreements will hold is game theory. Memory is the axis around which everything else revolves, because more memory per unit of compute means more tokens out, which lowers cost per token, which is why demand has shown essentially no negative elasticity. Market share among the four buyers that matter (Amazon with Trainium, Google with TPUs, AMD, and an Nvidia bigger than all of them combined) will be determined for years by supply chain allocation. Break a long-term agreement to chase a lower price in an oversupply year and the supplier can break the volume commitment in return and hand your allocation to a competitor. Since oversupply in this industry is reliably followed by undersupply, that is a decision that can end a franchise. Apple could get away with this historically because its volume was overwhelming and it had no equivalent competitor. That world is gone.

    Nvidia’s New Playbook

    Baker finds Nvidia’s low multiple hard to reconcile with how thoroughly the current environment favors it. If chips need to be financed, nothing on earth is more financeable than an Nvidia GPU. If land and power are the constraint, Nvidia has been playing the matchmaking chess game well. On top of that they have rolled out what Baker describes as a credit wrapper with a revenue share that kicks in when GPU prices sit above a floor. It is not vendor financing, since someone else lends the buyer the money. What it does is give Nvidia a royalty on recurring compute revenue, which could amount to a very large cloud business built entirely out of royalties, while helping bridge the cash flow mismatch between an industry that has gone free cash flow negative and a supplier collecting all the cash.

    Asked what he would do as a memory CEO, Baker says he would do exactly what Nvidia is doing: approach GPU and accelerator buyers, participate in the credit wrapper, perhaps put up cash upfront to make lenders comfortable, and take a cut of ongoing revenue. He expects firms like Blackstone and Apollo are pitching variants of this to the memory companies already. He also thinks the arrangement quietly widens Nvidia’s competitive moat, since startup accelerator companies pay more at the foundry, pay more for high bandwidth memory, and cannot finance their chips at Nvidia’s rate. And he notes that essentially every time Nvidia has declined to take an equity stake in something, it has turned out to be a mistake.

    What Could Break the Thesis

    Pressed for the scenario that would flip him, Baker names two. The first is operating cash flow failing to accelerate, which would force the buildout onto debt and validate the credit bears. That outcome depends largely on whether the combined trajectory of Anthropic, OpenAI, Grok, Cursor, and open source keeps compounding. The second is a sustained sharp contraction in GPU rental prices. The market would react instantly, and it would mean the compute shortage had broken. As of the recording, not a single person he has spoken with says they have too many GPUs.

    The technical wildcard is continual learning and sample efficient learning. Many researchers believe both are close. A human learns effectively on something like 20 billion tokens while frontier models train on 300 trillion, so a model that could be trained on 10 trillion tokens and then learn efficiently in the world would represent a discontinuity in training demand. Baker thinks training will asymptote to a small but nonzero share of compute regardless, and that the development would be extraordinarily good for the world. He also notes Nvidia is deeply involved with essentially all of the labs pursuing it.

    China, Lithography, and Decoupling

    On China’s deep ultraviolet lithography machine, Baker holds both views at once. It is a genuine phase transition, comparable to going from having no propeller plane to having one, because they did not have it before and now allegedly they do. It is also roughly 25 years behind extreme ultraviolet, and lithography is learning by doing, so you cannot teleport through the required cycles. He suspects the market overreacted and that if it ever affects ASML’s order book it will be years out, by which time the market will have forgotten and rediscovered the concern several times.

    He is careful about certainty here. It is very hard for an American to have real clarity on what is happening inside China, the people there are extremely capable and work brutally hard, and they consider this existential for the country. There are unverified reports that an extreme ultraviolet machine was smuggled in, which he treats as noise. His larger point is that decoupling is now self-reinforcing on both sides, it is unfortunate, and neither side is going to stop.

    Regulation, Data Centers, and a Failure of Storytelling

    Asked for the worst thing that could happen to AI, Baker answers regulation without hesitation. New York’s data center moratorium feels like the first of many, and even deep red pro-growth states are telling the industry it is doing a poor job explaining itself. The political narrative among ordinary Americans is that data centers will raise electricity prices, drain water supplies, and eliminate jobs. Baker’s counter is that behind the meter deals generally lower local electricity prices, that community agreements now routinely include hospitals, schools, police stations, and fire stations rather than the old model of buying the fire department new trucks, and that the jobs are ongoing rather than one-time because of continuous maintenance, replacement, and upgrade cycles.

    The water claim is the clearest case of a myth outrunning the correction. An author overstated data center water usage by roughly 10,000 times, has acknowledged the error repeatedly, and the figure still circulates. Patrick offers the parallel of the spinach iron myth, created by a misplaced decimal point in an academic text and still believed 80 years later. Baker’s proposed remedy is blunt: a foundation or political action committee running ads during the Final Four, NFL games, and the World Series explaining what a data center actually does for a community, alongside the story of AI accelerating medical research and improving outcomes for people with serious illness. The people building this find the benefits so obvious that they assume everyone already knows, and they cannot process how divergent their view is from most Americans.

    SRAM Accelerators and Disaggregated Inference

    An underdiscussed development, Baker argues, is what happens when SRAM-based accelerators arrive at scale. These chips are not constrained by high bandwidth memory and are often built on older nodes, so they do not compete for the leading edge capacity that GPUs consume. Inference disaggregates into prefill and decode, and decode splits further into attention and the feed forward network. The holy grail is running prefill on a chip without high bandwidth memory, attention on a high-memory chip, and the feed forward network on SRAM, which nothing beats for that workload. Since workloads keep changing, no single chip can get the ratio of compute to high bandwidth memory to on-die SRAM permanently right, which is precisely the argument for disaggregation. Baker expects this to be strongly positive for the return on investment across the installed base and on new compute.

    SpaceX, Orbital Compute, and Dark Horses

    Baker does not think the market understands SpaceX as a company yet, and he considers it the most important new public company. The fundamentals have improved since the IPO, and the compute story is the part being missed. Only the hyperscalers, CoreWeave, Crusoe, and SpaceX have ever brought more than 500 megawatts of power online in a single year, and SpaceX has done it fastest and cheapest while building clusters customers actually like. When SpaceX dumped a large block of compute into the market, it was absorbed without a blip, which Baker treats as one of the most bullish demand datapoints available. A circulating Substack report claims eight gigawatts within 18 months. He doubts that figure and quotes it only because it is public, but at roughly 50 billion dollars of monetization per gigawatt against a 73 billion dollar consensus estimate, even partial delivery would overwhelm expectations. There is a well-known New York hedge fund short case built on spot compute prices falling 90 percent.

    On orbital compute, Baker says time at Starbase left him thinking it feels more real every day, and the Starship landing reinforced it. His sanity check is that Benchmark, from entirely outside the Elon ecosystem and without the benefit of internal launch costs, chose to fund StarCloud at a real valuation, with SpaceX partnering to provide the Starlink laser technology that orbital compute requires. As he puts it, maybe he is crazy, maybe Elon is crazy, maybe Benchmark is crazy, and maybe the SpaceX engineers are crazy too, but all of that being true simultaneously does not seem probable. Asked for dark horses who could become as consequential as the current giants, he names Lip-Bu Tan, Lin Qiao at Fireworks, and Scott Wu at Cognition. The episode was recorded at Benchmark’s offices, at the table where their dinners are held.

    Notable Quotes

    “I want to be scared. I don’t want to feel like a lunatic watching these stocks get cheaper thinking the expected forward returns are going up.”

    Gavin Baker, on why he spent the week in Silicon Valley hunting for bearish data

    “I would describe July as 2022 in a month.”

    Gavin Baker, characterizing a 40 to 60 percent drawdown in AI names that happened in a straight line

    “Have you heard a single negative quantitative metric about AI? A single instance of deceleration?”

    Gavin Baker to Patrick O’Shaughnessy, framing the central contradiction of the month

    “A token is a token, and you need the exact same amount of compute to make a token. It takes the same amount of flops, the same amount of memory, the same amount of watts.”

    Gavin Baker, on why the open source panic misread infrastructure demand

    “Open source is kind of dark matter to the public markets. It’s hard for public markets to measure it.”

    Gavin Baker, on why the fastest-growing part of inference demand is invisible in audited financials

    “Claude is kind of Walter Cronkite for the stock market and everybody just believes whatever it says. And by the way, it’s really smart, but it’s not always right.”

    Gavin Baker, on the collapse of interpretive diversity among investors

    “Nvidia is actually, as we record this, at its lowest forward PE of the last 10 years.”

    Gavin Baker, noting the only cheaper moments were the DeepSeek shock and Liberation Day, both V-bottoms

    “If you break your LTA and then in the next two or three years for any reason leverage shifts back to the memory guys, you’re out of business.”

    Gavin Baker, on why long-term agreements will hold through the next memory cycle

    “If you need to be able to finance the chips, and you do, nothing’s more financeable than an Nvidia GPU. Nothing.”

    Gavin Baker, on why the current environment favors Nvidia more than its multiple suggests

    “Data centers are in a lot of ways the best thing to happen for blue collar wages in my lifetime.”

    Gavin Baker, on the gap between the political narrative and the local economics

    “A lie could go around the world faster than truth gets out of bed.”

    Gavin Baker, on a data center water usage figure overstated by roughly 10,000 times that still circulates

    “One of Elon’s phrases is we specialize in making the impossible late.”

    Gavin Baker, on why he doubts the eight gigawatt figure without betting against SpaceX

    Watch the full conversation here: Why the Markets Are Pricing AI Wrong with Gavin Baker on Invest Like the Best.

    Related Reading

    • Invest Like the Best on Colossus the show’s home, where the full episode archive and transcripts live.
    • Atreides Management Gavin Baker’s firm and the vantage point behind these compute and semiconductor calls.
    • More Than You Know by Michael Mauboussin, the source of the diversity breakdown framework Baker invokes to explain why markets crash when everyone reasons the same way.
    • High Bandwidth Memory (Wikipedia) background on the memory technology that sits at the center of the long-term agreement game theory.
    • Fireworks AI the inference cloud whose routing and fine-tuning stack Baker credits with making open source models competitive for production workloads.
  • Elon Musk Announces SpaceX AI Satellites, Starship Mass to Orbit, and a Moon Mass Driver to Climb the Kardashev Scale

    Elon Musk sat down with the SpaceX Starlink team for a wide ranging update that connects every recent SpaceX move into one thesis: harness far more of the sun’s energy by putting AI compute in orbit. In this SpaceX conversation, the group walks from galaxy sized framing (the Kardashev scale) all the way down to the engineering specifics of a new AI satellite, the manufacturing buildout in Bastrop, Texas, and a long term plan that ends with a mass driver on the moon. The pitch is that none of it requires magic, just scaling technology SpaceX already flies.

    TLDW

    Musk frames civilizational progress with the Kardashev scale, a measure of how much power a species harnesses, and points out that humanity uses less than a trillionth of the sun’s output, barely registering even on the Type 1 (planet) level. Because most of Earth is water and the usable sunlit land is limited, the only way to capture a meaningful fraction of the sun’s energy is to go to space, where cooling is also easier since heat radiates straight into the vacuum. Three limiting factors must be solved: mass to orbit (handled by fully and rapidly reusable Starship, which already beats the Saturn V on thrust and aims for millions of tons to orbit per year), solar power plus radiators, and AI chips. SpaceX unveils its first AI satellite design, AI1, a roughly 70 meter wingspan craft at 150 kW peak and 120 kW sustained power that matches an Nvidia GB300 rack, reuses Starlink V3 solar technology, links by laser, and runs at only a few milliseconds of latency from low orbit. Chips start as off the shelf Nvidia GB300 and Rubin parts plus a TPU reference design, then scale through a planned 100 million square foot “Terafab” toward a terawatt per year of compute, about twice current US electricity use. The endgame pushes another 1,000x by manufacturing on the moon and using a lunar mass driver to fling satellites into deep space without rockets.

    Thoughts

    The most important reframe in this conversation is that Starlink, Starship, the xAI acquisition, and a new chip factory are not separate bets. They are one bet expressed as a single number: the percentage of the sun’s energy that civilization can capture and put to work. By anchoring everything to the Kardashev scale, Musk turns “build more satellites” into a measurable physics goal rather than a product roadmap. It is a rhetorically powerful move because it makes today’s hyperscale AI buildout, which already strains terrestrial grids, look like the obvious forcing function for going to space. If you accept that compute demand keeps compounding, then the constraint stops being chips and becomes power and cooling, and space genuinely is better at both.

    The cleverest engineering insight is almost understated: an AI satellite is simpler than a Starlink satellite, not harder. A Starlink craft carries complex phased array and parabolic antennas to talk to millions of dispersed users. An orbital data center mostly needs solar cells, radiators, some laser links, and the chips. SpaceX has already industrialized the hard parts (mass produced solar arrays, constellation flight operations at 10,000 satellites, laser mesh networking), so the new product is closer to a remix of proven subsystems than a clean sheet program. That is the real argument for why SpaceX, specifically, can do this when “data center in space” has sounded like science fiction for a decade.

    The numbers are where skepticism should live, and to his credit Musk says to take the timeline with a grain of salt. An annualized gigawatt of space compute by the end of next year, scaling roughly 10x per year toward a terawatt, is an extraordinary ramp. A terawatt is about twice the entire electricity consumption of the United States, delivered as orbiting hardware. Getting there leans on Starship hitting rapid reusability and on a 100 million square foot chip fab that is ten times Gigafactory Texas. Each of those is itself a moonshot, and stacking them multiplies the risk. The honest read is that the architecture is coherent even if the schedule is aspirational.

    The moon segment is where the talk turns from aggressive to genuinely speculative, and it is the part worth watching. A lunar mass driver, essentially a long linear motor that accelerates payloads to escape velocity, only makes sense once you are already moving enormous mass and want to escape Earth’s gravity well and atmosphere entirely. It is a classic Musk pattern: solve the near term problem (mass to orbit with Starship) in a way that creates the precondition for the next, larger problem (local production on the moon). Whether or not the dates hold, the dependency chain is logical, and it explains why SpaceX keeps investing in capabilities that look excessive for today’s market.

    One underrated takeaway for readers outside aerospace: this is as much a manufacturing story as a space story. The bottleneck is not whether a single AI satellite works, it is whether you can stamp out thousands to a million of them, plus the solar, plus the chips, at volume and low cost. That is why so much of the conversation is about Bastrop production lines, a solar manufacturing facility already under construction, and the Terafab. The space hardware is the visible part; the factories are the actual product.

    Key Takeaways

    • The whole strategy is framed around the Kardashev scale, a measure of how much power a civilization harnesses, named for Russian physicist Nikolai Kardashev.
    • Type 1 harnesses a planet’s available power, Type 2 a star’s full output, and Type 3 a galaxy’s; humanity sits at the very bottom of even Type 1.
    • We currently use much less than a trillionth of the sun’s power output, and a trillion is a million times a million.
    • The sun is about 99.86% of all mass in the solar system; most of the remaining 0.14% is Jupiter, and Earth is a tiny dust mote by comparison.
    • Incident solar energy on Earth’s cross section is roughly a half billionth of the sun’s total power output.
    • Most of that sunlight is unusable because about 70% of Earth is water and much of the land is at the poles or far north where solar is weak.
    • Reaching one millionth of the sun’s output, a “micro” on the Kardashev 2 scale, would be an epic achievement relative to today, and 1% would make a civilization vastly more powerful than ours.
    • Space avoids building massive ground power plants and makes cooling easier, because waste heat can radiate directly into the vacuum.
    • Three limiting factors must be solved to scale: mass to orbit, solar power plus radiators, and AI chips.
    • Starship provides the mass to orbit and is the first rocket designed for full and rapid reusability, the breakthrough behind both multiplanetary life and ascending the Kardashev scale.
    • SpaceX catches the booster with the launch tower instead of adding heavy landing legs, an extreme mass optimization measure.
    • Starship V3 already produces more than double the thrust of the Saturn V; V4 will be roughly three times, making it the largest, heaviest, most powerful moving object ever built.
    • Starship is targeted to eventually fly more than once per hour.
    • SpaceX already delivers roughly 85 to 90% of all Earth mass to orbit with Falcon 9 and Falcon Heavy.
    • The plan is to go from around 2,500 tons to orbit per year to millions of tons per year, reaching a million tons per year in about three years.
    • The AI satellite, called AI1, is actually simpler than a Starlink satellite because it lacks the complex phased array and parabolic antennas.
    • AI1 targets 150 kW peak power and 120 kW sustained power, roughly matching an Nvidia GB300 rack of 72 GPUs.
    • Design assumptions are about 250 watts per square meter for the solar array and about 1,400 watts per square meter for the double sided radiators, both expected to improve over time.
    • Radiators are oriented knife edge to the sun and radiate from both sides; each satellite has roughly a 70 meter wingspan.
    • Each satellite carries on the order of a terabit of laser link connectivity.
    • Satellites connect to each other or to the Starlink constellation by laser, and Starlink relays data to the ground over existing Ka and Ku antennas plus laser to ground links.
    • At 600 to 800 km altitude latency is only around 3 milliseconds, since light travels about 300 km per millisecond.
    • SpaceX has about 10,000 Starlinks in orbit and is the only operator with experience flying constellations at that scale.
    • The constellation could eventually grow to thousands or even up to a million satellites; space is big enough to pack and fly them safely.
    • The satellites and solar will be built in Bastrop, Texas, where a solar manufacturing facility is already under construction.
    • The AI satellite production building and solar production are expected to be operating at reasonable volume by the end of next year.
    • SpaceX keeps making Starlink user terminals in Bastrop and is turning on new, higher volume production lines, with possibly a few hundred million terminals eventually, plus a direct to cell constellation that connects straight to phones.
    • Initial chips are off the shelf: the reference design targets Nvidia GB300 or Rubin chips, with a TPU reference design as well, and essentially any existing chip can be put into orbit.
    • The chip industry looks set to reach maybe 100 gigawatts a year of AI compute, far short of the terawatt SpaceX wants.
    • To close that gap, SpaceX plans a “Terafab,” a chip factory around 100 million square feet, roughly 10 times the size of Tesla Gigafactory Texas.
    • A terawatt of chip output per year is like a billion full reticle equivalent chips, each running about a kilowatt, plus a lot of memory.
    • The timeline targets an annualized rate of a gigawatt per year of space compute by the end of next year, scaling roughly 10x per year: 10 GW in about 2.5 years, 100 GW in about 3.5 years, then a terawatt per year, which is 1,000 GW and about twice current US electricity consumption.
    • Beyond a terawatt, the only path to another 1,000x is the moon, using local production of photovoltaics, solar, and radiators so most mass does not have to be shipped from Earth.
    • A lunar mass driver (a linear electric motor or rail gun) could accelerate AI satellites into deep space without rockets, thanks to the moon’s lack of atmosphere and one sixth gravity.
    • Bringing that much mass to the moon would also make it possible for anyone who wants to go to the moon to go, and even live there.
    • Musk stresses none of this requires magic; the AI satellite reuses Starlink V3 solar technology, and he frames the timelines as a best guess rather than a promise.
    • SpaceX has acquired xAI, now referred to as SpaceX AI, folding its AI ambitions directly into the space company.

    Detailed Summary

    The Kardashev Scale and Why Earth Barely Registers

    Musk opens with the question of how you objectively measure a civilization’s progress, the metric an alien species would use to calibrate us. The answer he reaches for is the Kardashev scale, named for the Russian physicist who proposed it, which ranks civilizations by the power they harness: a planet’s worth (Type 1), a star’s worth (Type 2), or a galaxy’s worth (Type 3). Humanity is extremely low even on Type 1. To dramatize the scale of the sun, he notes it is about 99.86% of all the mass in the solar system, with most of the rest being Jupiter and Earth a tiny dust mote in the miscellaneous category. The incident solar energy hitting Earth’s cross section is only about a half billionth of the sun’s total output, and we capture a vanishingly small slice of even that.

    Why Energy at Scale Means Going to Space

    Because roughly 70% of Earth is water and much of the remaining land sits at the poles or in far northern regions where solar is weak and few people live, the usable area for ground solar is small. To reach any meaningful percentage of the sun’s energy, you have to go to space. Musk sets the aspiration at a millionth of the sun’s output as a first “micro” milestone, noting that even 1% would make a civilization vastly more powerful than today’s. Orbit also solves two practical problems at once: you avoid building enormous terrestrial power plants, and cooling becomes easier because waste heat can be radiated straight into the vacuum rather than fought against in an atmosphere.

    The Three Limiting Factors

    Scaling to space based compute comes down to three things: a large mass to orbit capability, a lot of solar power and radiators, and a lot of AI chips. To put a hundred gigawatts and ultimately a terawatt into space, you need a terawatt of solar generation, the radiators to reject the heat, and a terawatt of AI chips. The rest of the conversation works through each limiting factor in turn, starting with the one SpaceX has spent two decades on.

    Starship and the Reusability Breakthrough

    Starship supplies the mass to orbit. Musk argues that full and rapid reusability is the fundamental breakthrough required for both multiplanetary life and climbing the Kardashev scale, since expendable rockets are simply too expensive and you cannot build enough of them. Every other mode of transport, from cars to planes to bicycles, is reusable; rockets are uniquely hard because Earth has a deep gravity well and thick atmosphere, which is why many prior reusable rocket attempts were abandoned. SpaceX pushes mass optimization to the extreme, even catching the booster with the launch tower instead of carrying heavy landing legs. The goal beyond catching the rocket is reflying it with no refurbishment, like an aircraft. Starship V3 already more than doubles the Saturn V’s thrust, V4 will be roughly triple, and the vehicle is the largest and most powerful moving object ever made, targeted to fly more than once per hour. SpaceX already lifts an estimated 85 to 90% of all Earth mass to orbit, and plans to scale from about 2,500 tons per year to millions of tons per year, reaching a million tons per year in roughly three years.

    Inside the AI Satellite (AI1)

    The team explains that a data center in space is not a building with engines bolted on; it reduces to chips plus the power and cooling to run them. The AI satellite, dubbed AI1, is actually simpler than a Starlink satellite because it skips the complex phased array and parabolic antennas, leaving mostly solar cells, a radiator, and some laser links. The draft version targets 150 kW peak power and 120 kW sustained, matching roughly what an Nvidia GB300 rack of 72 GPUs draws. Design assumptions are about 250 watts per square meter of solar array and about 1,400 watts per square meter for double sided radiators oriented knife edge to the sun, both numbers expected to improve. The result is a craft with around a 70 meter wingspan and roughly a terabit of laser connectivity. Compute racks link to each other or to the Starlink constellation by laser, and data reaches the ground via existing Ka and Ku antennas or laser to ground links. From 600 to 800 km up, latency is only about 3 milliseconds, since light travels 300 km per millisecond, so the common worry about high latency does not apply.

    Operating a Constellation of a Million Satellites

    The satellites are large, but space is enormous, so even thousands or up to a million of them would not crowd orbit; viewed against the Earth they are nearly invisible. SpaceX leans on hard won operational experience, with about 10,000 Starlinks already flying and a unique track record of operating constellations at that scale safely. Knowing how tightly satellites can be packed and flown without collisions is treated as the number one constraint when designing the constellation.

    Manufacturing in Bastrop, Texas

    The satellites and solar will be built in Bastrop, Texas, in a facility the hosts describe as already massive and about to be dwarfed by what comes next. A solar manufacturing facility is already under construction, and the AI satellite production building will follow, with both expected to operate at reasonable volume by the end of next year. The same site keeps producing Starlink user terminals and is spinning up new, higher volume lines. Musk projects there could eventually be a few hundred million Starlink terminals, alongside a direct to cell constellation that connects straight from a phone to space for high bandwidth communication.

    Chips, the Terafab, and the Road to a Terawatt

    In the near term, SpaceX simply launches chips that already exist. The current reference design targets Nvidia GB300 or Rubin chips, with a TPU reference design as well, and essentially any existing chip can be flown. The problem is that the chip industry as a whole may only reach about 100 gigawatts a year of AI compute, which does not answer how you get to a terawatt. The answer is a gigantic chip factory, a “Terafab” around 100 million square feet, roughly ten times the size of Tesla Gigafactory Texas, big enough that Musk jokes about needing Starship point to point to cross it. Even with no new fundamental breakthroughs, scaling existing chip technology to a terawatt of output per year is, from a logic die standpoint, like a billion full reticle equivalent chips each running a kilowatt, plus a lot of memory. The stated timeline is an annualized gigawatt per year of space compute by the end of next year, then scaling roughly an order of magnitude per year: about 10 GW in 2.5 years, 100 GW in 3.5 years, and eventually a terawatt per year, which is 1,000 GW, about twice the current electricity consumption of the United States. Musk repeatedly flags these as best guesses, not promises.

    The Moon, a Mass Driver, and the Next 1,000x

    Asked why stop at a terawatt, Musk says a terawatt is actually very small. Getting another three orders of magnitude, a 1,000x jump, points to the moon. The plan is local lunar production of photovoltaics, solar, and radiators, so that most of the mass does not have to be transported from Earth, with chips either shipped up or eventually made on the moon. Because the moon has no atmosphere and only one sixth of Earth’s gravity, you can accelerate AI satellites into deep space without a rocket, using an electromagnetic mass driver, essentially a rail gun or linear electric motor. A side benefit of moving that much mass to the moon is that anyone who wants to go to the moon would be able to, and could even live there. The team closes on the excitement of building a whole new kind of satellite and the sci fi prospect of a mass driver on the moon.

    Notable Quotes

    “We currently use much less than a trillionth of the power output of the sun. And a trillion is a million times a million.”

    Elon Musk, on how far humanity sits from harnessing the sun’s energy

    “The sun is about 99.86% of all mass in the solar system.”

    Elon Musk, dramatizing the scale of the star we orbit

    “You’re an extremely kick-ass civilization if you get to 1% of the sun’s energy.”

    Elon Musk, on what a meaningful Kardashev milestone would look like

    “Reusability is the fundamental breakthrough that is necessary to make life multiplanetary, as well as to ascend the Kardashev scale.”

    Elon Musk, on why Starship matters

    “An AI satellite is essentially a lot of solar cells, a radiator, and you still need some laser links, but you don’t have all of the super complex antennas that you have on a Starlink satellite.”

    Elon Musk, on why the orbital data center is simpler than Starlink

    “There’s not some magic that’s necessary that doesn’t exist for the AI satellites.”

    Elon Musk, on reusing existing Starlink technology

    “We expect that the Terafab is going to be around 100 million square feet, which is 10 times the size of the Tesla Gigafactory Texas.”

    Elon Musk, on the chip factory needed to reach a terawatt

    “The only way that we can really see that you can achieve that is on the moon with a mass driver.”

    Elon Musk, on scaling another 1,000x beyond a terawatt

    Watch the full conversation here: Elon Musk and the SpaceX team on AI satellites and climbing the Kardashev scale.

    Related Reading

    • Kardashev scale (Wikipedia), background on the Type 1, 2, and 3 framework that anchors the entire conversation.
    • Starship (SpaceX), the official page for the fully reusable vehicle behind the mass to orbit numbers.
    • Starlink, the constellation whose solar arrays, laser links, and operations the AI satellites are built on.
    • Mass driver (Wikipedia), the electromagnetic launch concept proposed for flinging satellites off the moon.
    • Nvidia GB300 (Nvidia), the GPU rack whose power profile defines the first AI satellite’s compute target.
  • How GPT-5, Claude, and Gemini Are Actually Trained and Served: The Real Math Behind Frontier AI Infrastructure

    Reiner Pope, CEO of MatX and former TPU architect at Google, sat down with Dwarkesh Patel for a different kind of episode: a chalk-and-blackboard lecture on how frontier LLMs like GPT-5, Claude, and Gemini are actually trained and served. With nothing but a handful of equations and public API prices, Reiner reverse engineers an astonishing amount of what the labs are doing. If you have ever wondered why Fast Mode costs more, why context length stalls around 200k tokens, why models seem 100x over-trained, or why hyperscalers are pouring half a trillion dollars into memory, this is the most lucid explanation on the internet.

    TLDW

    Frontier LLM economics come down to two simple budgets: compute time and memory time. Once you write the rooflines on a blackboard, almost everything else falls out of them. Optimal batch size is roughly 300 times your sparsity ratio (around 2,000 to 3,000 tokens for a DeepSeek-style model). A new batch “train” departs every 20 milliseconds because that is how long it takes to read HBM end to end. Mixture of experts strongly favors staying inside a single rack, which is why scale-up domains went from 8 GPUs (Hopper) to 72 (Blackwell) to 500-plus (Rubin). Pipeline parallelism solves weight capacity but does nothing for KV cache, and adds painful per-hop latency, which is why Ilya famously said pipelining is not wise. Because of reinforcement learning and inference economics, frontier models are roughly 100x over-trained versus Chinchilla optimal, and a well-tuned model should output roughly as many tokens during deployment as went into its pre-training corpus. API prices leak the rest: Gemini’s 50% premium above 200k tokens reveals where KV memory time crosses weight memory time, prefill being 5x cheaper than decode confirms decode is memory bandwidth bound, and cache hit pricing tiers map directly to HBM, DDR, flash, and (yes) spinning disk. The lecture closes on a beautiful detour about the convergent evolution of neural nets and cryptographic ciphers.

    Key Takeaways

    • Two equations explain almost everything. A roofline analysis comparing compute time to memory fetch time predicts cost, latency, and architectural choices with shocking accuracy.
    • Optimal batch size is about 300 times sparsity. For a DeepSeek model that activates 32 of 256 experts, that lands around 2,000 to 3,000 tokens per batch. Real deployments go a bit higher to leave headroom.
    • The 20 millisecond train. A new batch departs every 20ms because that is how long it takes to read all of HBM once. Worst-case queue latency is roughly 40ms.
    • Fast Mode is just smaller batches. Pay 6x more, get 2.5x faster decode by amortizing weights over fewer users. There is a hard latency floor at the HBM read time.
    • Slow Mode would not save much. Once you are past the optimal batch size, the cost-per-token plateau is dominated by compute, not weight fetches. You cannot meaningfully amortize KV cache because it is unique per sequence.
    • One rack is the natural MoE unit. Expert parallelism wants all-to-all communication, which strongly favors the scale-up network (NVLink) over the scale-out network (roughly 8x slower).
    • Bigger scale-up domains drove model scaling. The jump from 8 (Hopper) to 72 (Blackwell) to 500-plus (Rubin) GPUs per rack increased aggregate memory bandwidth by 8x, which is why trillion-plus parameter models only became viable recently.
    • Pipeline parallelism is overrated for inference. It saves on weight memory capacity but does nothing for KV cache memory. It also adds milliseconds of latency per hop in decode.
    • Why Ilya said pipelining is not wise. Architectural constraints (cross-layer residuals like in Kimi) and the inability to amortize weight loads across micro-batches make pipelining a hassle in training too.
    • The memory wall is real and paradoxical. Hyperscalers reportedly spend 50% of CapEx on memory, yet racks have far more HBM than a trillion-parameter model needs. The capacity is there for KV cache and batch size, not for weights.
    • Frontier models are roughly 100x over-trained vs Chinchilla. When you minimize total cost across pre-training plus RL plus inference, smaller models trained on more data win.
    • Each model should output roughly all human knowledge. If you equalize pre-training and inference compute, the total tokens served by a model during its lifetime should approximate its training corpus. Roughly 150 trillion in, 150 trillion out.
    • API pricing reveals architecture. Gemini’s 50% premium above 200k context, the 5x decode-vs-prefill ratio, and cache duration tiers all leak detailed information about KV size, memory bottlenecks, and storage hierarchy.
    • KV cache is roughly 2KB per token. Solving Gemini’s pricing equation gives a plausible 1.6 to 2 kilobytes per token at 100B active parameters and 200k context.
    • Decode is memory bandwidth bound, prefill is compute bound. The 5x price gap is direct evidence.
    • Cache pricing maps to memory tiers. The 5-minute and 1-hour cache durations probably correspond to flash and spinning disk drain times respectively. LLM serving uses spinning disk.
    • Context length is stuck near 200k. Memory bandwidth, not compute, is the binding constraint. Sparse attention gives a square-root improvement but is not infinite.
    • Cryptography and neural nets are mathematical cousins. Both rely on jumbling information across inputs. Feistel ciphers led directly to RevNets (reversible neural networks). Adversarial attacks mirror the cipher avalanche property.

    Detailed Summary

    The Roofline: Compute Time vs Memory Time

    Reiner starts with the simplest possible model of LLM inference. The time to do a forward pass is bounded below by the maximum of compute time and memory fetch time. Compute time is the batch size times active parameters divided by FLOPs. Memory time is total parameters divided by memory bandwidth, plus a KV cache term that scales with batch size and context length. From these two equations, almost every economic and architectural fact about modern LLMs can be derived.

    Plotting cost per token against batch size gives a clean picture: at low batch you pay enormous overhead because you cannot amortize the weight fetches, and at high batch you hit a compute floor. There is a sweet spot where memory bandwidth time equals compute time. That sweet spot is what Fast Mode and Slow Mode are tuning around.

    Why Fast Mode Costs More: The Batch Trade-Off

    When Claude Code or Codex offers Fast Mode at 6x the price for 2.5x the speed, what is really happening is that they are running you at a smaller batch size. Smaller batch means weight loads are amortized over fewer users, so cost per token goes up. But latency goes down because each forward pass touches less data. There is a hard floor on latency because you have to read every byte of HBM at least once per token, and that takes about 20 milliseconds on Blackwell-class hardware. There is also a soft ceiling on Slow Mode savings because the unamortizable parts (KV cache fetches, compute) eventually dominate.

    The 20 Millisecond Train

    HBM capacity divided by HBM bandwidth lands consistently around 20 milliseconds across generations of Nvidia hardware. That is the natural cadence at which a frontier model can run a forward pass over all its weights. Reiner uses a memorable analogy: a train departs every 20 milliseconds. Any users whose requests are ready board the train. If the train is full, they wait. If it is empty, it leaves anyway. This is why you do not need millions of concurrent users to saturate a model’s batch. You only need enough to fill a 2,000-token train every 20ms.

    Why Optimal Batch Size Is About 300 Times Sparsity

    Setting compute time equal to weight fetch time and rearranging gives a beautiful result: batch size needs to be greater than (FLOPs / memory bandwidth) times (total params / active params). The hardware ratio is a dimensionless 300 on most GPUs and has stayed remarkably stable from A100 through Hopper, Blackwell, and Rubin. The model term is just the sparsity ratio. For DeepSeek with 32 of 256 experts active, that is 8. So optimal batch is around 2,400 tokens. Real deployments push this to 3x to leave headroom for non-ideal efficiency. At 64 trains per second, that is roughly 128,000 tokens per second per replica, or about 1/1000 of Gemini’s reported global throughput.

    Mixture of Experts Wants to Live Inside a Rack

    MoE all-to-all routing means every token can be sent to any expert on any GPU. The communication pattern strongly prefers the fast scale-up network (NVLink) inside a rack to the slower scale-out network between racks. Scale-out is roughly 8x slower in bandwidth. This is why one rack ends up being the natural unit for an expert layer, and why Nvidia’s progression from 8 GPUs per rack (Hopper) to 72 (Blackwell) to 500-plus (Rubin) has been such a big deal for model size scaling.

    Reiner walks through the physical constraints: cable density, bend radius, weight, power, cooling. Modern racks are pushing every dimension to the limit. Stuffing more GPUs into the scale-up domain is genuinely a hardware engineering problem.

    Pipeline Parallelism: Why Ilya Said It Is Not Wise

    Pipelining splits model layers across racks. It is the natural way to scale beyond the scale-up domain for very large models. But it has problems. In inference, pipelining does not save runtime, it only saves memory capacity per rack, which already is not the binding constraint because trillion-parameter models only need a terabyte and racks have 10x that. In training, pipelining creates the famous bubble (idle GPU time at the start and end of each pipeline pass) and forces micro-batching, which kills your ability to amortize weight loads across the global batch.

    There is also an architectural cost. Models like Kimi use cross-layer residual connections where attention attends to layers a few back, and pipelining makes those patterns very hard to implement cleanly. Ilya’s quip “as we now know, pipelining is not wise” captures all of this.

    The Memory Wall Paradox

    Industry analysts report that hyperscalers are spending 50% of CapEx on memory this year, while smartphones and laptops are seeing 30% volume drops because there is not enough HBM and DDR to go around. Yet a Blackwell rack already has tens of terabytes of HBM, far more than a trillion-parameter model needs. The reason is that all that extra capacity goes to KV cache, batch size, and longer context. The bandwidth, not the capacity, is what matters most for weight loading. This also implies that hardware could be designed with less HBM per GPU if you commit to pipelining the weights, which is a real architectural option for a chip startup like MatX.

    Reinforcement Learning and the 100x Over-Training of Frontier Models

    Chinchilla scaling laws say a model with N active parameters should be trained on roughly 20N tokens for compute-optimal training. But frontier labs do not just minimize training cost. They minimize training plus inference cost across the model’s deployment lifetime. With reinforcement learning added to the mix, the cost equation has three terms: pre-training (6 times active params times tokens), RL (somewhere between 2x and 6x times active params times RL tokens, with a 30% efficiency penalty for decode-heavy rollouts), and inference (2 times active params times inference tokens).

    If you assume those three roughly equalize at the optimum (a heuristic that holds for many cost curves), you get a clean conclusion: the data going into pre-training should be roughly equal to the data going into RL, which should be roughly equal to the tokens served at inference. With 100 billion active parameters and roughly 150 trillion training tokens, that is about 75x past Chinchilla optimal. Reiner rounds it to 100x. This is the most concrete first-principles argument for why frontier models are so deeply over-trained, and it implies that as inference traffic grows, models should keep getting smaller and longer-trained.

    Each Model Should Output All of Human Knowledge

    The most jaw-dropping consequence: if you equalize pre-training and inference compute, then the total tokens generated by a model across its deployment lifetime should approximate the size of its training corpus. GPT-5, served to hundreds of millions of users for two months, will collectively output something on the order of 150 trillion tokens. That is roughly the sum of human knowledge in textual form. Each frontier model is, in this sense, a one-shot universal author of a corpus the size of its source material.

    API Prices Leak Architecture

    This is where the lecture gets really fun. Gemini 3.1 charges 50% more for context above 200k tokens. Setting memory time equal to compute time at exactly 200k context and solving for KV cache size gives roughly 1.6 to 2 kilobytes per token, which is plausible for a model with 8 KV heads, dense attention, and head dimension of 128.

    The 5x premium for output (decode) tokens versus input (prefill) tokens is direct evidence that decode is severely memory bandwidth bound and prefill is compute bound. Prefill processes many tokens per weight load, so it amortizes memory cost over the whole sequence. Decode processes one token per weight load, so it pays full memory cost every time.

    Cache hits priced at one tenth of cache misses tell you that storing the KV cache in HBM (or DDR or flash) is much cheaper than recomputing it from scratch. The two cache duration tiers (5 minutes and 1 hour) probably correspond to memory tiers whose drain times match those durations: flash for the 5-minute tier, spinning disk for the 1-hour tier. Yes, spinning disk is in the modern LLM serving stack, despite being decades-old technology.

    Why Context Length Has Plateaued at 200k

    Context lengths shot up from 8k to roughly 200k during the GPT-3 to GPT-4 era and have stayed roughly flat for the past two years. Reiner argues this is the natural balance point where memory bandwidth cost crosses compute cost. Going to a million tokens is expensive. Going to 100 million tokens (which Dario has hinted is needed for true continual learning via in-context learning) is essentially impossible without either a memory technology breakthrough or a much more aggressive sparse attention scheme. Sparse attention helps with a square-root improvement, but it is not unlimited. Going too sparse trades off too much quality.

    Cryptography Meets Neural Nets

    The episode ends with a lovely intellectual detour. Cryptographic protocols and transformer architectures both rely on jumbling information across all inputs. They are doing inverse versions of the same operation: ciphers take structured input and produce randomness, while neural nets take noisy input and extract structure. Both fields use differentiation as their primary attack vector (differential cryptanalysis on ciphers, gradient descent on neural nets). Adversarial attacks on image classifiers exploit exactly the avalanche property that good ciphers are designed for.

    The most concrete crossover: Feistel ciphers, which let you build invertible functions out of non-invertible ones, were ported into deep learning as RevNets (reversible networks) in 2017. RevNets let you run the entire network backwards during the backward pass, eliminating the need to store activations and dramatically reducing training memory footprint. It is the opposite trade-off of KV caching: spending compute to save memory rather than spending memory to save compute.

    Thoughts

    The most striking thing about this episode is how much can be deduced from a few equations and the public API price sheets of the major labs. The labs treat their architectures as trade secrets, but the moment they price tokens to be close to cost (which competition forces them to do), the prices themselves leak the underlying ratios. Anyone with a pen and paper can reverse engineer the KV cache size, the memory tier hierarchy, and the compute-vs-memory bottleneck profile of a frontier model. There is a lesson here for builders: in competitive markets, the prices tell you almost everything.

    The 100x over-training result has interesting implications for what comes next. If the optimal balance shifts further toward inference (as adoption keeps growing), models should get smaller and longer-trained. That is good news for serving costs and bad news for training-compute-as-moat. The biggest determinant of model quality might increasingly be data quality and RL environment design, not raw pre-training compute. This squares with what is visible publicly: the leading labs are investing heavily in RL infrastructure, evaluations, and synthetic data pipelines.

    The memory wall is the most underrated infrastructure story in AI. Most people think of compute as the bottleneck, but Reiner makes it clear that memory bandwidth is what actually limits context length, which limits how agentic a model can be in practice. If you cannot get to 100 million token contexts, you probably cannot have an AI agent that has been working with you for a month and remembers everything. Either some sparse attention scheme has to give us cheap effective context length, or we need a memory hardware breakthrough, or we have to invent some form of continual learning that does not rely on context windows. None of those paths are obviously easy, and the fact that context length has been flat for two years despite enormous investment suggests we are stuck against a real wall.

    The cryptography parallel is the kind of cross-disciplinary insight that does not show up enough in AI discourse. Treating neural networks as a kind of differentiable cipher reframes a lot of the architecture choices (residual connections, layer normalization, attention) as deliberate efforts to make the function smooth and invertible enough to learn, in contrast to ciphers, which are deliberately designed to resist exactly that. Adversarial robustness research probably has a lot more to learn from cryptanalysis than it currently does.

    Finally, the format itself is a win. Most AI podcasts are conversational, which is great for personality but bad for technical depth. A blackboard lecture with an interlocutor who asks naive questions at the right moments is a much higher bandwidth medium. More of this, please.