PJFP.com

Pursuit of Joy, Fulfillment, and Purpose

Tag: attention layer

  • Why the Markets Are Pricing AI Wrong: Gavin Baker on the July 2026 Selloff, GPU Spot Prices, Memory LTAs, and Nvidia’s Credit Wrapper

    Gavin Baker of Atreides Management returned to Invest Like the Best with Patrick O’Shaughnessy days after one of the strangest months the AI trade has ever produced. AI and semiconductor names fell 40 to 60 percent in a straight line while, by Baker’s account, not a single quantitative metric on the ground deteriorated. He spent the week in Silicon Valley hunting for a bearish data point and came back with almost nothing except credit. This conversation is the result: a detailed argument that the market has mispriced the gap between contracted compute and spot compute, that open source is growing the infrastructure pie rather than shrinking it, and that the one risk actually worth fearing is political rather than financial.

    TLDW

    Gavin Baker describes July 2026 as “2022 packed into a single month,” a violent AI and semiconductor drawdown that happened while hyperscaler operating cash flow accelerated from roughly 28 percent growth to 32 percent, or closer to 35 percent adjusting for unusual legal charges. His core claim is that the installed base of GPU compute is locked into long-term contracts priced far below the current spot market, so as those contracts roll off, compute reprices higher, operating cash flow accelerates, and the buildout can be funded internally rather than with the debt that widening credit default swap spreads and a poorly received Meta bond have made look expensive. He walks through each catalyst of the selloff: Meta renting out compute (misread as a capex cut), the open source capability leap from GLM 5.2 and Kimi K3 (misread as deflationary when a token is a token and costs the same flops, watts, and memory to produce), China acquiring a domestic deep ultraviolet lithography machine (real but 25 years behind), and rising real yields (the only genuine negative). He covers the game theory of breaking a memory long-term agreement in a world where market share is set by supply allocations, Nvidia’s new credit wrapper plus revenue share model and why it is misunderstood, the router and fine-tuning stack from Fireworks and Baseten that turns “ChatGPT wrappers” into defensible AI natives, continual learning as the one technical development that could disrupt training demand, SRAM accelerators for disaggregated inference, SpaceX as an underappreciated compute company with orbital ambitions, and his view that regulation, not fundamentals, is the biggest risk because the industry has done a terrible job telling its own story. He also makes an unusual observation about market structure: everyone now feeds news into Claude, and Claude has become a kind of Walter Cronkite for the stock market, collapsing the diversity of interpretation that normally keeps markets stable.

    Thoughts

    The load-bearing claim in this episode is the spread between contracted and spot compute, and to Baker’s credit it is falsifiable in a way most bull cases are not. He is not arguing that AI will be transformative or that demand feels strong. He is arguing something narrow and checkable: hyperscalers and neoclouds signed multi-year GPU contracts in 2024 and 2025 at prices that assumed a gentle decline, prices instead went vertical, and the installed base is therefore systematically under-earning. A startup rented several thousand B200s in the mid two dollars per GPU hour range and expects to pay just under four dollars for an identical cluster seven months later. If that repricing is real and broad, hyperscaler operating cash flow mechanically accelerates and roughly 700 billion dollars of projected credit demand evaporates. If GPU rental prices roll over and stay down for two consecutive quarters, the thesis is dead. That is the number to watch rather than any earnings headline. The caveat he steps past quickly is that the open source mix shift he describes as bullish does not eliminate margin, it relocates it, out of the frontier labs and down into the infrastructure layer. Excellent if you sell GPUs, power, and memory. Considerably more awkward for the labs whose projected cash flows are the reason anyone believes the compute gets paid for at all.

    The Claude as Walter Cronkite observation deserves more attention than it got, where it passed as a joke. Baker is describing a genuine change in market microstructure. Every institutional and retail participant now feeds the same news into roughly the same models, and while those models are probabilistic, they are not producing meaningfully diverse readings of the same headline. He connects this to Michael Mauboussin’s argument that a breakdown in diversity, not leverage alone, is what produces bubbles and crashes. If that is what happened in July, then the Japanese capacitor stock chart he cites, an entire three-year cycle compressed into six weeks before the fundamentals had even arrived, is not a curiosity. It is the signature of a market where thousands of participants share one interpretive engine. That makes drawdowns faster and deeper without making them more informative, which argues for holding through machine-generated narrative cascades rather than trading them.

    The middle of the conversation contains the most consequential business idea in it, and it is one that got almost no coverage during the selloff: memory long-term agreements and Nvidia’s credit wrapper are the same move executed at two different layers of the stack. Both trade near-term upside for durability. The memory companies stopped maximizing spot price and started signing prepaid agreements with floors and ceilings, and the reason those agreements will hold is that the penalty for breaking one has changed category. Apple could renege on memory pricing for years because its volume was overwhelming and it had no equivalent competitor. In a world with four buyers that matter and where AI market share is set by supply allocation rather than product quality, a supplier can answer a broken price agreement by breaking the volume commitment and handing your allocation to a rival, in an industry where oversupply is always followed by undersupply. Nvidia is running the same play one layer up. The credit wrapper with a revenue share above a price floor converts a cyclical one-time chip sale into a royalty on recurring compute revenue, financed on someone else’s balance sheet, which is a materially better business than selling hardware. It also widens the moat, because a startup accelerator pays more at the foundry, pays more for high bandwidth memory, and cannot finance its chips at Nvidia’s rate. Baker is right that this is misunderstood, and it is a strange thing for a stock at a ten-year-low forward multiple to be quietly doing.

    The technical material in the back half reveals an asymmetry worth naming. Baker treats two efficiency developments very differently. Continual learning and sample efficient learning, which several labs believe are close, would collapse the token budget required to produce a capable model, and he handles this by asserting that training asymptotes to a small but nonzero share of compute and that the outcome would be wonderful for the world anyway. SRAM-based accelerators for disaggregated inference, running prefill on one chip, attention on a high-memory chip, and the feed forward network on SRAM, he embraces enthusiastically as a return-on-investment improvement across the installed base. Both are efficiency gains. One is treated as neutral, the other as clearly positive, and Jevons paradox is doing all the work in both directions. That is probably correct given everything we have observed so far, but it is an assumption rather than a finding, and it is the assumption on which the entire “cheaper compute is bullish for compute” framework rests. Worth noting too that the SRAM disaggregation point is genuinely underdiscussed: those chips sit on older nodes and do not compete for leading-edge capacity, so they are additive supply rather than substitute supply.

    The final twenty minutes hold both the largest unpriced upside and the largest unpriced risk, and neither is in consensus estimates. On the upside, only the hyperscalers, CoreWeave, Crusoe, and SpaceX have ever brought more than 500 megawatts online in a single year, and SpaceX has done it fastest and cheapest. When it dumped a large block of compute into the market, the market absorbed it without a blip, which tells you more about demand than any survey. Baker’s sanity check on orbital compute is the sharpest reasoning move in the episode: Benchmark, from entirely outside the Elon ecosystem and without the benefit of internal launch costs, funded StarCloud at a real valuation, so the set of people who would all have to be wrong keeps growing. On the downside, regulation is the risk he names first and it is the one his own framework cannot arbitrage. New York’s data center moratorium is not a fundamentals problem, and no amount of operating cash flow acceleration fixes a permitting ban. His diagnosis is that the industry finds the benefits so obvious that it never learned to explain them, which is how a water usage figure overstated by four orders of magnitude became conventional wisdom. Proposing a foundation that buys World Series ad time is a tell about how far behind he thinks the industry is. Every other risk in this conversation is priced somewhere. That one is not.

    Key Takeaways

    • Baker characterizes July 2026 as “2022 in a month,” with AI names down 40 to 60 percent from their highs in a straight line while underlying fundamentals improved.
    • He spent the week in Silicon Valley explicitly hunting for a negative quantitative metric and found essentially one: third-party data suggesting Anthropic’s growth curve came slightly off trajectory, a data point Anthropic shareholders reportedly dispute.
    • Nvidia was trading at its lowest forward price to earnings multiple in ten years at the time of recording. The only cheaper moments were the DeepSeek shock and Liberation Day, both of which proved to be V-bottoms.
    • A low forward multiple means the market believes these companies are significantly over-earning. Baker’s counter is that they are under-earning because their installed compute is contracted below spot.
    • Combined operating cash flow at Microsoft, Meta, and Amazon accelerated from roughly 28 percent to 32 percent growth, or to about 35 percent after adjusting for an unusual quarter of legal and regulatory charges.
    • Nobody in 2024 or 2025 modeled old GPU prices going vertical in 2026. The bull case assumed a slow decline in rental rates and the bear case assumed a steep one.
    • A concrete example: a well-known startup rented several thousand Blackwell B200s in the mid two dollars per GPU hour range and expects to pay just under four dollars for an identical cluster seven months later, a 50 to 60 percent increase.
    • One inference cloud stated publicly that it plans to pay roughly 100 percent more for Blackwells when its current contract expires.
    • Neoclouds were often forced into below-market long-term contracts because they needed an offtake agreement to finance the GPUs in the first place.
    • Consensus models hyperscalers monetizing Blackwell and Rubin at roughly Ampere rates, two generations behind, producing about 1.3 to 1.4 trillion dollars of hyperscale operating cash flow. Assuming monetization merely at a discount to current Blackwell rates pushes that closer to two trillion and removes roughly 700 billion dollars of credit demand.
    • The credit concerns are real and undeniable: real yields are up, spreads have widened, credit default swap levels for the large buyers have blown out, and a recent Meta bond did not price where a Meta bond should price.
    • Baker’s response is that debt-fueled buildouts demand immediate repayment and unwind violently, which is what happened in the internet buildout, but this buildout is still overwhelmingly funded from operating cash flow.
    • If credit is not available, he argues the existing flops simply become more valuable, which is self-correcting rather than catastrophic.
    • The Meta selloff catalyst was a misread. Meta renting out compute was interpreted as excess capacity and a capex cut. Meta did not cut capex, and the actual motivation appears to have been demonstrating strong internal rates of return on a small slice of capacity ahead of a capital raise.
    • The open source panic was also a misread. Open source taking token share moves margin dollars out of the frontier model layer, but a token still requires the same flops, memory, and watts to produce, so infrastructure demand rises rather than falls.
    • Frontier tokens carry gross margins somewhere in the 80 to 95 percent range. Open source tokens might carry 30 percent. The customer’s savings come almost entirely out of that margin, not out of compute consumption.
    • Baker calls open source “dark matter to the public markets,” growing rapidly through GLM 5.2, Kimi K3, and Nvidia’s Nemotron, but nearly impossible for public investors to measure since it runs through private inference clouds.
    • Jensen Huang being the world’s loudest supporter of open source is itself evidence that open source is good for Nvidia’s business.
    • Enterprises that blow through their AI budget in three months set up a router, which cuts their spend but often increases total GPU hours consumed by shifting volume to cheaper open source tokens.
    • Adoption is happening in staggered waves: AI natives are all in and hiring very few humans, coastal public companies are optimizing, East Coast and non-coastal companies have barely adopted, and Europe is trying to regulate AI before using it.
    • Roughly 500,000 people worldwide use agentic AI, and perhaps half that number use it seriously, yet the world is already in an acute compute shortage. The relevant question is what happens at 100 million or 500 million users.
    • Token spend at the most AI-forward companies now runs 20 to 25 percent of total compensation spend, with individual examples at 30 percent and reports as high as 50 percent, against a roughly 25 trillion dollar global knowledge work market.
    • Founder-controlled companies are not conducting large-scale layoffs, which suggests the cash flow to pay for AI is expected to come from growth rather than from labor substitution.
    • Memory is the dominant variable in token economics. More memory per unit of compute yields more tokens out, which lowers cost per token, which is why demand has shown no negative elasticity to memory pricing.
    • Memory suppliers have shifted from maximizing near-term price to signing long-term agreements with prepayments, floors, and ceilings, trading short-term upside for durability.
    • Breaking a memory long-term agreement is now potentially fatal. With four buyers that matter at scale and market share determined by supply allocation, a supplier can respond by breaking the volume commitment and handing your allocation to a competitor.
    • This is structurally different from the Apple era, when a single dominant buyer could break pricing agreements without consequence.
    • Nvidia’s new model is best described as a credit wrapper with a revenue share triggered when GPU prices exceed a floor. It is not vendor financing, since a third party lends the money, and it could produce a very large cloud-scale royalty business quickly.
    • Baker thinks this model is badly misunderstood, meaningfully increases Nvidia’s revenue per gigawatt, and strengthens its competitive position against startup accelerators that pay more at the foundry, pay more for high bandwidth memory, and cannot finance their chips as cheaply.
    • Nvidia has taken equity stakes across the ecosystem, and Baker’s read is that every time they have not taken a stake it has proven to be a mistake.
    • The scenario that would genuinely frighten him: hyperscaler operating cash flow stops accelerating, forcing the buildout onto debt, or a sustained sharp contraction in GPU rental prices. Nobody he has spoken to says they have too many GPUs.
    • Continual learning and sample efficient learning are the technical developments most likely to disrupt training demand, and several new labs including Safe Superintelligence are focused on them. Baker still thinks training asymptotes to a small share of compute rather than to zero, and that the change would be enormously good for the world regardless.
    • Fireworks launched a product called Nexus that plugs into Claude Code, OpenAI Codex, or Grok in roughly three lines of code, ingests a customer’s data, applies reinforcement learning to a model, and routes queries appropriately.
    • This stack is what converts an alleged “ChatGPT wrapper” into a defensible company. Shifting 30 to 60 percent of token consumption to a customized open model on top of frontier orchestration produces better outcomes at roughly half the cost.
    • Cheap, capable open source models may actually inflate the value of the very best frontier model, since a 160 IQ orchestrator becomes more valuable when it has an army of cheap 120 IQ models to direct.
    • The inference clouds are growing almost as fast as the frontier labs did in their early days while burning very little cash, which is extraordinary by any conventional software metric.
    • China obtaining a domestic deep ultraviolet lithography machine is a genuine phase transition and should not be dismissed, but the technology is roughly 25 years behind extreme ultraviolet, and lithography progress is learning by doing that cannot be teleported through.
    • Baker considers regulation the biggest single risk to AI, citing New York’s data center moratorium as the first of many and describing the current environment as post-factual and post-logical.
    • The public narrative that data centers raise power bills, drain water, and destroy jobs is largely wrong. Behind the meter deals typically lower local electricity prices, and modern community agreements include hospitals, schools, police and fire stations.
    • The widely cited data center water figure originated in a published error overstating usage by roughly 10,000 times, since acknowledged by the author, which Baker likens to the decimal point error that created the myth that spinach is exceptionally high in iron.
    • He argues data centers are among the best things to happen to blue collar wages in his lifetime, with ongoing rather than one-time employment from maintenance, replacement, and upgrade cycles.
    • SRAM-based accelerators built on older nodes and free of high bandwidth memory constraints could substantially improve return on investment by allowing disaggregated inference: prefill on one chip, attention on a high-memory chip, and the feed forward network on SRAM.
    • SpaceX has improved fundamentally since going public, and Baker believes the market does not yet understand it as a compute company. Only the hyperscalers, CoreWeave, Crusoe, and SpaceX have ever brought on more than 500 megawatts of power in a single year, and SpaceX has done it fastest and cheapest.
    • A widely circulated report claims SpaceX intends to bring on eight gigawatts of compute in 18 months. Baker doubts the number but notes that at roughly 50 billion dollars of monetization per gigawatt, even a fraction of it dwarfs the current consensus estimate.
    • When SpaceX dumped a large block of compute into the market, it was absorbed without a blip, which Baker reads as one of the more bullish demand signals of the year.
    • Orbital compute feels more real every day. Benchmark funding StarCloud, from outside the Elon ecosystem and without access to internal launch costs, functions as a useful sanity check on the idea.
    • Dark horse names Baker flags for the next phase: Lip-Bu Tan, Lin Qiao at Fireworks, and Scott Wu at Cognition.

    Detailed Summary

    A Selloff That Contradicted Every Fundamental

    Baker opens by describing July 2026 as 2022 compressed into a single month. AI names fell 40 to 60 percent from their highs in a nearly straight line. What made the month unusual was not the magnitude but the absence of a legible cause. In 2022 the market feared recession, rising rates, and inflation. During the DeepSeek shock and Liberation Day you knew exactly what the market was reacting to. This time the fundamentals moved in the opposite direction from the tape. GPU availability tightened, GPU rental pricing rose, DRAM spot prices rose, and token growth accelerated. Baker asked Patrick, who had also spent the summer in Silicon Valley, whether he had heard a single negative quantitative metric or a single instance of deceleration. The answer was nothing.

    Part of the problem is visibility. Public markets cannot see Anthropic or OpenAI directly, and they cannot see the American open source inference clouds like Fireworks, Baseten, Modal, and Together that monetize inference. Everyone stares at the same chart of semiconductor cash flow rising while hyperscaler free cash flow falls, and that chart omits the private companies entirely. It also omits the repricing dynamic Baker considers the most important fact in the market.

    The Spot Versus Contract Gap

    In 2024 and 2025 every serious forecast assumed GPU rental prices would decline, with the only debate being how fast. Neoclouds locked in long-term contracts partly out of prudence and partly because they needed offtake agreements to finance the hardware at all. The result is a large installed base of contracted compute trading at a steep discount to today’s spot market. Baker’s argument is that as those contracts roll off, compute reprices higher even if spot itself declines from current levels, and that repricing flows directly into hyperscaler operating cash flow.

    The anecdotes are stark. A prominent startup rented several thousand B200s in the mid two dollar per GPU hour range and expects to pay just under four dollars for an identical cluster seven months later. One inference cloud said publicly it plans to pay roughly double for Blackwells at contract renewal. Baker’s read is that hyperscalers are therefore under-earning across the board, which is the exact opposite of what a ten-year-low forward multiple implies the market believes.

    Financing the Buildout and the Credit Question

    Credit is the one bearish input Baker concedes is real. Real yields have risen, spreads have widened, credit default swap levels have blown out across the large buyers, and a recent Meta bond did not price the way a Meta bond should. Sophisticated private capital investors told him this is just banks hedging commitments, but he acknowledges the optics are bad and the facts are undeniable. His concern is the classic capital cycle: debt-financed buildouts demand immediate repayment, so when supply and demand slip out of alignment the unwind is fast and brutal, exactly as it was in the internet buildout.

    The math he ran is the counterweight. Consensus effectively models hyperscalers monetizing Blackwell and Rubin at Ampere rates, two generations behind, producing 1.3 to 1.4 trillion dollars of operating cash flow. Assume instead that they monetize merely at a modest discount to current Blackwell rates and the figure approaches two trillion, taking about 700 billion dollars of credit demand off the table. Better cash flow also improves the credit ratios, which makes debt cheaper if they choose to use it. And if credit disappears entirely, the flops already installed simply become more valuable. Microsoft brought on a large slug of capacity in June that did not even appear in second quarter results.

    How the Month Actually Unfolded

    Baker walks the sequence of catalysts. First, Meta announced it would rent out compute, which the market read as excess capacity and an imminent capex cut. Meta did not cut capex. What Meta appears to have seen was SpaceX selling trading-optimized clusters into the market at an enormous premium to contracted rates, and the plan was likely to demonstrate strong returns on a small slice of capacity before raising equity capital and increasing capex. Shortly afterward Meta released its best model in a long time, overshadowed by a competing release but a clear signal it was not easing off.

    Next came the open source freakout. Kimi K3 arrived, the widely watched token index dipped and flattened, and the two were connected: the index captures mix, and a shift from expensive frontier tokens toward open source tokens looks like weakness even when total compute consumption is rising. Then China’s deep ultraviolet lithography news triggered a broad selloff in semicap equipment. Finally, rising real yields and widening spreads gave the market a genuine reason to worry. Baker’s summary is that with the sole exception of credit, every one of these narratives was factually wrong, and a friend at Fidelity described the winning strategy of the past three years as doing the dumbest, most superficial thing as fast as possible and cycling between them.

    Open Source as Dark Matter

    The most important conceptual argument in the episode is that a token is a token. Regardless of which model produces it, a token consumes the same flops, the same memory, and the same watts. Open source taking share therefore does not reduce compute demand. It transfers margin from the frontier model layer, where gross margins might be 90 percent, to open weights inference at perhaps 30 percent, and the resulting price decline drives elasticity in token volume. Since frontier labs and open source models both run on the same underlying cloud infrastructure at the same compute cost, the effect is to push margin dollars down into the infrastructure layer.

    Baker calls open source dark matter to public markets. It is real, it is accelerating on the back of capability leaps from GLM 5.2 and Kimi K3, Nvidia continues to push Nemotron closer to the frontier, and yet none of it appears in audited financials that public investors can underwrite. He also notes the tell that should have settled the debate: Jensen Huang is the world’s most vocal supporter of open source, which would be an odd position for the largest beneficiary of frontier concentration to hold if open source actually threatened the business. Baker adds a normative point, that a world with only one or two dominant frontier models charging 90 percent margins is not good for humanity, and that many models is the better outcome.

    Routers, Fine-Tuning, and the End of the Wrapper Insult

    The practical mechanism behind the open source surge is the router plus fine-tuning stack. Inference clouds have become genuinely good at supervised fine-tuning and reinforcement learning, so a company can take its proprietary data, customize an open weights model, put it behind a router, and have the router send most queries to that model while escalating to a frontier model for verification or harder work. The result is often slightly better outcomes at half the cost. Fireworks shipped a product called Nexus that connects to Claude Code, OpenAI Codex, or Grok in roughly three lines of code and handles ingestion, reinforcement learning, and routing.

    This changes the durability question for AI natives. Two years ago the criticism was that these companies were thin wrappers with no defensibility. Now a company with domain-specific proprietary data can train on it, own the model serving 30 to 60 percent of its tokens, and get off the frontier lab treadmill it previously had no choice but to accept. Baker points to Cursor, Harvey, and others leaning hard into this. He also raises the counterargument fairly: some believe that once a frontier model achieves recursive self-improvement it will serve every intelligence level more cheaply through distillation, leaving no room for open source. He does not dismiss it, but he thinks the proprietary data held by AI natives and the orchestration value of the single smartest model make the multi-model future more likely. Cheap 120 IQ models arguably make a 160 IQ orchestrator more valuable, not less.

    Where the Money Comes From

    The pushback Baker gets on X is fair: even if hyperscalers are under-earning, where does the customer revenue ultimately come from? Definitionally it must come from faster economic growth through productivity or from labor substitution. He sees labor substitution happening at AI natives, though not through firing. They simply never hire the humans, and gross profit dollars per full-time employee at these companies is vertical compared with prior startup generations. Token spend now runs 20 to 25 percent of total compensation spend at the most aggressive companies, with individual examples at 30 percent and reports as high as 50 percent, against a roughly 25 trillion dollar global knowledge work market.

    The encouraging signal is that founder-controlled companies, the ones most likely to move fast on efficiency, are not conducting large-scale layoffs once you adjust for pandemic-era overhiring. That suggests they see continued opportunity for people plus large token budgets rather than a straight substitution. Data from Cognition, Ramp, and Stripe indicates that companies spending the most on AI are growing meaningfully faster, though Baker acknowledges the skeptics’ point that these datasets do not control for industry.

    The Memory Supply War and LTA Game Theory

    Everything is currently in shortage, and Baker argues the constraint is energizing gigawatts rather than manufacturing. Turbine makers and diesel generator makers are ramping, old aircraft turbines are being stripped and reconditioned for data center power, and regulatory policy is moving favorably. The transition he says he got wrong is the shift, especially in memory, from maximizing short-term pricing to signing long-term agreements with customer prepayments, price floors, and price ceilings.

    The reason those agreements will hold is game theory. Memory is the axis around which everything else revolves, because more memory per unit of compute means more tokens out, which lowers cost per token, which is why demand has shown essentially no negative elasticity. Market share among the four buyers that matter (Amazon with Trainium, Google with TPUs, AMD, and an Nvidia bigger than all of them combined) will be determined for years by supply chain allocation. Break a long-term agreement to chase a lower price in an oversupply year and the supplier can break the volume commitment in return and hand your allocation to a competitor. Since oversupply in this industry is reliably followed by undersupply, that is a decision that can end a franchise. Apple could get away with this historically because its volume was overwhelming and it had no equivalent competitor. That world is gone.

    Nvidia’s New Playbook

    Baker finds Nvidia’s low multiple hard to reconcile with how thoroughly the current environment favors it. If chips need to be financed, nothing on earth is more financeable than an Nvidia GPU. If land and power are the constraint, Nvidia has been playing the matchmaking chess game well. On top of that they have rolled out what Baker describes as a credit wrapper with a revenue share that kicks in when GPU prices sit above a floor. It is not vendor financing, since someone else lends the buyer the money. What it does is give Nvidia a royalty on recurring compute revenue, which could amount to a very large cloud business built entirely out of royalties, while helping bridge the cash flow mismatch between an industry that has gone free cash flow negative and a supplier collecting all the cash.

    Asked what he would do as a memory CEO, Baker says he would do exactly what Nvidia is doing: approach GPU and accelerator buyers, participate in the credit wrapper, perhaps put up cash upfront to make lenders comfortable, and take a cut of ongoing revenue. He expects firms like Blackstone and Apollo are pitching variants of this to the memory companies already. He also thinks the arrangement quietly widens Nvidia’s competitive moat, since startup accelerator companies pay more at the foundry, pay more for high bandwidth memory, and cannot finance their chips at Nvidia’s rate. And he notes that essentially every time Nvidia has declined to take an equity stake in something, it has turned out to be a mistake.

    What Could Break the Thesis

    Pressed for the scenario that would flip him, Baker names two. The first is operating cash flow failing to accelerate, which would force the buildout onto debt and validate the credit bears. That outcome depends largely on whether the combined trajectory of Anthropic, OpenAI, Grok, Cursor, and open source keeps compounding. The second is a sustained sharp contraction in GPU rental prices. The market would react instantly, and it would mean the compute shortage had broken. As of the recording, not a single person he has spoken with says they have too many GPUs.

    The technical wildcard is continual learning and sample efficient learning. Many researchers believe both are close. A human learns effectively on something like 20 billion tokens while frontier models train on 300 trillion, so a model that could be trained on 10 trillion tokens and then learn efficiently in the world would represent a discontinuity in training demand. Baker thinks training will asymptote to a small but nonzero share of compute regardless, and that the development would be extraordinarily good for the world. He also notes Nvidia is deeply involved with essentially all of the labs pursuing it.

    China, Lithography, and Decoupling

    On China’s deep ultraviolet lithography machine, Baker holds both views at once. It is a genuine phase transition, comparable to going from having no propeller plane to having one, because they did not have it before and now allegedly they do. It is also roughly 25 years behind extreme ultraviolet, and lithography is learning by doing, so you cannot teleport through the required cycles. He suspects the market overreacted and that if it ever affects ASML’s order book it will be years out, by which time the market will have forgotten and rediscovered the concern several times.

    He is careful about certainty here. It is very hard for an American to have real clarity on what is happening inside China, the people there are extremely capable and work brutally hard, and they consider this existential for the country. There are unverified reports that an extreme ultraviolet machine was smuggled in, which he treats as noise. His larger point is that decoupling is now self-reinforcing on both sides, it is unfortunate, and neither side is going to stop.

    Regulation, Data Centers, and a Failure of Storytelling

    Asked for the worst thing that could happen to AI, Baker answers regulation without hesitation. New York’s data center moratorium feels like the first of many, and even deep red pro-growth states are telling the industry it is doing a poor job explaining itself. The political narrative among ordinary Americans is that data centers will raise electricity prices, drain water supplies, and eliminate jobs. Baker’s counter is that behind the meter deals generally lower local electricity prices, that community agreements now routinely include hospitals, schools, police stations, and fire stations rather than the old model of buying the fire department new trucks, and that the jobs are ongoing rather than one-time because of continuous maintenance, replacement, and upgrade cycles.

    The water claim is the clearest case of a myth outrunning the correction. An author overstated data center water usage by roughly 10,000 times, has acknowledged the error repeatedly, and the figure still circulates. Patrick offers the parallel of the spinach iron myth, created by a misplaced decimal point in an academic text and still believed 80 years later. Baker’s proposed remedy is blunt: a foundation or political action committee running ads during the Final Four, NFL games, and the World Series explaining what a data center actually does for a community, alongside the story of AI accelerating medical research and improving outcomes for people with serious illness. The people building this find the benefits so obvious that they assume everyone already knows, and they cannot process how divergent their view is from most Americans.

    SRAM Accelerators and Disaggregated Inference

    An underdiscussed development, Baker argues, is what happens when SRAM-based accelerators arrive at scale. These chips are not constrained by high bandwidth memory and are often built on older nodes, so they do not compete for the leading edge capacity that GPUs consume. Inference disaggregates into prefill and decode, and decode splits further into attention and the feed forward network. The holy grail is running prefill on a chip without high bandwidth memory, attention on a high-memory chip, and the feed forward network on SRAM, which nothing beats for that workload. Since workloads keep changing, no single chip can get the ratio of compute to high bandwidth memory to on-die SRAM permanently right, which is precisely the argument for disaggregation. Baker expects this to be strongly positive for the return on investment across the installed base and on new compute.

    SpaceX, Orbital Compute, and Dark Horses

    Baker does not think the market understands SpaceX as a company yet, and he considers it the most important new public company. The fundamentals have improved since the IPO, and the compute story is the part being missed. Only the hyperscalers, CoreWeave, Crusoe, and SpaceX have ever brought more than 500 megawatts of power online in a single year, and SpaceX has done it fastest and cheapest while building clusters customers actually like. When SpaceX dumped a large block of compute into the market, it was absorbed without a blip, which Baker treats as one of the most bullish demand datapoints available. A circulating Substack report claims eight gigawatts within 18 months. He doubts that figure and quotes it only because it is public, but at roughly 50 billion dollars of monetization per gigawatt against a 73 billion dollar consensus estimate, even partial delivery would overwhelm expectations. There is a well-known New York hedge fund short case built on spot compute prices falling 90 percent.

    On orbital compute, Baker says time at Starbase left him thinking it feels more real every day, and the Starship landing reinforced it. His sanity check is that Benchmark, from entirely outside the Elon ecosystem and without the benefit of internal launch costs, chose to fund StarCloud at a real valuation, with SpaceX partnering to provide the Starlink laser technology that orbital compute requires. As he puts it, maybe he is crazy, maybe Elon is crazy, maybe Benchmark is crazy, and maybe the SpaceX engineers are crazy too, but all of that being true simultaneously does not seem probable. Asked for dark horses who could become as consequential as the current giants, he names Lip-Bu Tan, Lin Qiao at Fireworks, and Scott Wu at Cognition. The episode was recorded at Benchmark’s offices, at the table where their dinners are held.

    Notable Quotes

    “I want to be scared. I don’t want to feel like a lunatic watching these stocks get cheaper thinking the expected forward returns are going up.”

    Gavin Baker, on why he spent the week in Silicon Valley hunting for bearish data

    “I would describe July as 2022 in a month.”

    Gavin Baker, characterizing a 40 to 60 percent drawdown in AI names that happened in a straight line

    “Have you heard a single negative quantitative metric about AI? A single instance of deceleration?”

    Gavin Baker to Patrick O’Shaughnessy, framing the central contradiction of the month

    “A token is a token, and you need the exact same amount of compute to make a token. It takes the same amount of flops, the same amount of memory, the same amount of watts.”

    Gavin Baker, on why the open source panic misread infrastructure demand

    “Open source is kind of dark matter to the public markets. It’s hard for public markets to measure it.”

    Gavin Baker, on why the fastest-growing part of inference demand is invisible in audited financials

    “Claude is kind of Walter Cronkite for the stock market and everybody just believes whatever it says. And by the way, it’s really smart, but it’s not always right.”

    Gavin Baker, on the collapse of interpretive diversity among investors

    “Nvidia is actually, as we record this, at its lowest forward PE of the last 10 years.”

    Gavin Baker, noting the only cheaper moments were the DeepSeek shock and Liberation Day, both V-bottoms

    “If you break your LTA and then in the next two or three years for any reason leverage shifts back to the memory guys, you’re out of business.”

    Gavin Baker, on why long-term agreements will hold through the next memory cycle

    “If you need to be able to finance the chips, and you do, nothing’s more financeable than an Nvidia GPU. Nothing.”

    Gavin Baker, on why the current environment favors Nvidia more than its multiple suggests

    “Data centers are in a lot of ways the best thing to happen for blue collar wages in my lifetime.”

    Gavin Baker, on the gap between the political narrative and the local economics

    “A lie could go around the world faster than truth gets out of bed.”

    Gavin Baker, on a data center water usage figure overstated by roughly 10,000 times that still circulates

    “One of Elon’s phrases is we specialize in making the impossible late.”

    Gavin Baker, on why he doubts the eight gigawatt figure without betting against SpaceX

    Watch the full conversation here: Why the Markets Are Pricing AI Wrong with Gavin Baker on Invest Like the Best.

    Related Reading

    • Invest Like the Best on Colossus the show’s home, where the full episode archive and transcripts live.
    • Atreides Management Gavin Baker’s firm and the vantage point behind these compute and semiconductor calls.
    • More Than You Know by Michael Mauboussin, the source of the diversity breakdown framework Baker invokes to explain why markets crash when everyone reasons the same way.
    • High Bandwidth Memory (Wikipedia) background on the memory technology that sits at the center of the long-term agreement game theory.
    • Fireworks AI the inference cloud whose routing and fine-tuning stack Baker credits with making open source models competitive for production workloads.
  • Jonathan Ross on Groq’s $20 Billion NVIDIA Deal, Faster Inference, and Why Asking the Right Questions Wins the AI Age

    Jonathan Ross, the founder of Groq and the inventor of Google’s Tensor Processing Unit (TPU), sits down with David Senra (host of the Founders podcast) to walk through Groq’s roughly $20 billion partnership with NVIDIA and the decade of near-death struggle that preceded it. You can watch the full conversation here. Ross, now a senior executive at NVIDIA following the deal, is unusually candid about being one of the world’s worst leaders when he started, about coming three weeks from running out of money, and about the single contrarian bet (that faster inference would make AI both faster and smarter) that almost everyone, including his own engineers, told him was pointless.

    TLDW

    Ross explains the structure of the NVIDIA deal (a call to Jensen Huang about buying 100,000 GPUs turned, in three weeks, into NVIDIA’s largest deal by nearly 3x) and why pairing Groq’s LPU with the GPU defeats the many different bottlenecks inside an LLM the way you would use both 18-wheelers and delivery vans in a logistics network. He unpacks the AlphaGo moment that revealed faster inference makes models smarter, the shift from the information age (answering questions) to the AI age (asking the right questions), and a leadership philosophy built on autonomy, one brutally clear priority (25 million tokens per second on a challenge coin), and giving people the fewest constraints so they can surprise you. He shares hard-won lessons from Jensen and NVIDIA (the least political large org he has seen, no secret one-on-ones), his concepts of reality quotient and the dominant game, return on luck and the GitHub opportunity he let his team talk him out of, intentional leadership (“I intend to do this”), the Grok bonds that traded salary for equity and saved the company, hiring for negatives instead of positives, loss bias and manufactured discontent, and a closing case for radical optimism: code is becoming free, software creation is being democratized like literacy, and education should stop teaching kids to answer questions and start teaching them to ask.

    Thoughts

    The technical spine of this interview is a genuinely counterintuitive claim: you can make a model smarter by making it faster. Ross’s proof is the AlphaGo anecdote, where the exact same model, ported from GPUs to his TPU, saw its ELO jump by hundreds of points and beat the world champion, because more compute per unit of time let it search deeper and surface moves like the famous Move 37 that were too far down the tree to find otherwise. Once you internalize that inference speed is not a convenience but a capability multiplier, the entire Groq thesis, and the logic of the NVIDIA deal, snaps into focus. The industry spent years treating fast inference as a nice-to-have. Ross treated it as the whole game, and was nearly alone in doing so for a very long time.

    The most transferable material is the leadership arc, precisely because Ross is willing to say he was bad at it. His core insight is that there is no single correct way to lead, any more than there is one way to invest, and the founder’s first job is to know which way is true to them. Ross is a delegator who hires autonomous people and gives them a single, poetically compressed objective, then gets out of the way. The reason that matters is subtle: if you over-constrain the goal, your team can never surprise you with a better answer than the one you already had, which means they can never actually innovate. The Kelly Johnson line Senra offers (“extreme performance often comes from one brutally clear priority”) is the same idea from the Skunk Works side. A challenge coin that reads “25 million tokens per second” is not a slogan, it is a mechanism that lets every engineer connect their work to one dominant game.

    Two ideas deserve to be lifted out and used directly. The first is intentional leadership, borrowed from David Marquet’s submarine turnaround: replace “should I do this?” with “I intend to do this.” Asking for opinions invites pessimism and hands your most timid people a veto. Declaring intent still lets someone shout “the hatch is open” when it truly matters, but it stops the reflexive no. Ross traces years of stalled progress to the simple error of asking instead of declaring. The second is his inversion of hiring: hire for negatives, not positives. Growing talent means showing people the path, so you emphasize positives. Selecting talent means screening people out, so you hunt for the disqualifying negatives, because one person’s negative trait infects the whole team. Most founders, Ross included for years, are clever enough to talk themselves into any candidate. A versioned “people spec” and a deliberate loss-averse posture are the antidote.

    The Grok bonds story is the emotional center and a small masterpiece of change management. Facing a layoff list that would have killed the company (because the people slated to be cut were exactly the ones needed to make the product work at all), Ross instead asked the team to trade salary for equity, framed with World War II war-bond imagery. Eighty percent participated, half went to statutory minimum wage, and attrition actually fell. His phrase for why is “put everyone’s hands on the steering wheel.” Passengers fear a windy road, drivers feel in control. It is a reminder that morale under existential stress is often a function of agency, not comfort, and that the Phil Knight move of converting employee sacrifice into ownership is a recurring pattern in company survival stories for a reason.

    Where the conversation turns almost spiritual is manufactured discontent. Ross observes that the entrepreneurs in a room of successful people were the least happy with their wealth, and that this very dissatisfaction was the fuel that kept them building. His own current discontent is stark and worth sitting with: the world does not have enough compute, and if it takes an extra year to cure cancer or slow aging because of that shortage, he considers it his fault. Whether or not you accept the moral weight he assigns himself, the mechanism is instructive. Edwin Land wrote “300 people died today” on the whiteboard while inventing anti-glare technology. A concrete, human cost attached to delay is a far more durable motivator than a revenue target. Paired with his closing optimism about code becoming free and software creation democratizing like literacy, it makes for one of the more clear-eyed and yet hopeful founder conversations in recent memory.

    Key Takeaways

    • The NVIDIA deal began as a request to buy about 100,000 GPUs; Jensen saw what Groq had built pairing GPUs and LPUs and decided to make it available to all NVIDIA customers, closing what Ross calls the firm’s biggest deal by nearly 3x in roughly three weeks from first call to wired money.
    • GPUs and LPUs are complementary: inside an LLM’s decoder layer, the GPU is better at the compute-bound attention portion and the LPU is better at the memory-throughput-bound weights, so combining them defeats bottlenecks across the whole performance curve, like using both 18-wheelers and last-mile vans.
    • As AI increasingly talks to AI, speed dominates, because agents kick off other agents and compound; a human tolerates a one-second wait, but AI is just sitting there idle.
    • Agentic micro payments will make the number of payments skyrocket, but payments infrastructure is not yet built for AI operating inside an allocated budget.
    • Ross prototypes cutting-edge ideas as personal hobby projects first, then brings them to work; his personalized “daily brief” evolved from long text into headlines he can interrogate with follow-up questions, like the game of 20 questions.
    • The information age rewarded answering questions; the AI age rewards asking the right ones, as everyone shifts from individual contributor to leader of AI, and good leaders ask the question no one else did.
    • There is no single right way to lead, just as there are many ways to invest; the founder’s job is to know themselves and pick the leadership form that is true to them (inspiration versus fear, control versus delegation).
    • Ross was, by his own account, one of the world’s worst leaders at the start, which cost Groq three to four years; his fix was to define one goal simple enough to fit on a challenge coin: 25 million tokens per second.
    • The fewer constraints you give a person (or an AI agent), the more freedom they have to surprise you with a better solution; over-constraining the goal makes real innovation impossible.
    • Lessons from Jensen and NVIDIA: it is the least political large organization Ross has seen, Jensen never runs secret one-on-ones (tell everyone at once, copy everyone on email), and the whole strategy reduces to “what does the customer actually need?”
    • Jensen manages around 60 direct reports, each smarter than him in their own domain, which he offers as the model for orchestrating AI agents that may be smarter than you.
    • Asking a sharp question that makes an expert say “I didn’t think of that” is a universal founder skill (it appears in every Bezos book) and can be honed.
    • Confidence, not competence, was Ross’s early bottleneck: shadowing a leader of 2,000 people, he realized he would have made the same decisions, and acting with confidence made people follow his direction without changing the decisions themselves.
    • The better and more creative your people, the harder they are to manage; running 450 highly creative scientists felt more like managing 5,000.
    • Reality quotient (RQ), distinct from IQ, is the ability to recognize reality and, in its extreme form, to choose the dominant game; MySpace optimized accounts signed up while Facebook optimized monthly active users and won.
    • The first principle of change management is to make it feel like it is not a change; people who seem fine with change are usually anchored to something that did not change.
    • Return on luck (from Jim Collins): the most successful companies do not get more lucky breaks, they seize the ones they get; Ross let his team talk him out of powering GitHub’s LLMs on Groq chips, then vowed never again.
    • People adopt fast inference only when they experience it personally; an Anthropic demo three months before ChatGPT drew no reaction because the answers were not the audience’s own, and Groq later went viral off a fast-LLM video posted on X.
    • Great innovators often experience a problem before others do; the future is already here, just not evenly distributed, and Ross saw fast inference’s value first because of AlphaGo.
    • Intentional leadership (from David Marquet’s USS Santa Fe turnaround): say “I intend to do this” instead of asking for an opinion, which stops reflexive pessimism while still letting people flag a real problem.
    • Grok bonds: three weeks from running out of money, Ross swapped a layoff for a war-bond-style salary-for-equity exchange; 80% participated, about half took statutory minimum wage, and it bought roughly two months of runway.
    • “Put everyone’s hands on the steering wheel”: participation in saving the company cut attrition to under 10% during the crisis, echoing Phil Knight converting employee loans into Nike equity.
    • West Coast VCs behave like lemmings (one pass triggers all passes), while East Coast VCs run independent analysis; the herd missed what became NVIDIA’s biggest deal ever, a live example of the Keynesian beauty contest.
    • For the first time, top startups are not starved for cash, so putting in more money is no longer an advantage even though investors still behave as if it is.
    • Hiring flip: move from hiring for positives (how you grow talent) to hiring for negatives (how you select talent), because one negative trait poisons the team; write a versioned “people spec” like a product spec.
    • Loss bias (a loss feels roughly six times more painful than an equal gain) can be a hiring signal: Ross looks for people who “book the win early,” treating any missed improvement as a loss.
    • Poetic design (maximum meaning in minimal expression, “every word matters”) was a positive on the people spec; its negative is maximalist, cluttered design.
    • Michael Jordan manufactured pressure by taunting opponents so a loss would be humiliating, forcing superhuman performance (per his trainer Tim Grover), a deliberate version of throwing your keys over the fence.
    • Manufactured discontent (David Ogilvy’s “divine discontent”): the best entrepreneurs never rest on wins; the least happy people with their wealth were the ones who kept building.
    • Ross’s discontent today is the world’s lack of compute; he treats every delayed medical breakthrough as partly his responsibility, the way Edwin Land wrote a daily death count on the whiteboard while fighting headlight glare.
    • Software has run on “code rationing” because code was expensive to write, enforced by “no engineers”; as the marginal cost of code approaches zero, you just implement, experience, and re-implement.
    • AI democratizes software creation like the alphabet democratized literacy: Ross’s executive assistant now builds working apps, and individual founders with taste but no coding background will create valuable companies.
    • Education should be revamped around asking questions and solving real community problems; if a kid can look up or prompt the answer, the assignment taught nothing, but making them ask the right questions to get AI to solve a real problem does.

    Detailed Summary

    The $20 Billion NVIDIA Deal and Why LPUs and GPUs Belong Together

    The deal’s most striking feature is speed: the idea was first floated on a call roughly three weeks before the money was in the bank. Groq had been integrating GPUs and LPUs and went to Jensen Huang wanting to buy about 100,000 GPUs to deploy themselves. Jensen saw the combined system and decided it should be offered to all of NVIDIA’s customers. The technical logic is that processing an LLM token involves many matrix multiplies with different bottlenecks, some compute-constrained (better on the GPU, especially the attention portion) and some memory-throughput-constrained (better on the LPU, applying the trained weights). There is no single perfect architecture, so putting the two together defeats bottlenecks across the whole curve. Ross adds that as AI talks to AI, speed becomes everything, because agents spawn agents and compound exponentially.

    Asking Questions, Daily Briefs, and the Shift to Leading AI

    Ross builds cutting-edge tools as personal hobby projects before bringing them to work, including a personalized “daily brief” that functions like a presidential daily brief. He redesigned it from long text into headlines he can interrogate, because interactivity, like 20 questions, distills straight to what you actually care about. This grounds one of his signature ideas: success in the information age meant answering questions, but success in the AI age means asking the right questions. As people move from individual contributors to leaders of AI, the skill that matters is the leader’s skill of asking the question everyone else missed or was afraid to raise, since the question you ask determines the output you get.

    Knowing Your Leadership Style and the Challenge Coin

    Ross frames leadership like investing: the first principle is simply having followers, but there are infinite valid styles. New founders fail by copying advice that is not true to them. Ross is a natural delegator (he has not held a driver’s license since his teens because he would rather think than control the car) who hires unusually autonomous people. Early on this backfired badly, because he entrusted people who needed direction, and he calls himself one of the world’s worst early leaders, a gap that cost Groq years. His breakthrough was distilling the mission onto a challenge coin reading “25 million tokens per second,” which let everyone connect their work to one dominant game. He references David Marquet’s Turn the Ship Around later, but the coin embodies Kelly Johnson’s Skunk Works principle that extreme performance comes from one brutally clear priority, plus the rule that fewer constraints give people more room to surprise you, turning a team from Superman into the Avengers.

    Lessons from Jensen: Killing Politics and Serving the Customer

    Working at NVIDIA taught Ross how much further he could have pushed lessons he half-learned at Groq. NVIDIA is, in his experience, the least political large organization anywhere, and a big reason is that Jensen never tells different people different things in private one-on-ones. When you address a room, everyone hears the same message; separate conversations breed side cliques. Ross’s practical rules: hold big meetings for anything you want a group to know, and copy everyone on email so no one can route politics through you. The other Jensen lesson is to stop playing 3D chess and just ask what the customer needs, tell them only what you believe and can support, and refuse to sell them something they do not need. Senra notes he has covered roughly 19 ideas from The Nvidia Way on his Founders podcast, and Jensen’s line that he already manages 60 reports smarter than him is the template for managing AI agents.

    Reality Quotient, the Dominant Game, and Change Management

    Groq hired for reality quotient, not just IQ, because plenty of very smart people construct elaborate stories disconnected from reality. In its extreme form, RQ is the ability to choose the dominant game, the way Facebook’s focus on monthly active users beat MySpace’s focus on accounts signed up. The founder’s job is to help everyone connect their activity to that dominant game (for Groq, tokens per second), then manage the change. Ross’s first principle of change management is to make it feel like it is not a change: nobody likes change, and people who tolerate it well are usually focused on something that stayed constant. If your team is anchored to the dominant goal, a new tactic does not feel like change; if they are anchored to a narrow task, it does.

    Return on Luck, the AlphaGo Insight, and the GitHub Miss

    From Jim Collins’s Great by Choice, Ross took the idea that winners seize luck better, not that they get more of it. He experienced it first-hand with AlphaGo: after a DeepMind team asked whether his TPU was as fast as rumored (he said yes, Ghostbusters-style), porting the identical model from GPUs to TPUs pushed its ELO from around 3,200 to roughly 3,900 and it crushed the world champion. As Thinking Fast and Slow by Daniel Kahneman frames it, more compute lets the model virtually play out more moves and occasionally find a better second-best line, which is how the famous Move 37 surfaced. Faster thinking is smarter thinking. Yet Ross also let his own engineers talk him out of powering GitHub’s LLMs on Groq chips, twice, because they focused on why it could not be done rather than why it could. He eventually did the math himself, hit the numbers, and learned to stop inviting that pessimism.

    Selling Speed and Intentional Leadership

    Customers could not grasp fast inference until they felt it. Ross recalls an Anthropic demo three months before ChatGPT that drew no reaction, because seeing someone else’s answer appear is not magical, but getting your own question answered instantly is. So Groq simply put fast inference online, and it went viral after someone posted a video of a blazing-fast LLM on X (Ross noticed his own demo slowing in Norway because usage had skyrocketed). The deeper fix for internal resistance came from Turn the Ship Around, David Marquet’s account of turning the USS Santa Fe from worst to best in nuclear readiness by replacing command-and-control with intentional leadership. Saying “I intend to do this” rather than “should I?” stops people from reflexively supplying negative opinions, while still letting someone shout “the hatch is open” when there is a genuine problem.

    Grok Bonds: Three Weeks From Zero

    With three weeks of cash left and a layoff list on the table, Ross realized the cuts targeted exactly the people needed to finish an unprecedented compiler and reach the critical mass where the product would even work. Layoffs would not save the company; only reducing burn without losing people could. So Groq held an all-hands, put up World War II war-bond imagery, and launched “Grok bonds,” an exchange of salary for equity. Ross expected heavy attrition; instead 80% participated and about half dropped to statutory minimum wage, real pain for engineers used to six-figure salaries. It bought closer to two months of runway. His framing, “put everyone’s hands on the steering wheel,” explains why attrition actually fell below 10%: drivers feel more in control than passengers, and it echoes Phil Knight in Shoe Dog converting employee loans into Nike equity on the edge of collapse.

    Hiring for Negatives, Loss Bias, and Manufactured Discontent

    Ross was good at spotting smart, talented people but kept hiring ones who caused organizational problems, because he could always talk himself into a candidate. Watching a sharp head of HR screen people out, he realized he had been hiring wrong: growing talent means showing positives, but selecting talent means hunting for disqualifying negatives, since one bad trait spreads to the whole team. He formalized a versioned “people spec” with positives like return on luck and poetic design, each paired with a negative. He also hired for loss bias, the fact that a loss feels roughly six times more painful than an equal gain, seeking people who “book the win early.” That competitive, pressure-seeking wiring links to Michael Jordan manufacturing humiliation stakes (per Tim Grover in Relentless) and to David Ogilvy’s divine discontent. Ross’s own manufactured discontent today is the world’s shortage of compute, which he frames in life-and-death terms.

    The Optimistic Close: Free Code and Universal Software Literacy

    Ross ends on aggressive optimism. Software has long run on “code rationing” because code was expensive to write, policed by “no engineers” whose job is to say no. As the marginal cost of code approaches zero, the workflow flips to implement, experience, then re-implement. More important is accessibility: just as alphabets and universal education turned reading and writing from a scribe’s monopoly into a question of quality, AI is making software creation universal. His executive assistant now builds working apps, and a wave of individual founders with taste but no coding background will create valuable companies. The corollary for education is to stop teaching kids to answer questions and start teaching them to ask, revamping curricula around real community problems where the point is asking the right questions to get AI to solve something that matters.

    Notable Quotes

    “Success in the information age was about being able to answer questions. Success in the AI age will be about being able to ask the right questions.”

    Jonathan Ross, on the fundamental shift AI creates

    “The fewer constraints that you give someone, the more freedom they have to solve the problem, and the more freedom they have to surprise you with the solution.”

    Jonathan Ross, on leading creative teams

    “Being able to think faster makes you think smarter.”

    Jonathan Ross, on why faster inference produces more capable models

    “There are plenty of really smart people who wouldn’t recognize reality if it tapped them on the shoulder.”

    Jonathan Ross, defining reality quotient versus IQ

    “If you express intentional leadership, you say, ‘I intend to do this.’ People don’t tend to offer their opinion, but if it’s very wrong and there’s a reason, they will push back.”

    Jonathan Ross, on the lesson from Turn the Ship Around

    “When people are passengers in a car, they’re more nervous about a windy road or a scary road. But when they’re the driver, they feel more in control.”

    Jonathan Ross, on why Grok bonds kept the team together

    “The biggest flip in my hiring was when I went from looking for positives, which is what you do when you’re trying to grow talent, to looking for negatives, which is what you do when you’re trying to select talent.”

    Jonathan Ross, on inverting his approach to hiring

    “If it takes us an extra year to cure cancer because we don’t have enough compute, that’s my fault.”

    Jonathan Ross, on the discontent that drives him today

    Watch the full conversation between Jonathan Ross and David Senra here on YouTube.

    Related Reading

    • Groq the company Ross founded and the LPU behind the fast-inference story and the NVIDIA partnership.
    • AlphaGo versus Lee Sedol (Wikipedia) the match, including Move 37, that showed Ross how much faster hardware raises a model’s capability.
    • The Keynesian Beauty Contest (Wikipedia) the dynamic Ross uses to explain why West Coast VCs herded past what became NVIDIA’s biggest deal.
    • Zero to One by Peter Thiel, the source of the first-principles thinking Ross applied to the contrarian bet on fast inference.
    • Founders podcast by David Senra the host’s biography-driven show, source of the Jensen, Michael Jordan, and Edwin Land ideas referenced throughout.