PJFP.com

Pursuit of Joy, Fulfillment, and Purpose

Tag: prefill

  • Jonathan Ross on Groq’s $20 Billion NVIDIA Deal, Faster Inference, and Why Asking the Right Questions Wins the AI Age

    Jonathan Ross, the founder of Groq and the inventor of Google’s Tensor Processing Unit (TPU), sits down with David Senra (host of the Founders podcast) to walk through Groq’s roughly $20 billion partnership with NVIDIA and the decade of near-death struggle that preceded it. You can watch the full conversation here. Ross, now a senior executive at NVIDIA following the deal, is unusually candid about being one of the world’s worst leaders when he started, about coming three weeks from running out of money, and about the single contrarian bet (that faster inference would make AI both faster and smarter) that almost everyone, including his own engineers, told him was pointless.

    TLDW

    Ross explains the structure of the NVIDIA deal (a call to Jensen Huang about buying 100,000 GPUs turned, in three weeks, into NVIDIA’s largest deal by nearly 3x) and why pairing Groq’s LPU with the GPU defeats the many different bottlenecks inside an LLM the way you would use both 18-wheelers and delivery vans in a logistics network. He unpacks the AlphaGo moment that revealed faster inference makes models smarter, the shift from the information age (answering questions) to the AI age (asking the right questions), and a leadership philosophy built on autonomy, one brutally clear priority (25 million tokens per second on a challenge coin), and giving people the fewest constraints so they can surprise you. He shares hard-won lessons from Jensen and NVIDIA (the least political large org he has seen, no secret one-on-ones), his concepts of reality quotient and the dominant game, return on luck and the GitHub opportunity he let his team talk him out of, intentional leadership (“I intend to do this”), the Grok bonds that traded salary for equity and saved the company, hiring for negatives instead of positives, loss bias and manufactured discontent, and a closing case for radical optimism: code is becoming free, software creation is being democratized like literacy, and education should stop teaching kids to answer questions and start teaching them to ask.

    Thoughts

    The technical spine of this interview is a genuinely counterintuitive claim: you can make a model smarter by making it faster. Ross’s proof is the AlphaGo anecdote, where the exact same model, ported from GPUs to his TPU, saw its ELO jump by hundreds of points and beat the world champion, because more compute per unit of time let it search deeper and surface moves like the famous Move 37 that were too far down the tree to find otherwise. Once you internalize that inference speed is not a convenience but a capability multiplier, the entire Groq thesis, and the logic of the NVIDIA deal, snaps into focus. The industry spent years treating fast inference as a nice-to-have. Ross treated it as the whole game, and was nearly alone in doing so for a very long time.

    The most transferable material is the leadership arc, precisely because Ross is willing to say he was bad at it. His core insight is that there is no single correct way to lead, any more than there is one way to invest, and the founder’s first job is to know which way is true to them. Ross is a delegator who hires autonomous people and gives them a single, poetically compressed objective, then gets out of the way. The reason that matters is subtle: if you over-constrain the goal, your team can never surprise you with a better answer than the one you already had, which means they can never actually innovate. The Kelly Johnson line Senra offers (“extreme performance often comes from one brutally clear priority”) is the same idea from the Skunk Works side. A challenge coin that reads “25 million tokens per second” is not a slogan, it is a mechanism that lets every engineer connect their work to one dominant game.

    Two ideas deserve to be lifted out and used directly. The first is intentional leadership, borrowed from David Marquet’s submarine turnaround: replace “should I do this?” with “I intend to do this.” Asking for opinions invites pessimism and hands your most timid people a veto. Declaring intent still lets someone shout “the hatch is open” when it truly matters, but it stops the reflexive no. Ross traces years of stalled progress to the simple error of asking instead of declaring. The second is his inversion of hiring: hire for negatives, not positives. Growing talent means showing people the path, so you emphasize positives. Selecting talent means screening people out, so you hunt for the disqualifying negatives, because one person’s negative trait infects the whole team. Most founders, Ross included for years, are clever enough to talk themselves into any candidate. A versioned “people spec” and a deliberate loss-averse posture are the antidote.

    The Grok bonds story is the emotional center and a small masterpiece of change management. Facing a layoff list that would have killed the company (because the people slated to be cut were exactly the ones needed to make the product work at all), Ross instead asked the team to trade salary for equity, framed with World War II war-bond imagery. Eighty percent participated, half went to statutory minimum wage, and attrition actually fell. His phrase for why is “put everyone’s hands on the steering wheel.” Passengers fear a windy road, drivers feel in control. It is a reminder that morale under existential stress is often a function of agency, not comfort, and that the Phil Knight move of converting employee sacrifice into ownership is a recurring pattern in company survival stories for a reason.

    Where the conversation turns almost spiritual is manufactured discontent. Ross observes that the entrepreneurs in a room of successful people were the least happy with their wealth, and that this very dissatisfaction was the fuel that kept them building. His own current discontent is stark and worth sitting with: the world does not have enough compute, and if it takes an extra year to cure cancer or slow aging because of that shortage, he considers it his fault. Whether or not you accept the moral weight he assigns himself, the mechanism is instructive. Edwin Land wrote “300 people died today” on the whiteboard while inventing anti-glare technology. A concrete, human cost attached to delay is a far more durable motivator than a revenue target. Paired with his closing optimism about code becoming free and software creation democratizing like literacy, it makes for one of the more clear-eyed and yet hopeful founder conversations in recent memory.

    Key Takeaways

    • The NVIDIA deal began as a request to buy about 100,000 GPUs; Jensen saw what Groq had built pairing GPUs and LPUs and decided to make it available to all NVIDIA customers, closing what Ross calls the firm’s biggest deal by nearly 3x in roughly three weeks from first call to wired money.
    • GPUs and LPUs are complementary: inside an LLM’s decoder layer, the GPU is better at the compute-bound attention portion and the LPU is better at the memory-throughput-bound weights, so combining them defeats bottlenecks across the whole performance curve, like using both 18-wheelers and last-mile vans.
    • As AI increasingly talks to AI, speed dominates, because agents kick off other agents and compound; a human tolerates a one-second wait, but AI is just sitting there idle.
    • Agentic micro payments will make the number of payments skyrocket, but payments infrastructure is not yet built for AI operating inside an allocated budget.
    • Ross prototypes cutting-edge ideas as personal hobby projects first, then brings them to work; his personalized “daily brief” evolved from long text into headlines he can interrogate with follow-up questions, like the game of 20 questions.
    • The information age rewarded answering questions; the AI age rewards asking the right ones, as everyone shifts from individual contributor to leader of AI, and good leaders ask the question no one else did.
    • There is no single right way to lead, just as there are many ways to invest; the founder’s job is to know themselves and pick the leadership form that is true to them (inspiration versus fear, control versus delegation).
    • Ross was, by his own account, one of the world’s worst leaders at the start, which cost Groq three to four years; his fix was to define one goal simple enough to fit on a challenge coin: 25 million tokens per second.
    • The fewer constraints you give a person (or an AI agent), the more freedom they have to surprise you with a better solution; over-constraining the goal makes real innovation impossible.
    • Lessons from Jensen and NVIDIA: it is the least political large organization Ross has seen, Jensen never runs secret one-on-ones (tell everyone at once, copy everyone on email), and the whole strategy reduces to “what does the customer actually need?”
    • Jensen manages around 60 direct reports, each smarter than him in their own domain, which he offers as the model for orchestrating AI agents that may be smarter than you.
    • Asking a sharp question that makes an expert say “I didn’t think of that” is a universal founder skill (it appears in every Bezos book) and can be honed.
    • Confidence, not competence, was Ross’s early bottleneck: shadowing a leader of 2,000 people, he realized he would have made the same decisions, and acting with confidence made people follow his direction without changing the decisions themselves.
    • The better and more creative your people, the harder they are to manage; running 450 highly creative scientists felt more like managing 5,000.
    • Reality quotient (RQ), distinct from IQ, is the ability to recognize reality and, in its extreme form, to choose the dominant game; MySpace optimized accounts signed up while Facebook optimized monthly active users and won.
    • The first principle of change management is to make it feel like it is not a change; people who seem fine with change are usually anchored to something that did not change.
    • Return on luck (from Jim Collins): the most successful companies do not get more lucky breaks, they seize the ones they get; Ross let his team talk him out of powering GitHub’s LLMs on Groq chips, then vowed never again.
    • People adopt fast inference only when they experience it personally; an Anthropic demo three months before ChatGPT drew no reaction because the answers were not the audience’s own, and Groq later went viral off a fast-LLM video posted on X.
    • Great innovators often experience a problem before others do; the future is already here, just not evenly distributed, and Ross saw fast inference’s value first because of AlphaGo.
    • Intentional leadership (from David Marquet’s USS Santa Fe turnaround): say “I intend to do this” instead of asking for an opinion, which stops reflexive pessimism while still letting people flag a real problem.
    • Grok bonds: three weeks from running out of money, Ross swapped a layoff for a war-bond-style salary-for-equity exchange; 80% participated, about half took statutory minimum wage, and it bought roughly two months of runway.
    • “Put everyone’s hands on the steering wheel”: participation in saving the company cut attrition to under 10% during the crisis, echoing Phil Knight converting employee loans into Nike equity.
    • West Coast VCs behave like lemmings (one pass triggers all passes), while East Coast VCs run independent analysis; the herd missed what became NVIDIA’s biggest deal ever, a live example of the Keynesian beauty contest.
    • For the first time, top startups are not starved for cash, so putting in more money is no longer an advantage even though investors still behave as if it is.
    • Hiring flip: move from hiring for positives (how you grow talent) to hiring for negatives (how you select talent), because one negative trait poisons the team; write a versioned “people spec” like a product spec.
    • Loss bias (a loss feels roughly six times more painful than an equal gain) can be a hiring signal: Ross looks for people who “book the win early,” treating any missed improvement as a loss.
    • Poetic design (maximum meaning in minimal expression, “every word matters”) was a positive on the people spec; its negative is maximalist, cluttered design.
    • Michael Jordan manufactured pressure by taunting opponents so a loss would be humiliating, forcing superhuman performance (per his trainer Tim Grover), a deliberate version of throwing your keys over the fence.
    • Manufactured discontent (David Ogilvy’s “divine discontent”): the best entrepreneurs never rest on wins; the least happy people with their wealth were the ones who kept building.
    • Ross’s discontent today is the world’s lack of compute; he treats every delayed medical breakthrough as partly his responsibility, the way Edwin Land wrote a daily death count on the whiteboard while fighting headlight glare.
    • Software has run on “code rationing” because code was expensive to write, enforced by “no engineers”; as the marginal cost of code approaches zero, you just implement, experience, and re-implement.
    • AI democratizes software creation like the alphabet democratized literacy: Ross’s executive assistant now builds working apps, and individual founders with taste but no coding background will create valuable companies.
    • Education should be revamped around asking questions and solving real community problems; if a kid can look up or prompt the answer, the assignment taught nothing, but making them ask the right questions to get AI to solve a real problem does.

    Detailed Summary

    The $20 Billion NVIDIA Deal and Why LPUs and GPUs Belong Together

    The deal’s most striking feature is speed: the idea was first floated on a call roughly three weeks before the money was in the bank. Groq had been integrating GPUs and LPUs and went to Jensen Huang wanting to buy about 100,000 GPUs to deploy themselves. Jensen saw the combined system and decided it should be offered to all of NVIDIA’s customers. The technical logic is that processing an LLM token involves many matrix multiplies with different bottlenecks, some compute-constrained (better on the GPU, especially the attention portion) and some memory-throughput-constrained (better on the LPU, applying the trained weights). There is no single perfect architecture, so putting the two together defeats bottlenecks across the whole curve. Ross adds that as AI talks to AI, speed becomes everything, because agents spawn agents and compound exponentially.

    Asking Questions, Daily Briefs, and the Shift to Leading AI

    Ross builds cutting-edge tools as personal hobby projects before bringing them to work, including a personalized “daily brief” that functions like a presidential daily brief. He redesigned it from long text into headlines he can interrogate, because interactivity, like 20 questions, distills straight to what you actually care about. This grounds one of his signature ideas: success in the information age meant answering questions, but success in the AI age means asking the right questions. As people move from individual contributors to leaders of AI, the skill that matters is the leader’s skill of asking the question everyone else missed or was afraid to raise, since the question you ask determines the output you get.

    Knowing Your Leadership Style and the Challenge Coin

    Ross frames leadership like investing: the first principle is simply having followers, but there are infinite valid styles. New founders fail by copying advice that is not true to them. Ross is a natural delegator (he has not held a driver’s license since his teens because he would rather think than control the car) who hires unusually autonomous people. Early on this backfired badly, because he entrusted people who needed direction, and he calls himself one of the world’s worst early leaders, a gap that cost Groq years. His breakthrough was distilling the mission onto a challenge coin reading “25 million tokens per second,” which let everyone connect their work to one dominant game. He references David Marquet’s Turn the Ship Around later, but the coin embodies Kelly Johnson’s Skunk Works principle that extreme performance comes from one brutally clear priority, plus the rule that fewer constraints give people more room to surprise you, turning a team from Superman into the Avengers.

    Lessons from Jensen: Killing Politics and Serving the Customer

    Working at NVIDIA taught Ross how much further he could have pushed lessons he half-learned at Groq. NVIDIA is, in his experience, the least political large organization anywhere, and a big reason is that Jensen never tells different people different things in private one-on-ones. When you address a room, everyone hears the same message; separate conversations breed side cliques. Ross’s practical rules: hold big meetings for anything you want a group to know, and copy everyone on email so no one can route politics through you. The other Jensen lesson is to stop playing 3D chess and just ask what the customer needs, tell them only what you believe and can support, and refuse to sell them something they do not need. Senra notes he has covered roughly 19 ideas from The Nvidia Way on his Founders podcast, and Jensen’s line that he already manages 60 reports smarter than him is the template for managing AI agents.

    Reality Quotient, the Dominant Game, and Change Management

    Groq hired for reality quotient, not just IQ, because plenty of very smart people construct elaborate stories disconnected from reality. In its extreme form, RQ is the ability to choose the dominant game, the way Facebook’s focus on monthly active users beat MySpace’s focus on accounts signed up. The founder’s job is to help everyone connect their activity to that dominant game (for Groq, tokens per second), then manage the change. Ross’s first principle of change management is to make it feel like it is not a change: nobody likes change, and people who tolerate it well are usually focused on something that stayed constant. If your team is anchored to the dominant goal, a new tactic does not feel like change; if they are anchored to a narrow task, it does.

    Return on Luck, the AlphaGo Insight, and the GitHub Miss

    From Jim Collins’s Great by Choice, Ross took the idea that winners seize luck better, not that they get more of it. He experienced it first-hand with AlphaGo: after a DeepMind team asked whether his TPU was as fast as rumored (he said yes, Ghostbusters-style), porting the identical model from GPUs to TPUs pushed its ELO from around 3,200 to roughly 3,900 and it crushed the world champion. As Thinking Fast and Slow by Daniel Kahneman frames it, more compute lets the model virtually play out more moves and occasionally find a better second-best line, which is how the famous Move 37 surfaced. Faster thinking is smarter thinking. Yet Ross also let his own engineers talk him out of powering GitHub’s LLMs on Groq chips, twice, because they focused on why it could not be done rather than why it could. He eventually did the math himself, hit the numbers, and learned to stop inviting that pessimism.

    Selling Speed and Intentional Leadership

    Customers could not grasp fast inference until they felt it. Ross recalls an Anthropic demo three months before ChatGPT that drew no reaction, because seeing someone else’s answer appear is not magical, but getting your own question answered instantly is. So Groq simply put fast inference online, and it went viral after someone posted a video of a blazing-fast LLM on X (Ross noticed his own demo slowing in Norway because usage had skyrocketed). The deeper fix for internal resistance came from Turn the Ship Around, David Marquet’s account of turning the USS Santa Fe from worst to best in nuclear readiness by replacing command-and-control with intentional leadership. Saying “I intend to do this” rather than “should I?” stops people from reflexively supplying negative opinions, while still letting someone shout “the hatch is open” when there is a genuine problem.

    Grok Bonds: Three Weeks From Zero

    With three weeks of cash left and a layoff list on the table, Ross realized the cuts targeted exactly the people needed to finish an unprecedented compiler and reach the critical mass where the product would even work. Layoffs would not save the company; only reducing burn without losing people could. So Groq held an all-hands, put up World War II war-bond imagery, and launched “Grok bonds,” an exchange of salary for equity. Ross expected heavy attrition; instead 80% participated and about half dropped to statutory minimum wage, real pain for engineers used to six-figure salaries. It bought closer to two months of runway. His framing, “put everyone’s hands on the steering wheel,” explains why attrition actually fell below 10%: drivers feel more in control than passengers, and it echoes Phil Knight in Shoe Dog converting employee loans into Nike equity on the edge of collapse.

    Hiring for Negatives, Loss Bias, and Manufactured Discontent

    Ross was good at spotting smart, talented people but kept hiring ones who caused organizational problems, because he could always talk himself into a candidate. Watching a sharp head of HR screen people out, he realized he had been hiring wrong: growing talent means showing positives, but selecting talent means hunting for disqualifying negatives, since one bad trait spreads to the whole team. He formalized a versioned “people spec” with positives like return on luck and poetic design, each paired with a negative. He also hired for loss bias, the fact that a loss feels roughly six times more painful than an equal gain, seeking people who “book the win early.” That competitive, pressure-seeking wiring links to Michael Jordan manufacturing humiliation stakes (per Tim Grover in Relentless) and to David Ogilvy’s divine discontent. Ross’s own manufactured discontent today is the world’s shortage of compute, which he frames in life-and-death terms.

    The Optimistic Close: Free Code and Universal Software Literacy

    Ross ends on aggressive optimism. Software has long run on “code rationing” because code was expensive to write, policed by “no engineers” whose job is to say no. As the marginal cost of code approaches zero, the workflow flips to implement, experience, then re-implement. More important is accessibility: just as alphabets and universal education turned reading and writing from a scribe’s monopoly into a question of quality, AI is making software creation universal. His executive assistant now builds working apps, and a wave of individual founders with taste but no coding background will create valuable companies. The corollary for education is to stop teaching kids to answer questions and start teaching them to ask, revamping curricula around real community problems where the point is asking the right questions to get AI to solve something that matters.

    Notable Quotes

    “Success in the information age was about being able to answer questions. Success in the AI age will be about being able to ask the right questions.”

    Jonathan Ross, on the fundamental shift AI creates

    “The fewer constraints that you give someone, the more freedom they have to solve the problem, and the more freedom they have to surprise you with the solution.”

    Jonathan Ross, on leading creative teams

    “Being able to think faster makes you think smarter.”

    Jonathan Ross, on why faster inference produces more capable models

    “There are plenty of really smart people who wouldn’t recognize reality if it tapped them on the shoulder.”

    Jonathan Ross, defining reality quotient versus IQ

    “If you express intentional leadership, you say, ‘I intend to do this.’ People don’t tend to offer their opinion, but if it’s very wrong and there’s a reason, they will push back.”

    Jonathan Ross, on the lesson from Turn the Ship Around

    “When people are passengers in a car, they’re more nervous about a windy road or a scary road. But when they’re the driver, they feel more in control.”

    Jonathan Ross, on why Grok bonds kept the team together

    “The biggest flip in my hiring was when I went from looking for positives, which is what you do when you’re trying to grow talent, to looking for negatives, which is what you do when you’re trying to select talent.”

    Jonathan Ross, on inverting his approach to hiring

    “If it takes us an extra year to cure cancer because we don’t have enough compute, that’s my fault.”

    Jonathan Ross, on the discontent that drives him today

    Watch the full conversation between Jonathan Ross and David Senra here on YouTube.

    Related Reading

    • Groq the company Ross founded and the LPU behind the fast-inference story and the NVIDIA partnership.
    • AlphaGo versus Lee Sedol (Wikipedia) the match, including Move 37, that showed Ross how much faster hardware raises a model’s capability.
    • The Keynesian Beauty Contest (Wikipedia) the dynamic Ross uses to explain why West Coast VCs herded past what became NVIDIA’s biggest deal.
    • Zero to One by Peter Thiel, the source of the first-principles thinking Ross applied to the contrarian bet on fast inference.
    • Founders podcast by David Senra the host’s biography-driven show, source of the Jensen, Michael Jordan, and Edwin Land ideas referenced throughout.
  • Gavin Baker on Orbital Compute, TSMC, Frontier AI Models, Anthropic’s Vertical Take Off, and the Coming Wafer Shortage

    Gavin Baker, founder and CIO of Atreides Management, returns to Patrick O’Shaughnessy’s Invest Like the Best for his sixth appearance. He calls the current AI moment the most extraordinary moment in the history of capitalism, walks through what Anthropic’s vertical takeoff in revenue actually means, lays out why orbital compute is closer than skeptics believe, dissects the TSMC bottleneck that may be the only thing standing between today’s market and a full-on AI bubble, and rates every hyperscaler on how they have positioned for a world where frontier model providers may stop selling API access altogether.

    TLDW

    Anthropic added eleven billion dollars of ARR in a single month, which is roughly the combined business of Palantir, Snowflake, and Databricks built over a decade. That is the setup. From there Gavin Baker covers the March and April selloff, the contrarian read that a closed Strait of Hormuz was actually bullish for American manufacturing competitiveness, why Anthropic and OpenAI multiples may be misleadingly cheap on an unconstrained run rate basis, why Elon Musk’s discipline on SpaceX valuation created a superpower of permanent access to capital, the practical engineering case for orbital compute as racks in space rather than Pentagon sized space stations, why TSMC’s capacity discipline is the single most important variable in whether the AI cycle becomes a bubble, what Terafab in Texas changes, why the Pareto frontier of AI models has flipped from Google dominance to Anthropic and OpenAI dominance in nine months, the shift from all you can eat AI subscriptions to usage based pricing and what that means for revenue scaling, Richard Sutton’s bitter lesson as the largest risk to the AI trade, why frontier tokens still capture an overwhelming share of economic value, the role of continual learning as the third great open question, why most new chip startups should not try to build a better GPU, why Cerebras did something different and hard, why disaggregated inference may extend GPU useful lives to ten or fifteen years and rescue the private credit industry, why being in the token path is the new venture filter, the new prisoner’s dilemma around releasing frontier models via API, an honest rating of Google, Meta, Amazon, and Microsoft, why personal safety is becoming a real AI era risk, and why he remains an AI optimist maximalist who believes this could be the next Pax Americana.

    Key Takeaways

    • Anthropic added eleven billion dollars of ARR in one month, more than the combined businesses of Palantir, Snowflake, and Databricks built across a decade. There is no precedent for this in the history of capitalism.
    • The SaaS and cloud revolution created between five and ten trillion dollars of value over twenty years. AI is replaying that compression on a timeline measured in months.
    • The March selloff was a drawdown driven by disagreement with price action, not invalidated thesis. That is the kind of drawdown an investor can lean into.
    • Deep Seek Monday in January 2025 was a similar setup. By the day of the selloff, AWS Asia GPU prices had already doubled, GPU availability had fallen, and it was obvious reasoning models would be vastly more compute hungry at inference. The market priced the opposite.
    • The Strait of Hormuz closing was actually positive for America. US natural gas (the primary input into US electricity, which feeds AI) fell twenty percent on Bloomberg while Asian and European natural gas doubled or tripled. American manufacturing competitiveness improved overnight.
    • The US is now the world’s largest producer and exporter of oil and gas. The economy is dramatically less energy intensive than in the 1970s. The shortage trauma comparison does not hold.
    • Tech as a sector traded as cheaply versus the rest of the market in early April as at any point in the last ten years, into the single most bullish moment for AI fundamentals on record.
    • Anthropic is dramatically more capital efficient than OpenAI, having burned roughly eighty percent less to reach a similar revenue scale. They have very different structural returns on invested capital.
    • Anthropic at roughly nine hundred billion for fifty billion of ARR (growing a thousand percent) is striking. Adjusted for compute constraint, the unconstrained run rate could be one hundred fifty to two hundred billion, putting the implied multiple closer to five times.
    • Claude Opus generates roughly seventy percent fewer tokens for the same question than previously, with token quantity tied to answer quality. Subscribers on flat-fee plans are getting a lobotomized model.
    • Elon Musk’s superpower is twenty years of making investors money. He never pushes valuation. SpaceX compounded low thirty percent per year for a decade because Musk treats fair pricing as a sacred covenant.
    • Capitalism will solve the watts shortage. The current bottleneck has shifted from chips and energy to zoning and political approval. Many capex decisions are paused until after the US midterms.
    • The watts shortage probably begins to alleviate in 2027 and 2028. Orbital compute solves it longer term.
    • Orbital compute is not Pentagon sized data centers in space. It is racks in space. A Blackwell rack is three thousand pounds, eight feet tall, four feet deep, three feet wide. SpaceX has shown a satellite roughly that size.
    • The satellites operate in sun synchronous orbit so solar wings (around five hundred feet per side) always face the sun and the radiator on the dark side always points to deep space.
    • Starlink V3 satellites already run at around twenty kilowatts. A Blackwell rack runs at one hundred kilowatts. SpaceX engineers express genuine confidence they have already solved cooling and radiator design at these scales.
    • Racks in space are connected with lasers traveling through vacuum, the same lasers already on every Starlink. SpaceX operates the world’s largest satellite fleet and, via xAI Colossus, the world’s largest data center on Earth.
    • Inference will move to orbit. Training will stay on Earth for a long time. Terrestrial data centers remain valuable for the rest of an investor’s career.
    • The wafer bottleneck is structural and political. TSMC is essentially Taiwan’s GDP, water, and electricity. The leaders see themselves as inheritors of Morris Chang’s sacred legacy and they do not behave like a Western public company.
    • Jensen Huang has never had a contract with TSMC. The relationship is run on handshakes and the assumption that things will be fair over time.
    • If TSMC did everything Jensen wanted, Nvidia could be selling two to three trillion dollars of GPUs in 2026 and 2027. TSMC’s discipline is the single largest factor preventing a true AI bubble.
    • Historically, foundational technologies always get a bubble. Railroads, canals, the internet. The current AI buildout is overwhelmingly funded out of operating cash flow, GPUs are running at one hundred percent utilization, and that is fundamentally different from the year 2000 fiber overbuild.
    • If one of Intel or Samsung Foundry catches up at the leading node, the other will follow, and TSMC’s discipline collapses. Watch TSMC capacity decisions to predict a bubble.
    • Terafab, the SpaceX and Tesla joint venture to build the world’s largest fab in America, has a partnership with Intel that grants access to fifty years of institutional foundry knowledge. The A teams at ASML, KLA, Lam Research, and Applied Materials will follow Elon’s reputation in hardware engineering.
    • The hiring playbook for Terafab includes building Taiwan Town, Japan Town, and Korea Town next to the fab. Recruit the engineers and import their families, their restaurants, and their staff.
    • Frontier tokens still capture an overwhelming share of all economic value created at the model layer. This is surprising and is one of the three big open questions for AI investing.
    • The Pareto frontier of intelligence versus cost has flipped. Nine months ago Google’s TPU dominated every point on the frontier. Today Anthropic and OpenAI dominate, with Grok 4.3 on the frontier and Gemini 3.1 hanging on.
    • Google’s conservative TPU V8 design (partly an attempt to reduce dependence on Broadcom and Nvidia) is the leading explanation for the loss of per token cost leadership.
    • AI pricing is shifting from all you can eat to usage based, mirroring the cellular and long distance industries. Cellular stopped being a great growth industry when it went all you can eat. AI just made the opposite move.
    • OpenAI and Anthropic together could exceed two hundred billion in ARR this year if compute keeps coming online and frontier token pricing holds.
    • The two hundred fifty dollar a month consumer AI plan is no longer enough to evaluate frontier capability. Enterprise plans with usage based billing are required because rate limits are now severe.
    • The three biggest open questions for AI investors are: violation of the bitter lesson via ASI or human ingenuity, whether frontier tokens keep commanding their premium, and when continual learning arrives.
    • Today’s continual learning is crude reinforcement learning during mid training on verifiable tasks. True continual learning means weights updating dynamically, like a human who learns the first time they touch fire.
    • Trying to build a better GPU is a losing strategy. Jensen will copy any one to three percent share design. Startups should target one percent share, do something different, and make it hard enough that Nvidia cannot fast follow.
    • Disaggregated inference (separating prefill and decode) opens new design canvases. Prefill is memory capacity bound. Decode is memory bandwidth bound. Each can be optimized independently.
    • Cerebras did something different and hard with wafer scale computing. Three generations of chips and real grit to get there.
    • Disaggregation of inference may stretch GPU useful lives to ten or fifteen years, dropping financing costs from low sevens to five or six percent, mathematically lowering the cost of the AI buildout and likely saving the private credit industry from its SaaS loan exposure.
    • Sellers of shortage outperform buyers of shortage. But owning the largest installed base of what is currently in shortage (hyperscaler CPU fleets, for example) is also a strong position.
    • Most of the economic value at the application layer of AI has been destroyed, not created. The exceptions are companies in the token path or in niches small enough that frontier labs ignore them.
    • Coding may be the shortest path to ASI. If you can write code, you can write code that does anything. Cursor, Cognition, and Anthropic correctly focused on it.
    • Jensen could probably get close to the frontier with his own Nemotron family of models whenever he wants. The fact that he chooses not to is a strategic decision about not commoditizing his customers.
    • The new prisoner’s dilemma in AI is whether frontier labs release their best model via API. If everyone agrees not to, Chinese open source falls behind. If anyone defects, the defector pulls ahead on revenue and resources, forcing everyone else to defect.
    • Google still owns the largest compute installed base. Without TPU’s prior cost advantage, this matters more. YouTube data has real value in a world of robotics. GCP is going crazy.
    • Meta deserves credit for becoming AI first internally faster than any other internet giant. Musa, their first MSL model, is impressively close to the Pareto frontier.
    • Amazon is strong because of Trainium and robotics driven retail P&L efficiency. Nova is better than it gets credit for.
    • Microsoft flinched on capex in early 2025 and lost position. Satya Nadella’s current decision to use Microsoft compute for Microsoft products rather than reselling to OpenAI is a courageous and probably correct call, even at the cost of an eight hundred dollar stock price.
    • The hyperscalers most engaged with startups are Amazon and Nvidia by a mile, followed by Google. Broadcom is the favorite ASIC partner. AMD, Microsoft, and Meta have minimal startup engagement and that will cost them as the best teams are now at startups.
    • Personal safety in an AI era requires a family or company safe word that cannot be socially engineered. Deepfake voice and video extortion at the speed of FaceTime is already feasible.
    • Ukraine is winning largely on the back of having the best battlefield AI outside America and Israel. Adversaries are starting to internalize what AI dominance means geopolitically.
    • An optimistic read is that this becomes a new Pax Americana, the way the post 1945 American nuclear monopoly was used to rebuild Germany and Japan rather than dominate.
    • AI cured a friend’s daughter’s rare disease by spinning up a research effort that identified a market drug capable of impacting her condition. That is the upside that keeps Gavin an AI optimist maximalist.

    Detailed Summary

    The most extraordinary moment in the history of capitalism

    Gavin’s framing of the current moment is unusually direct. Anthropic added eleven billion dollars of annual recurring revenue in a single month. The three highest profile SaaS companies of the last decade plus, Palantir, Snowflake, and Databricks, took a decade and tens of thousands of employees collectively to build the combined business that Anthropic added in thirty days. He has been investing through every major tech cycle and says there is no historical analog. Not the dotcom era, not the cloud transition, not mobile. This is its own thing.

    The market response, then, was peculiar. The NASDAQ sold off into the single most bullish moment for AI fundamentals on record. Tech traded at roughly its widest discount versus the rest of the market in a decade. Investors who said they wished they had bought into AI during 2022, during COVID, or during Deep Seek Monday got the same valuation setup again in early April, this time with an even clearer inflection.

    Why the Strait of Hormuz closing was secretly bullish for America

    One reason the macro fear in March may have been mispriced is that the same geopolitical event that drove the selloff was, in practice, a relative benefit to the United States. American natural gas, the input into American electricity, which is the input into American AI training and inference, fell roughly twenty percent. Asian and European natural gas prices doubled or tripled. The US emerged with sharply improved relative manufacturing competitiveness, which is exactly what the current administration cares about.

    The 1970s comparison does not hold. The US economy is dramatically less energy intensive, it is now the world’s largest producer and largest exporter of oil and gas, and there are no shortages, only price moves. That backdrop made it easier for disciplined investors to stay focused on AI fundamentals through the volatility.

    Anthropic and OpenAI valuations on an unconstrained run rate

    Anthropic at roughly nine hundred billion for fifty billion of ARR sounds rich until you adjust for the fact that the company is severely compute constrained. Gavin estimates that, unconstrained, Anthropic might be at one hundred fifty to two hundred billion in run rate revenue, putting the implied multiple closer to five times. He also points out that Claude Opus now generates roughly seventy percent fewer tokens for the same question than it used to. Token quantity correlates with answer quality, and Anthropic is rate limiting and shrinking outputs to ration capacity across its user base.

    Anthropic and OpenAI are also structurally very different. Anthropic has burned around eighty percent less cash than OpenAI to reach a comparable revenue scale. That implies very different long term returns on invested capital, though OpenAI has done a better job locking in compute and Sarah Friar is one of the most exceptional CFOs Gavin has worked with.

    Why neither lab is raising at a three trillion dollar valuation

    The answer Gavin gives is that both labs are deliberately leaving valuation on the table the way Elon has done for two decades. SpaceX compounded at low thirty percent annually for a decade because Elon never pushed price. The result is a permanent superpower of access to capital. Investors trust him because they have made money with him for twenty years. That is a moat that compounds with every round.

    Anthropic could probably raise at a one hundred percent premium to its rumored latest mark. They are choosing not to. In an uncertain world (Ukraine, Russia, Iran, Taiwan), preserving the ability to raise more capital later at fair prices is more valuable than maximizing this round.

    Watts and wafers, the two real constraints

    Capitalism is solving the watts problem. The leading PE infrastructure investors now say zoning and political approval, not chips or energy, are the gating factors. Companies are deferring big capex announcements until after the US midterms. Turbine capacity is being doubled at the manufacturers. Companies like Boom Aerospace are repurposing jet engines for grid use. Watts probably ease meaningfully in 2027 and 2028 and then orbital compute does the rest.

    Wafers are the harder problem because they live in Taiwan, run on handshakes, and depend on a corporate culture that does not respond to public market incentives. TSMC is essentially the GDP, water consumption, and electricity consumption of Taiwan. Its leadership treats the company as the legacy of Morris Chang. The Silicon Shield doctrine is real and internal.

    Orbital compute as racks in space

    The biggest mental update Gavin asks listeners to make is to stop picturing data centers in space as Pentagon sized space stations. A Blackwell rack is three thousand pounds and roughly the size of a refrigerator. SpaceX has shown a concept satellite of about that size. Solar wings extend five hundred feet to each side and the radiator extends hundreds of feet behind, both possible because the orbit is sun synchronous and the orientation is fixed relative to the sun.

    SpaceX engineers Gavin has spoken to at Starbase express genuine confidence that they have solved cooling at these power levels. They have. Starlink V3 satellites already operate at twenty kilowatts. A Blackwell rack is one hundred kilowatts. The same company operates the world’s largest satellite fleet and the world’s largest data center on Earth via xAI Colossus. The racks are connected to each other with lasers traveling through vacuum, technology already deployed in every Starlink. The naysayers, Gavin observes, are armchair skeptics and Larry Ellison’s response (he is out there landing rockets, no one else is) is the right frame.

    Terafab in Texas and the threat to TSMC’s discipline

    Terafab, the SpaceX and Tesla joint venture, intends to be the largest fab in the world. The partnership with Intel grants access to fifty years of foundry institutional knowledge, allowing Terafab to start three to five quarters behind the leading node rather than fifteen years behind. The A teams at the semicap equipment companies (ASML, KLA, Lam Research, Applied Materials) will follow Elon’s reputation in hardware engineering the same way they followed TSMC twenty years ago when Intel stumbled.

    The talent strategy is the part most observers underestimate. Recruit the best engineers globally, then import their families, their restaurants, their staff. Build Taiwan Town, Japan Town, and Korea Town next to the fab. Optimize the human experience for the people whose work matters. Intel and Samsung do not think that way.

    Bubble watch and the year 2000 comparison

    Every foundational technology in modern history has had a bubble. Railroads, canals, the internet. Carlota Perez documented why. Markets correctly identify the importance, diversity of opinion collapses, supply gets ahead of demand, the bubble crashes. The current cycle has two important differences. The buildout is overwhelmingly funded out of operating cash flow, not debt. Every GPU is running at one hundred percent utilization, while at the peak of the fiber bubble ninety nine percent of fiber was unused.

    TSMC discipline is the single largest reason a bubble has not formed. If Jensen could buy everything TSMC could theoretically make, Nvidia could sell two to three trillion dollars of GPUs in 2026 and 2027. At some point that becomes more than the market can absorb. If Intel or Samsung Foundry catches up at the leading node, the other will too. TSMC’s pricing discipline collapses and the bubble starts.

    The Pareto frontier and the loss of Google’s cost advantage

    The most important chart in AI is the Pareto frontier of model intelligence versus per token cost. Nine months ago, Google’s TPU based models dominated every point on it. OpenAI, Anthropic, and xAI sat inside the frontier. Today the frontier is dominated by Anthropic and OpenAI, with Grok 4.3 on the frontier and Gemini 3.1 hanging on by subsidization more than economics. The most likely cause is Google’s conservative TPU V8 design, an attempt to reduce dependence on Broadcom and Nvidia that sacrificed per token economics.

    The bitter lesson, frontier tokens, and continual learning

    Three open questions dominate AI investing. The first is whether Richard Sutton’s bitter lesson (more compute beats human algorithmic cleverness) gets violated by ASI itself optimizing for efficiency. Closer observers of AI are more skeptical of a violation. Gavin thinks ASI’s first move will be to make itself more efficient and more resourced, which is technically a temporary violation.

    The second is whether frontier tokens keep capturing the overwhelming share of economic value at the model layer. Today they do, surprisingly. Gemini 3.1 Pro was mindblowing nine months ago and is intolerable today. The third is when continual learning arrives. Today’s models need a million fire touches to learn what a human learns from one. True continual learning would mean dynamic weight updates in real time and would produce a fast takeoff.

    From all you can eat to usage based AI pricing

    AI is shifting from flat fee plans to usage based pricing. The historical analogy is cellular and long distance. Both stopped being great growth industries when they went all you can eat. AI just made the opposite move. The consequence is that flat fee subscribers, even on premium consumer plans, get a rate limited and token throttled version of the frontier model. Enterprise plans with usage based billing are now required to evaluate true capability. Gavin thinks the combination of new compute coming online and usage based pricing is what gets OpenAI and Anthropic past two hundred billion in combined ARR this year.

    Chip startups, prefill decode disaggregation, and Cerebras

    Trying to build a better GPU is the wrong move. The four scaled players (Nvidia, AMD, Trainium, TPU) have copy capability for any one to three percent share design that looks attractive. The good news for startups is that disaggregated inference (separating prefill and decode) opens a richer design canvas. Prefill is memory capacity bound. Decode is memory bandwidth bound. Each can be optimized independently. Andrew Fox’s analogy is a British naval ship of the eighteenth century. Prefill is loading the cannon. Decode is firing it.

    Cerebras is the model. Wafer scale computing is genuinely different and genuinely hard. It took three generations of chips to get right. Andrew Feldman and his team had the grit to keep going through chip one being a failure. The design has a high ratio of on chip compute and memory relative to shoreline IO, which is why Cerebras is now experimenting with putting an optical wafer on top of the compute wafer to solve scale out.

    GPU useful lives and the rescue of private credit

    One of the strongest claims in the conversation is that disaggregated inference will stretch GPU useful lives to ten or fifteen years. The skeptical narrative (GPUs are obsolete in two years, companies are cooking their depreciation books) is wrong. You can put a Cerebras system or Groq LPU in front of older Hopper or Ampere parts, use them only for prefill, and run them until they physically melt. Private credit, which is in pain from SaaS loans and which underwrote GPU loans on three to four year lives, may be saved by this.

    If GPU financing rates can come down from low sevens to five or six percent, the mathematics of the AI buildout improves materially. That is a structural tailwind that compounds for years.

    The application layer, the token path, and a new prisoner’s dilemma

    Trillions of dollars of value have been destroyed at the application layer, not created. Cursor and Cognition are the rare scaled exceptions, and they got there by focusing on coding very early. As Amjad Masad noted, coding is plausibly the shortest path to ASI because a coding agent can write itself into any new domain. Jamin Ball’s frame is that the new venture filter is whether the company is in the token path. Data Bricks is. Most application layer startups are not.

    Jensen could probably get close to the frontier with Nemotron whenever he wants, and the strategic question of whether to do that is a new prisoner’s dilemma. If every frontier lab agrees not to release best models via API, Chinese open source falls steadily behind. If anyone defects, the defector gains revenue and resources, and everyone else has to defect. The same dynamic exists between TSMC, Intel, and Samsung. If Nvidia or AMD ever truly used an alternative foundry, that foundry would catch up rapidly.

    Rating the hyperscalers

    Google has the largest compute installed base, the YouTube data that matters in a robotics world, and a search business that prints. Their loss of TPU cost leadership is the surprise of the year. If Google IO in five days does not produce a leapfrog model, the Nvidia centric narrative gets even stronger.

    Meta deserves real credit. Zuckerberg made Meta AI first internally faster than any other internet giant, paid up for the talent contracts when no one else would, and shipped Musa as a first model from MSL that is close to the Pareto frontier. Amazon is well positioned on Trainium, robotics in retail, and a Nova model line that is better than it gets credit for. Microsoft flinched on capex in early 2025 and lost position. Satya Nadella’s current decision to use Microsoft compute for Copilot rather than reselling to OpenAI is courageous and probably correct, even at the cost of stock price.

    The most interesting cross hyperscaler metric is startup engagement. Nvidia and Amazon engage deeply with startups. Google is next. Broadcom is the favored ASIC partner. AMD, Microsoft, and Meta have minimal startup engagement, which Gavin believes will cost them as the best teams now sit at startups.

    Personal safety, geopolitics, and the Pax Americana case

    The closing section turns darker. Personal safety in an AI era requires a family or company safe word that cannot be socially engineered. Deepfake voice and video extortion via something that looks exactly like your child calling on FaceTime is already feasible. Political violence against AI leaders is a real concern. Geopolitically, Ukraine is winning largely because it has the best battlefield AI outside America and Israel. How adversaries respond to that asymmetry is the next great variable.

    Gavin’s optimistic frame is the Pax Americana. After 1945 the US had a nuclear monopoly and could have controlled the world. Instead it rebuilt Germany and Japan, both of which became the most reliable American allies for the next eighty years. If AI dominance plays out similarly, this is a generationally positive story rather than a destabilizing one. The personal anecdote that closes the conversation is a friend whose daughter was diagnosed with a rare genetic condition. He spun up agents, identified a drug already on the market that addresses her mutation, and her life is immeasurably different because of AI. That is the upside.

    Thoughts

    The Anthropic eleven billion in a month framing is the kind of stat that resets priors. The right way to interpret it is not as a one off but as a measure of how fast value can compound when the underlying technology improves on a curve steeper than the ability of the rest of the economy to absorb it. The skeptical question is whether that ARR is durable or whether it is heavily tied to a customer base of other AI companies that are themselves on a single venture funded year of runway. The bullish answer is that frontier coding, frontier research, and frontier enterprise tasks are not going to stop being valuable, and Anthropic is the best at all three. Both can be true. The number is still extraordinary.

    The argument that TSMC discipline is the only thing preventing a bubble is the analytically tightest part of the conversation. The implied trade is to watch TSMC capacity additions like a hawk and to be more, not less, cautious if Intel Foundry or Samsung Foundry ever announce real share at the leading node. The Terafab thesis is more speculative but more interesting. If Elon’s talent recruiting playbook works and the Intel partnership gives Terafab a real seat at the table within five years, the geometry of the global semiconductor industry shifts in a way that is bullish for American manufacturing, bullish for power and water infrastructure in Texas, and ambiguous for TSMC itself.

    The Pareto frontier discussion deserves more attention than it usually gets. Pricing leadership in AI is not a vanity metric. It determines who can subsidize free tier usage, who can absorb compute shortages, who can ship cheaper enterprise plans, and ultimately whose model becomes the default for any given workload. Google losing per token leadership in nine months is one of the most under analyzed events in the sector and it explains a lot about why Anthropic and OpenAI are growing the way they are. If Google IO does not produce a leapfrog model, the implied verdict on TPU V8 design choices gets a lot harsher.

    The application layer destruction point is worth sitting with. Founders building on top of frontier models are competing in a world where the model itself moves faster than any moat they can build, where the model lab can absorb their niche if it gets interesting, and where the only protection is either deep token path integration or a niche so small the lab does not bother. That is a much harsher venture environment than the early SaaS era. The compensating opportunity is that one human can now run a hundred agents, so the ceiling on what a small team can build is correspondingly higher. The bet is that productivity per founder rises faster than competitive pressure from the labs. We will find out.

    The orbital compute pitch is the section that will polarize listeners. The naive read is that this is science fiction. The closer read is that every component (sun synchronous orbit, laser interconnect, twenty kilowatt satellite buses, ten thousand satellite manufacturing cadence, full rocket reusability) already exists. The remaining engineering problems are repair, maintenance, and radiator scale, all of which are real but tractable on a five to ten year horizon. The strategic implication is that the political and zoning ceiling on terrestrial data centers becomes less binding if orbital compute is a credible alternative for inference workloads. The investor implication is that being short the watts and cooling complex on a five year horizon is a real trade, not a meme.

    Watch the full conversation here.