PJFP.com

Pursuit of Joy, Fulfillment, and Purpose

  • OpenAI’s Defender’s Window Cyber Security Keynote: GPT-6 Astra, Daybreak Red and Blue, Codex Security Red, the $1 Billion Defense Fund, Patch the Planet, and How OpenAI Built Its Internal Defense Factory

    In The Defender’s Window, a cyber security keynote held in London, OpenAI’s cyber team makes one argument from several angles. Frontier models like GPT-6 Astra now find and exploit real vulnerabilities, open weights models are not far behind, and defenders have a short head start to use that capability before attackers can. Over about 47 minutes, OpenAI’s EMEA general manager, head of cyber engineering Matt, the EMEA cyber go-to-market lead, deployment engineer Vanessa and field CTO Lou cover the models, the safeguards, a live Codex Security demo, and the internal “defense factory” OpenAI built to secure itself.

    TLDW

    OpenAI says the “defender’s window” is the gap between its frontier cyber models and the open weights models closing in behind them, and that defenders need to use it now. The keynote covers GPT-6 Astra, OpenAI’s most capable and, it says, most aligned model. Astra completes about 40% more Exploit Gym challenges than GPT-5.6 Sol while using far fewer tokens. It is OpenAI’s first model to reach the “cyber critical” threshold, and it exploited a forbidden out-of-scope shortcut in zero cases, against 48% for Sol without production safeguards. Access comes in two Daybreak tiers. Red is for approved offensive teams and gets a new managed pentesting product, Codex Security Red, with sandboxed agents and a guardian agent checking outbound traffic. Blue is for everyday defensive work. There is also a $1 billion fund that subsidizes access for critical infrastructure, nonprofits and open source. Real-world results include a 23-year-old OpenBSD flaw, MikroTik bugs going back to 2013, a two-bug Chrome JavaScript engine exploit chain, and 37 patches merged in the first week of Patch the Planet with Trail of Bits. Vanessa demos Codex Security on the Ladybird browser: a context-driven full scan, a validated finding, a generated patch, a security.md file, and sub-agents that open Jira tickets and GitHub PRs, then a CLI for bulk scans across thousands of repos and in CI. Lou walks through OpenAI’s internal “code red” defense factory: 250 people, agents running in VMs with dev containers, a false positive rate under 1%, 90% correct ownership assignment, and a fix rollback rate under 1%. Matt closes with a plea not to wait for the productized version.

    Thoughts

    The “defender’s window” is a sales pitch, but a useful one. OpenAI is telling a room of European security leaders that its closed frontier models are ahead of open weights models on offensive cyber capability, and that the lead is temporary. Both halves of that sentence do work. The first justifies gated access through Daybreak and a $1 billion subsidy. The second turns a procurement decision into a race against the clock. The framing is self-serving, and it may also be correct. If a model can take a crash and turn it into a working exploit today, a free model will probably do the same within a year. Then any organization that has not already worked through its backlog is facing attackers who have the same tools it does. The honest takeaway is that the window matters regardless of which vendor you use to exploit it.

    The safety section (around the 18 to 20 minute mark) is more interesting than most keynote safety slides because it gives a number. OpenAI gave the model a hard task with an out-of-scope target it could exploit as a shortcut. GPT-5.6 Sol took the shortcut in about 48% of runs without production safeguards, and Astra took it in zero. That is the property that actually matters when you hand an agent Daybreak Red access and a network path to your production application. Codex Security Red’s architecture takes the same idea further by not trusting the model alone: each agent gets its own sandbox, and a guardian agent reviews outbound traffic. Pairing a model trained to respect scope with infrastructure that enforces scope anyway is the right design. Security teams should ask every AI pentest vendor for both.

    Vanessa’s demo (roughly minutes 26 to 30) quietly makes the most practical point of the day. The difference between a reasoning model and a traditional static analyzer is that you can tell the model what matters to you. Compensating controls, business logic, how your company defines fraud, where your trust boundaries sit: all of that can be seeded into a scan and then committed as a security.md file, which she describes as “an agents.md but for security.” That turns your threat model into a versioned artifact that lives next to the code and shapes every future scan. Many organizations have never written their threat model down anywhere. The fact that writing one now pays off directly in scan quality may be the best reason they will ever get to do it.

    The metrics in Lou’s section (minutes 39 to 41) are the part to take most seriously, and most skeptically. A false positive rate under 1% after dynamic validation, against the 50 to 60% Lou says is typical, would change how security teams work, because false positives are what make developers ignore security tickets. But the number that says the most is ownership assignment at about 90%. Finding and even fixing a bug is now the easy part. Working out which team owns a piece of code, from Slack, PagerDuty and the source tree, is the part that still fails one time in ten, even at OpenAI with 250 people on it. That is an organizational problem that no model fixes. It is also a reminder that these figures describe OpenAI’s own unusually well-resourced estate, not the typical enterprise.

    The closing instruction, “don’t wait for us,” is the most honest line in the keynote. Lou has just described a productized defense factory that only a small handful of customers can get, so Matt’s point is that the pieces already exist: the models, the Codex Security plugin and CLI, and dev containers for reproducible agent environments. The architecture Lou describes is not proprietary magic. It is a VM for kernel-level isolation and monitoring, a dev container that packages tools and security hooks (a CrowdStrike Falcon sensor, Wiz), and an agent with skills specific to your organization. A competent platform team could build a first version in weeks. The organizations that come out of this window in good shape will be the ones that treated it as an engineering project, not a vendor rollout.

    Key Takeaways

    • OpenAI defines the “defender’s window” as the gap between its frontier cyber capabilities and the fast-improving open weights models behind them. The window is open now and will not stay open.
    • OpenAI’s EMEA general manager says cyber security went from absent to the opening topic in customer conversations within months. Nearly every meeting in the past two weeks started with it.
    • Adoption figures cited include more than 1 billion weekly users, and a German SMB, Stadler, running 145 agents alongside 650 employees for a 30 to 40% efficiency gain.
    • OpenAI says it grew its safety team, invested more compute in safety, and paused its most recent training runs for a couple of weeks in early August.
    • An open letter calling for a collective, ecosystem-wide response to AI-era cyber risk was signed by hundreds of organizations. Matt puts the count at 500.
    • OpenAI frames defense as an ecosystem play: company leadership making cyber the top priority, technology partners like Darktrace and Check Point, governments as accelerators, and OpenAI’s models.
    • OpenAI announced work with Ukraine on cyber defense, citing 6,000 attacks in the last year. ENISA, the EU cyber agency, reported finding unexpected vulnerabilities with OpenAI’s models.
    • A $1 billion fund subsidizes Daybreak access so hospitals, utilities, nonprofits and open source maintainers are not priced out of defense.
    • Matt opens with the Thames Barrier as a metaphor: infrastructure that protects a city for decades because someone takes responsibility for it.
    • OpenAI’s models helped find a 23-year-old flaw in OpenBSD and MikroTik vulnerabilities affecting releases back to 2013.
    • Findings are not the goal. Every finding triggers validation, deduplication and risk-acceptance checks, and only verified fixes actually protect anyone.
    • A government Cyber Shield initiative aims to use agents to find vulnerabilities and move toward automated remediation across critical infrastructure, in effect a national defense factory. Matt says the UK government is also pursuing this.
    • OpenAI’s position is that every organization needs its own defense factory, and that “their security is our security” for under-resourced defenders.
    • GPT-6 Astra is OpenAI’s most intelligent model, strong at software engineering and computer use, and trained specifically to find more vulnerabilities, including zero days.
    • On Exploit Gym, which only counts a working exploit (a crash is not enough), GPT-5.6 Sol completed about 30% of challenges. Astra completed about 40% more with far fewer output tokens.
    • OpenAI’s models, working with its researchers, found two unknown bugs in Chrome’s JavaScript engine and chained them into an exploit. Google shipped a fix.
    • Patch the Planet, run with Trail of Bits, funds researchers to use OpenAI’s models on open source projects like Python, curl and Go. It merged 37 patches in its first week, and the aiohttp maintainers fixed eight reported issues within hours.
    • Astra is OpenAI’s first model to reach the “cyber critical” threshold.
    • Safeguards are layered: refusal training, abuse detection and blocking, tighter limits on high-risk accounts, and monitors that watch the model’s reasoning and actions.
    • In a scope test with an exploitable out-of-scope shortcut, Sol without production safeguards took the shortcut about 48% of the time, and Astra never did.
    • Cyber capability is built by adding specialist training (vulnerability understanding, exploit development, result verification) on top of the frontier coding and reasoning models.
    • Daybreak Red is the highest access tier, for approved teams doing advanced red teaming and weaponized exploit development, such as turning crashes into exploits or chaining weaknesses.
    • Codex Security Red is a new managed penetration testing product. You define scope and rules of engagement in Codex, agents run in OpenAI-hosted isolated sandboxes, and a guardian agent reviews their outbound traffic.
    • Daybreak Blue is the starting point for most defenders: general-purpose frontier models (GPT-6 Sol, Luna, and soon Astra) with safeguards tuned for authorized code review, alert triage, vulnerability finding and patching.
    • The Codex Security demo runs a full scan on the Ladybird browser, reviewing about 28,000 files to build a threat model, surface candidate vulnerabilities, and validate them.
    • For a first scan, Vanessa recommends a full-codebase scan at high or extra-high reasoning on the most capable frontier model available.
    • Context is the differentiator. Attack vectors, areas of focus, compensating controls, business logic and your definition of fraud can all be seeded into a scan.
    • A security.md file works like an AGENTS.md for security. It gives repo-level context on threats, trust boundaries and priorities to every future scan.
    • A fix-finding skill generates a patch that can be applied locally and then verified to confirm the vulnerability is gone.
    • Sub-agents can patch findings in parallel, group them, open Jira tickets and draft GitHub PRs, and notify the right engineers on Slack, so the person who scans and the person who fixes can be different people.
    • At scale, the Codex Security CLI and SDK run bulk scans with configurable workers across a CSV of repositories, and plug into CI to catch vulnerable dependencies and fail checks on risky PRs. Vanessa says the tooling is open source.
    • OpenAI called an internal “code red” and pulled 250 people from engineering, security and research into a cross-functional team to build its first defense factory.
    • The defense factory loop is inventory, discovery, dynamic validation, ownership assignment, and verified remediation.
    • Existing security tools stay in place. What is new is agents, organization-specific skills, and isolated environments where agents can run the application and iterate.
    • The reference architecture runs inside your network, puts each agent in a VM for isolation and kernel-level monitoring, and uses a dev container for reproducible dependencies and security hooks such as a Falcon sensor or Wiz.
    • OpenAI’s internal results: a false positive rate under 1% after dynamic validation (Lou says 50 to 60% is common elsewhere), about 90% correct ownership assignment, and a fix rollback rate under 1%.
    • OpenAI is productizing the defense factory with a small group of early customers, with wider rollout over the coming weeks and months and more announcements expected around DevDay.
    • Matt’s practical next steps: apply for Daybreak, start with the Codex Security plugin on one repo, scale with the CLI and SDK in CI, and adopt dev containers for reproducible agent environments.
    • OpenAI plans to extend the work through a Daybreak Defense Network of training and implementation partners, and to move beyond prevention into investigation, response and security operations.

    Detailed Summary

    Opening: Trust, Safety and Why Everyone Showed Up

    OpenAI’s general manager for Europe, the Middle East and Africa opens by placing the event in a moment of fast AI progress. They cite the recent solution of a Navier-Stokes mathematics problem, the release of GPT-6 Astra, more than a billion weekly users, and enterprise examples like Stadler’s 145 agents. With that capability comes the question of trust. OpenAI says it has expanded its safety team, invested in safety compute, and paused its latest training runs in early August. The speaker notes that six months ago no customer led with cyber, and now every one does. That is why 150 invitations produced 150 attendees. The defender’s window is framed as an ecosystem problem needing company leaders, technology partners, governments and OpenAI together, with Ukraine, ENISA and a $1 billion fund as early examples.

    The Defender’s Window and the Case for Fixes Over Findings

    Matt, OpenAI’s head of engineering for cyber, opens with the Thames Barrier: a piece of infrastructure that has protected London for more than 40 years because someone owns it. This summer OpenAI saw its frontier models finding weaknesses that had gone unnoticed for decades while broadly available models caught up, so it pulled its security, applied engineering and research teams together to harden its own systems. That became the defense factory. Matt’s central point is that findings are cheap and fixes are what count. Every finding kicks off validation, deduplication and risk review, and the capabilities that surface bugs also make them easier to exploit. He points to the Cyber Shield initiative and the UK government as national versions of the idea, and presents Daybreak, Codex Security and the $1 billion fund as the pieces that make it possible for organizations without OpenAI’s resources.

    GPT-6 Astra: Capability, Efficiency and Real-World Results

    OpenAI’s EMEA cyber go-to-market lead presents GPT-6 Astra as a big step for everyday engineering and computer use, and a bigger one for cyber: more vulnerabilities found, zero days included, and stronger red teaming. The key evaluation is Exploit Gym, which requires a working exploit against targets such as Chrome’s JavaScript engine and the Linux kernel. Astra completes about 40% more challenges than GPT-5.6 Sol’s roughly 30% while using far fewer output tokens, which lets defenders run more thorough testing on the same budget. Real results include two chained bugs in Chrome’s JavaScript engine, reported to Google and fixed, and Patch the Planet with Trail of Bits, which funds researchers to harden Python, curl, Go and other widely used open source projects. That program merged 37 patches in its first week.

    Safeguards, the Cyber Critical Threshold and Alignment

    Astra is OpenAI’s first model to reach the cyber critical threshold, so the safeguards were strengthened to match. They include refusal training, abuse detection, tighter controls on high-risk accounts, better detection of attempts to bypass the safeguards, and monitors that watch the model’s reasoning and actions for moments when it strays from instructions. OpenAI reiterates that it paused frontier training to focus on monitoring and alignment, and that safety thresholds must be met before capability is pushed further. On alignment, the scope test gives the model a hard task with an out-of-scope target it could exploit as a shortcut. Sol without production safeguards did so in about 48% of cases. Astra did so in none.

    Daybreak Red, Daybreak Blue and Codex Security Red

    Cyber models are built by adding specialist training on vulnerabilities, exploit development and result verification on top of the frontier coding and reasoning models, then offered in two tiers. Daybreak Red gives approved teams specialist models for advanced red teaming and weaponized exploit development. Because giving an agent that much capability requires clear boundaries, OpenAI announced Codex Security Red, a managed penetration testing service. Teams define scope and rules of engagement in Codex. Investigation agents run in isolated, OpenAI-hosted sandboxes, their requests to your application pass through network controls and a guardian agent, and findings and evidence collect in a dashboard. Daybreak Blue is the default for most defenders, offering GPT-6 Sol, Luna and soon Astra with safeguards tuned for code review, alert triage and patch creation.

    Live Demo: Codex Security on the Ladybird Browser

    Vanessa, a cyber deployment engineer, runs Codex Security in the Codex desktop app against a local copy of Ladybird, the pre-alpha open source browser, which is now also part of Patch the Planet. She picks a full-codebase scan over a diff scan for the first run, sets high or extra-high reasoning, and stresses the context fields: attack vectors, focus areas, and enterprise context such as compensating controls, business logic and definitions of fraud. The scan builds a threat model from about 28,000 files, surfaces candidates, validates them, and reports a shared JavaScript bytecode cache vulnerability with its attack path, impact and likelihood. A fix-finding skill generates a patch, which is applied and verified. A security.md file is then generated and committed, so future scans inherit the repo’s security context.

    Scaling Remediation: Sub-Agents, Tickets, PRs and the CLI

    To go beyond one repo, Vanessa asks Codex to take the scan results and spawn sub-agents. Each one patches and validates a finding, groups related findings, opens Jira tickets and draft GitHub PRs, and notifies the right engineer on Slack. The PRs can be tuned to your CI pipeline so checks stay green before a human merges. The question she hears most from roughly a thousand customer meetings is scale: seven vulnerabilities in one repo is easy, 10,000 repos is not. The answer is the Codex Security CLI and SDK, which package the plugin’s composable skills for bulk scans with configurable workers across a repositories CSV, and for CI checks on dependency upgrades and new PRs.

    Inside OpenAI’s Defense Factory

    Lou, the field CTO for cyber, describes the internal code red: 250 people across engineering, security and research building a first version of continuous defense. The loop runs from inventory of the attack surface, to discovery with cyber models, to dynamic validation (duplicate? already tracked? reachable?), to ownership assignment through Slack, PagerDuty and source control, and finally to verified remediation. Existing tools stay in place. What is new is agents, organization-specific skills, and isolated environments where agents run the application itself. Each agent runs inside your network in a virtual machine, chosen over containers for stronger isolation and kernel-level monitoring of tool calls. On top of the VM sits a dev container built on Microsoft’s open specification for reproducible dependencies and security hooks. Results: under 1% false positives, about 90% correct ownership, and under 1% fix rollbacks. OpenAI is productizing this with a few early customers ahead of DevDay and further cyber announcements.

    Closing: Don’t Wait

    Matt returns to tie the pieces together: models provide capability, Codex and development tools put it to work, and security workflows give a starting point. His main lesson is not to wait for OpenAI’s product. He gives four steps: apply for Daybreak, start with the Codex Security plugin on one repository, scale through CI with the CLI and SDK, and package agent environments in dev containers for reproducible testing and clear ownership. OpenAI plans to go beyond prevention into investigation, response and security operations, and to bring in training and implementation partners through the Daybreak Defense Network. He closes on the hope that the industry will look back on this window as the moment it came together and made the world more secure.

    Notable Quotes

    “Six months ago the messages were not around cyber. Six days ago everybody started the conversation with cyber security concerns. Everybody.”

    OpenAI’s EMEA general manager, on how quickly customer priorities shifted

    “Fixes are what we want not findings in order to protect our institutions.”

    Matt, head of engineering for cyber at OpenAI, on the real bottleneck in vulnerability management

    “We can’t assume we’re going to stay ahead, but we do know we’re ahead for now.”

    Matt, on the temporary nature of the defender’s advantage

    “The defenders window is open. It’s not going to stay open for long, but it’s open right now, and we need to make it count.”

    Matt, framing the thesis of the keynote

    “As Sam has said previously, we shouldn’t be taking risks on behalf of humanity. People need to remain in control.”

    OpenAI’s EMEA cyber go-to-market lead, introducing the safeguards on GPT-6 Astra

    “We’re not interested in just having a raw dump of vulnerabilities. Our goal is to ultimately remediate and burn down risk.”

    Vanessa, cyber deployment engineer, during the Codex Security demo

    “It’s very very easy to do this for one repo when you have seven vulnerabilities but how do you do this when you have 10,000 repos?”

    Vanessa, introducing the Codex Security CLI and SDK

    “Across our dynamic validation we managed to get our false positive rate down below 1%. Which for a lot of organizations this number is vastly higher.”

    Lou, field CTO for cyber, reporting results from OpenAI’s internal defense factory

    “Don’t wait for us. All the pieces are here. We don’t have time.”

    Matt, urging security teams to build their own defense factory now

    Watch the full Defender’s Window keynote here.

    Related Reading

  • Higgsfield CEO Alex Mashrabov on $1B ARR in 18 Months, Burning $4M a Month on AI Models, 80% Margins on Open Weights, and Why Moats in AI Are BS (20VC)

    Higgsfield went from $1 million to $1 billion in annualized revenue in 18 months, faster than Cursor, and almost nobody outside the AI world has heard the story. In this 20VC interview, Harry Stebbings sits down with Higgsfield CEO Alex Mashrabov, a former top-three competitive programmer from Central Asia who sold his first company to Snap, to talk about the near-death pivot, how Higgsfield counts revenue, why the company spends over $4 million a month on models internally, the economics of open versus closed models, what a moat even means in AI, and what it costs him personally to run it.

    TLDW

    Alex Mashrabov grew up pushed toward competitive programming by parents who told him the United States was where technology mattered, sold AI Factory to Snap for $166 million, and led generative AI there before co-founding Higgsfield. The company burned over $10 million of a $16 million seed chasing hype before talking to eight creative directors, finding that camera control was missing from AI video, and launching into immediate product market fit with no paid marketing. Higgsfield now calculates ARR as the last four weeks of live, prorated revenue times 13, gets slightly over half its revenue from businesses, sees about 30% churn in month one but net revenue retention over 300% at month 12, and has one customer who went from a $99 subscription to a $6 million deal. Alex argues Google and OpenAI will crush $20 a month consumer subscriptions, that AI benchmarks are gamed and do not reflect real video workflows, that open weights models give Higgsfield 80%+ margins versus 20 to 30% on closed models, and that routing traffic between models (“tokenomics”) is a core feature. Internally, the roughly 400-person team spends over $4 million a month on models, more than $10,000 per person. He sees only two real moats (delivering outcomes and network effects), talks about raising from Yuri Milner, hiring in Kazakhstan, Europe’s strengths, working 80 to 90 hours a week, the lesson from Snap that momentum does not last, and a personal target of over $10 billion in revenue within 12 months.

    Thoughts

    The most useful part of the founding story is the admission about the pivot. Higgsfield spent more than a year and over $10 million of a $16 million seed optimizing for “what’s hype today, what’s the right narrative,” and Alex takes the blame for it directly. What saved the company was not a new narrative, it was eight conversations with creative directors who all said the same thing: AI video had no camera control, and you cannot tell a story without it. That is a very specific, very unglamorous insight, and it took Higgsfield from roughly $1 million to $20 million in ARR in about three months. The lesson for any founder with a shrinking runway is that the last attempt should come from customers, not from the timeline.

    The revenue section is worth reading closely because it is so unusual for a company at this stage to be this specific. Higgsfield counts ARR as the last four weeks of revenue times 13, prorates annual contracts, and excludes future contract value. Month-one churn is around 30%, which Alex openly says is below the old B2B SaaS bar, but month-12 net revenue retention above 300% is something SaaS basically never saw. Put those together with the $99 subscriber who became a $6 million a year customer, and his claim that Google and OpenAI will demolish the $20 a month consumer market, and you get a clear strategy: consumer signups are a funnel, and the business is expansion into companies that make hundreds or thousands of ads a week.

    The model economics in the middle of the interview are the most concrete numbers on AI app margins I have heard in a while. Open weights and in-house models run at over 80% gross margin for Higgsfield. Closed frontier models run at 20 to 30%. And because agentic ad workflows let Higgsfield choose the model in over 40% of cases, routing becomes a margin lever, which he calls “tokenomics.” This connects to his sharp critique of benchmarks: video benchmarks test text-to-video, while real production uses 3,000-word prompts and ten or more image references per scene, closer to driving a rendering engine than writing a sentence. If the real workload does not need PhD-level intelligence to make a viral ad, the cheapest model that does the job wins, and the company that owns the routing keeps the spread.

    On moats, Alex is more careful than Harry, who thinks they are mostly nonsense. Alex names two that still hold: delivering an outcome (helping businesses sell more through AI ads) and network effects, which “AI does not replace.” The system-of-record point is the less obvious part. Assets are scattered across Dropbox, Google Drive, Miro and Frame.io, and marketers want to search them semantically and check them against brand guidelines, something the pixel-first tools from Adobe and Canva were not built for. A harness that learns a brand’s visual style over time, plus a community that went from about 10 seeded open source projects to over 10,000 in eight weeks, is a more believable defensibility story than any single model. It is also a reminder that “wrapper” is not an insult if the wrapper becomes the place where the work lives.

    The last 15 minutes carry the most interesting tension. Alex’s lesson from Snap is that “the momentum doesn’t last forever,” and Snap is now under $15 billion in market cap, partly because it never told a convincing AI story. His finance team projects $4.5 billion in revenue in 12 months with deceleration built in, while he personally says over $10 billion. He also predicts that most social media content will be AI generated, but that authentic content like Harry’s will command 10 to 15x higher CPMs, and that the job of the future is a creative director talking to a computer and generating stories in real time, where taste matters most. The honest version of his outlook is that he knows the curve will flatten and is trying to capture as much as possible before it does, while working 80 to 90 hours a week and admitting he has had one full day with his son in three months. That is not a lifestyle anyone should copy, but it is a candid look at what hypergrowth actually costs.

    Key Takeaways

    • Higgsfield crossed $1 billion in annualized revenue 18 months after hitting $1 million. Cursor took 24 months, and Alex believes only OpenAI and Anthropic did it faster.
    • Alex’s father is from Uzbekistan and both parents are mechanical engineering professors. From age eight they told him he had to get to the United States because that is where technology matters.
    • His mother worked three jobs so he could compete in programming and attend training camps. By 19 he was top three in the world in competitive programming, and he was also top three in the world at checkers.
    • In 2014 he worked on pre-transformer neural nets for English to Russian translation. Many teammates were hired by DeepMind and Meta, but he was drawn to a future of AI-generated video on phones.
    • With co-founder Mahi de Silva he built AI Factory and sold it to Snap for $166 million, after heavy dilution in an era when AI multiples were near zero and a $12 million round was a big deal.
    • At Snap, his team’s face filters ran on-device, which made them nearly free for Snapchat, drove most daily new users, and reached hundreds of millions of people.
    • The original Higgsfield insight: most companies cannot keep up with the pace of social media content production, because trends change almost every day.
    • By 2023 it was clear scaling laws worked, and Alex bet they would work for video too, just two to three years behind language models and coding.
    • Higgsfield spent over a year searching for product market fit and burned more than $10 million of its $16 million seed. Alex blames himself for chasing hype and narrative instead of product.
    • With under $5 million left, the team talked to eight creative directors. All of them said AI video lacked camera control. Higgsfield launched on March 31 of the previous year and hit immediate product market fit.
    • Higgsfield does no paid advertising. It had problems after outsourcing influencer work to an agency, and the takeaway was to own distribution.
    • Higgsfield partnered with a major streamer on an AI-generated replica stream. Alex expects digital replicas to become a normal way for creators to monetize given how much pressure they are under.
    • ARR methodology: revenue over the last four weeks times 13, only live revenue, with annual contracts prorated to 28 days and no multi-year deal value included.
    • Video AI is about two years behind coding in adoption, and its share of on-demand usage revenue is well under the 50%+ seen at leading coding companies.
    • One customer went from a $99 a month subscription to a deal worth over $6 million a year in six months.
    • Big demand drivers come from Asia: direct-to-consumer brands rebuilding go-to-market to produce hundreds or thousands of ads a week, and the $10 billion+ short-form drama industry, where most new shows are made end to end with AI.
    • Business revenue is slightly over 50%. Pure consumer use is about 10%, roughly matching mobile’s share of revenue. The rest is aspiring creators and freelancers who are churny but tend to return within a year.
    • Alex believes Google and OpenAI will demolish the $20 a month consumer subscription market, so Higgsfield focuses on upgrading users to spend over $1,000 a year.
    • Consumer retention drops about 30% in month one and then stays flat. Business net revenue retention at month 12 is over 300%.
    • Higgsfield has over 150 in-house creative professionals, nearly half of the company, making launch videos, tutorials and an open-sourced AI-generated movie.
    • That movie needed over 100 hours of generated footage for 90 minutes of TV-quality content, so creative selection still matters a lot.
    • Alex calls chasing benchmarks a mistake and says researchers at large labs game them by leaking test data into training and using LLM-as-a-judge tricks to hit bonus targets.
    • Video benchmarks mostly test text-to-video, but real workflows use prompts averaging over 3,000 words and at least ten image references per scene. Video models are best thought of as modern rendering engines, like Unreal or Unity with different inputs.
    • VFX and camera control took Higgsfield from about $1 million to $20 million ARR in three months. Its own image model for aesthetic photo shoots and product consistency took it from $20 million to $100 million.
    • Most companies that say they build their own models post-train open weights models. The most valuable post-training uses customer decision sequences to teach models to compress ten steps into one.
    • Alex cites OpenRouter data showing the open source share of usage rose from under 30% to over 60% in a year, but expects OpenAI and Anthropic to keep over 50% of the market in dollars, driven by coding.
    • Gross margin on own and open weights models is over 80%. On closed models it is 20 to 30%. Higgsfield chooses the model in over 40% of cases, which it calls tokenomics.
    • Internal model spend is over $4 million a month across close to 400 people, more than $10,000 per person. One creative spent over $30,000 in a week vibe coding an asset workflow tool.
    • Alex expects top “10x” engineers and creatives to spend $50,000 to $100,000 a month on models, and to ask for matching raises.
    • Legal and customer support have not been replaced. Higgsfield has over 10 people in legal and over 40 in customer success, all heavy AI users. AI handles over 60% of first-line consumer support but does not work well for B2B.
    • High product velocity makes AI support harder: agents are only as good as their context and rules, and those change twice a week.
    • Engineers moved to Claude between March and June, then largely to Codex from mid June. Alex thinks tool preference is cyclical and model release velocity will not slow down.
    • Whoever builds the AI-native system of record wins. For Higgsfield that means semantic search of assets, brand guideline enforcement, and a harness that learns visual style over time.
    • Alex sees two real moats: delivering outcomes and network effects. Higgsfield’s open source community projects grew from about 10 to over 10,000 in eight weeks.
    • Yuri Milner was his best VC meeting. Alex has also had investors shake hands on a price and then try to syndicate the round at a 30% lower valuation the next day.
    • Excluding pharma and big tech, public companies spend more on sales and marketing than on R&D, and Alex expects much of that to become personalized video.
    • The West is over 70% of revenue, the US is the largest country, and Seoul is the largest city by usage. Higgsfield does not operate in China.
    • About 50 employees are in California, around 50 are remote, and over 300 are in Kazakhstan, which he says is top five in the world in physics olympiads and blends Soviet math with Singaporean education.
    • His management philosophy: hire the best people, empower them, retain them. He learns from Jensen Huang, Elon Musk and Nik Storonsky, who reject much of standard corporate management.
    • He works 80 to 90 hours a week, aims for at least 3 hours with his wife and 5 with his son, owns no property, drives a Tesla Model 3, and spent his first exit money buying apartments for family.
    • He changed his mind on HubSpot: familiar interfaces matter to go-to-market hires, so he no longer thinks everyone will build their own CRM.
    • The finance model projects $4.5 billion in revenue in 12 months. Alex personally believes over $10 billion is possible and is pushing for at least 30% month-over-month growth.
    • Hollywood sentiment has shifted from strictly negative to neutral or slightly negative, with AI increasingly used as a new form of CGI in hybrid production.

    Detailed Summary

    From Competitive Programming in Central Asia to Snap

    Alex describes a childhood built around competition. In Uzbekistan, a family of five earning $1,000 a month is considered wealthy, and his parents, both engineering professors, saw international rankings as the only way out. His mother worked three jobs and his father traveled with him to camps and competitions. By 19 he ranked top three in the world in competitive programming, but instead of academia he went into startups. After working on pre-transformer translation models, he became convinced that phones would become the dominant device and that AI would produce much of the video people watch on them. That led to AI Factory, co-founded with Mahi de Silva, which sold to Snap for $166 million. The dilution was heavy because AI companies were valued at close to nothing back then, but the deal got him to the US. He found San Francisco genuinely meritocratic, though its investors are more consensus-driven than he expected.

    The Near-Death Pivot and Product Market Fit

    At Snap, his team’s on-device face filters drove much of Snapchat’s new user growth. Afterward he focused on a different gap: most companies cannot produce social media content fast enough to stay relevant. Early tools like photo slideshows and long-to-short video clipping were not good enough. Believing scaling laws would come to video, he bet on AI video, but Higgsfield wandered for over a year and burned through most of its seed round. With under $5 million left, the team went back to basics, interviewed eight creative directors, heard that camera control was the missing piece, and launched. Product market fit was immediate, and Higgsfield still does no paid advertising. Asked about an influencer controversy, Alex says outsourcing creator work to an agency was a mistake and that owning distribution matters more than ever.

    How Higgsfield Counts $1 Billion in Revenue

    The interview was recorded the day Bloomberg reported Higgsfield crossing $1 billion in annualized revenue. Alex explains the method: the last four weeks of revenue times 13, which he says matches how leading AI labs report. Annual contracts are prorated so only one 28-day slice counts, and only live revenue is included. Business revenue is slightly over half, and pure consumer use (roughly the mobile share) is under 10%. A large middle group of aspiring creators and freelancers churns but often returns, which is why Higgsfield invests in education like Higgsfield Academy. The standout metric is expansion: one customer went from $99 a month to a $6 million a year contract, largely driven by e-commerce brands producing ads at massive scale and by AI-made short-form dramas.

    Why $20 a Month Subscriptions Are Doomed

    Alex’s contrarian view is that Google and OpenAI, with strong horizontal products, will destroy the consumer $20 a month subscription market for vertical apps. Harry offers Canva as an example of a company whose low-end consumer design use has been eaten by OpenAI. That is why Higgsfield judges success by whether a $20 subscriber can be shown enough value to spend over $1,000 a year. Month-one consumer retention drops about 30% before flattening, which Alex concedes is below the 80% logo-retention bar from B2B SaaS, but business net revenue retention at month 12 is over 300%.

    Content as Distribution and the 150-Person Creative Team

    Competitors told Harry that Higgsfield ran the most impressive influencer campaign in tech. Alex frames it differently: the goal is for the best commercial video content to be made on Higgsfield, with every workflow shown publicly. Over 150 in-house creative professionals, nearly half the workforce, produce launch videos, tutorials and even a fully AI-generated, open-sourced movie. The movie showed how much curation still matters: 90 minutes of TV-quality output required over 100 hours of generated footage.

    Benchmarks, Own Models and Open Weights

    Alex calls his early focus on benchmarks a mistake and describes how large labs game them. By OpenRouter usage, he argues, Google is the only relevant US incumbent in models (later adding Nvidia), while in China Tencent, Xiaomi and Alibaba are all relevant, with ByteDance catching up. For video specifically, benchmarks miss the real work: long prompts, many reference images, and precise control of characters and backgrounds. Higgsfield still builds its own models where customers need them, like its image model for product photo shoots that helped take it from $20 million to $100 million ARR. He says most companies that claim to build models are post-training open weights, and that the most valuable post-training teaches models to collapse multi-step customer workflows into one.

    Tokenomics and the $4 Million Monthly Model Bill

    Most social media marketing does not need frontier intelligence, so customers want cheaper, stable models. Open weights and own models give Higgsfield over 80% margins, while closed models give 20 to 30%, and Higgsfield picks the model in over 40% of workflows. Internally, the team of nearly 400 spends over $4 million a month on models. The creative team started vibe coding tools that do not exist in production, including one person who spent $30,000 in a week building an asset organization workflow over five straight nights. Alex admits spend sometimes goes out of control but calls that case net positive. He expects top engineers and creatives to reach $50,000 to $100,000 a month, while support functions like legal and finance will stabilize at much lower levels.

    Where AI Has Not Replaced Jobs

    Alex expected legal and customer support to be largely replaced and says not staffing them quickly enough was an operational mistake. Today Higgsfield has over 10 people in legal and over 40 in customer success. AI can handle over 60% of first-line support, but not complex B2B requests. Harry notes Revolut’s 92% AI resolution rate for consumers. Alex points out that shipping new products every week makes support agents harder to keep accurate, because their context and rules keep changing. His engineers moved to Claude from March to June and then mostly to Codex, and he expects model releases to keep coming fast, including specialized models like OpenAI’s legal push, which Harry is skeptical of.

    System of Record and the Two Moats

    Using Solve Intelligence as an example, Alex argues whoever builds the AI-native system of record wins. Creative assets are spread across many tools, and marketers want to search them in natural language and check them against brand identity and guidelines. Adobe and Canva built the best software for a pixel-first era, but that is not where the world is heading. Harry says moats are mostly nonsense and cites Lovable and Instinct as wrappers that won on speed. Alex responds that there are only two real moats now: delivering outcomes and network effects. Higgsfield’s open source community projects grew from about 10 to over 10,000 in eight weeks, which he hopes becomes a compounding moat much like forking on GitHub did for software.

    Fundraising, Valuation and the Size of the Market

    Alex names Yuri Milner as his best VC meeting because Milner understood that content trends like AI-native ads and short dramas flow from Asia to the West. He has scars from investors who agreed on a price and then shopped the deal at a 30% discount. Harry argues Higgsfield is discounted for not being a Silicon Valley insider and would easily be worth $25 billion otherwise. Alex says Higgsfield is building for the long term, with aspirations to be the distribution infrastructure for direct-to-consumer businesses the way Shopify became their commerce infrastructure. The West drives over 70% of revenue, and the biggest lesson from Asia is how hard companies there push for direct relationships with customers.

    Family, Sacrifice and Work Ethic

    The conversation turns personal when Harry asks if Alex will regret the time away from his son. Alex talks about his mother’s three jobs, his father’s devotion to his education, and his father’s Parkinson’s diagnosis, which no amount of money can fix. He works 80 to 90 hours a week, tries to spend at least 3 hours a week with his wife and 5 with his son, and has had one full disconnected day with his son in three months. He says there is no shortcut to hard work, citing the product leaders he worked with at Snap, and credits his wife’s patience and a cultural emphasis on mutual sacrifice.

    Hiring, Kazakhstan and Europe

    Over 300 of Higgsfield’s employees are in Kazakhstan. Alex rejects the idea that it is mostly labor arbitrage: Kazakhstan ranks top five in physics olympiads, combines Soviet math with Singaporean teaching methods, sends thousands of students abroad who often return, and offers a 15% personal income tax. He hopes Higgsfield creates more dollar millionaires in Central Asia than any other company. He also pushes back on Harry about Europe, pointing to neoclouds like Nscale and Nebius, application companies like Legora, ElevenLabs and Lovable, and ASML as proof that Europe competes at every layer. He values loyalty and contrasts it with Silicon Valley’s two-year job hopping. On management, he learns from Jensen Huang, Elon Musk and Nik Storonsky, and boils it down to hiring the best, empowering them and retaining them.

    Quick Fire: HubSpot, AI Content, Snap and $10 Billion

    Alex changed his mind on HubSpot, because experienced go-to-market hires value a familiar system of record. His contrarian belief is that most social media content will be AI generated, while authentic content will earn 10 to 15x higher CPMs. The job that does not exist yet is a creative director who talks to a computer and generates stories in real time, where taste is the key skill. He would most like Frank Slootman on his board after reading Amp It Up, though Harry warns that Snowflake’s go-to-market focus lost ground to Databricks’ product focus. His lesson from Snap is that momentum does not last, and he thinks Snap fell behind Meta because it never told a convincing AI story, while Mark Zuckerberg did. Higgsfield’s finance model projects $4.5 billion in revenue in 12 months, but Alex personally says over $10 billion, pointing to monetization-driven adoption and a Hollywood that is slowly warming to AI as a new kind of CGI. He wants to take Higgsfield public and believes it can be bigger than AppLovin and Shopify because distribution is what matters.

    Notable Quotes

    “We burned more than 10 million out of 16 million raised in seed fundraising. So we felt we have just one attempt left.”

    Alex Mashrabov, on the year Higgsfield spent searching for product market fit

    “I was so much optimizing for what’s hype today, what’s the right narrative, how we can hijack the attention, all these things really like everything instead of building a good product.”

    Alex Mashrabov, taking responsibility for the failed first year

    “One customer started 6 months ago spending just subscription $99 a month. $99 a month. And now we just signed a deal over 6 million.”

    Alex Mashrabov, on Higgsfield’s revenue expansion

    “The way to think about video models today, it’s just modern rendering engine. It think about this as like Unreal Engine or Unity but just different types of inputs.”

    Alex Mashrabov, on why text-to-video benchmarks miss real workflows

    “The margin on own models and open weights models is over 80%.”

    Alex Mashrabov, comparing it with 20 to 30% on closed source models

    “The agents are as good as context and rules which they have. And if context and rules change pretty much twice a week, it gets a little difficult.”

    Alex Mashrabov, on why fast product velocity makes AI customer support harder

    “Unfortunately, AI does not replace network effects.”

    Alex Mashrabov, on the two moats he still believes in

    “So yeah, I don’t believe that there is any shortcut to hard work.”

    Alex Mashrabov, on working 80 to 90 hours a week

    “People just still don’t fully appreciate that most of the content on social media is going to be AI generated.”

    Alex Mashrabov, on what he believes that others think is crazy

    “The momentum doesn’t last forever.”

    Alex Mashrabov, on his biggest lesson from Snap

    Watch the full 20VC conversation with Alex Mashrabov here.

    Related Reading

  • Testing the Best-Selling Point and Shoot Cameras From Every Decade: Kodak Instamatic, Canon Sure Shot, Olympus Stylus Epic, Canon PowerShot, iPhone and Fujifilm X100VI Compared

    Point and shoot cameras are having a genuine moment. The Fujifilm X100VI sells for hundreds over retail, and a lot of people are leaving their phones at home in favor of pocket cameras built 20 or 30 years ago. In this episode of Then vs. Now, James Pumphrey and Zach from Speeed take the best-selling point and shoot from every decade since the 1960s, shoot the same film through all of them, put them up against the iPhone, and try to settle three questions: what makes a point and shoot good, why this is all coming back now, and which era actually did it best.

    TLDW

    The lineup runs from the 1963 Kodak Instamatic 100 (126 cartridge, fixed f/11 lens, 70 million sold, scored 4/10 because the film no longer exists), through the 1972 Kodak Pocket Instamatic (tiny 110 film, very portable, photos judged “useless”), the Konica C35 AF that pioneered autofocus in 1977, and the 1979 Canon Sure Shot (Autoboy in Japan, 38mm f/2.8, 11 million sold, 8.5/10). The 1997 Olympus Stylus Epic (the mju II outside the US, 35mm f/2.8 clamshell, about $300 new and around $500 on eBay today) wins with a 9.5 for sharper optics, smarter exposure and speed out of a pocket. The 2006 Canon PowerShot is the cheap, fast, very usable millennial relic that needs an old 2GB SD card to work. The iPhone 6 represents the 2010s, the camera that helped cut point and shoot sales 90% from their peak, and it takes technically good HDR photos that feel flat next to film. The 2020s get both the iPhone 17 Pro Max and the X100VI (40 megapixels, film recipes, $1,800 retail), which scores 7.5: best image quality of the day, but “fake romantic.” The verdict: past a baseline of image quality it is all about experience, the resurgence is about getting off your phone, and the ’90s reign supreme.

    Thoughts

    The first half of this video is secretly a story about film formats, not cameras. The Instamatic and the Pocket Instamatic were both huge commercial wins because Kodak solved a real problem: people ruined film constantly by threading and sealing it wrong, so Kodak put the film in a drop-in cartridge. The trade-off was that each shrink in the cartridge (126 then 110) cost image quality, and both formats are now dead, which is why the Instamatic needed a 3D-printed cartridge loaded in a dark room between every shot. The moment the lineup hits the Canon Sure Shot and standard 35mm film, the photos jump from curiosities to things you would actually frame. If you are buying a vintage point and shoot today, the lesson is simple: buy for the film you can still get, not for the camera’s legend.

    The more subtle thread is how much of “point and shoot” is really the camera making decisions for you, and how that intelligence evolved. The Instamatic’s answer was to stop everything down to f/11 and hope, a middle-ground approach that only works outside in daylight. The Sure Shot opened up to f/2.8 and used autofocus and auto exposure to actually read the scene. The Stylus, shooting the exact same film at the same aperture, exposed the shadows better and found a kid in the corner that the Sure Shot crushed to black. Then the iPhone takes it to the extreme by shooting multiple exposures and merging them into HDR. Each step made the camera smarter and the photographer less involved, and the video makes a decent case that somewhere around 1997 the balance was perfect.

    The iPhone segment is the most honest part. The iPhone 6 photo of the river is objectively better than the Sure Shot’s: more shadow detail, a real sky, no noise. And yet the reaction is “I would rather hang this on my wall,” pointing at the film print. Technically better and emotionally flat is a real category, and the blind test at the end backs it up: James identified every film photo correctly and mixed up all the digital ones. Once digital processing gets good enough, cameras start to converge on the same look, while film keeps its fingerprint. That is a useful thing to know before paying a premium for a digital camera that promises character.

    The Fujifilm X100VI critique is the sharpest idea in the video and the one most reviews skip. It takes the best photos in the lineup, has recipes that preview a film look right in the viewfinder, and it still gets a 7.5, below a $120 camera from 1979. The complaint is not quality, it is that the X100VI sells the aesthetic of disconnection while keeping the habit of reviewing every frame on a screen. The old cameras give you a click and a question mark (“I don’t know if I got it”), and that uncertainty is exactly what keeps you in the moment. “Fake romantic” is harsh, but it names the gap between a camera that looks like 1970 and one that behaves like 1970.

    The closing answer is the one worth keeping. Past a bar for image quality, a good point and shoot “lets you capture the moment without taking you out of it,” and the resurgence is about people wanting off their phones: a phone camera is a door into doom-scrolling, and a point and shoot is disconnected by design. James keeps his Stylus by the front door and grabs it every time he leaves, and he ends by calling joy a weird metric before correcting himself: maybe joy is a normal metric. That is the right frame for any tool purchase. The best camera is not the one that wins the spec sheet, it is the one you actually carry and enjoy using.

    Key Takeaways

    • Point and shoot cameras are surging in popularity, with new models like the Fujifilm X100VI reselling above retail and older digital compacts from 20 to 30 years ago becoming fashionable everyday carry.
    • The video sets three questions: what makes a point and shoot good, why the resurgence is happening now, and which decade produced the best one.
    • Selection rules: the best-selling camera of each decade, priced between $100 and $1,500 adjusted for inflation, with no interchangeable lenses. That excludes disposables and high-end cameras like Leica and Ricoh.
    • The Kodak Instamatic 100 (1963) launched at $16 (about $160 today), used a 43mm fixed f/11 lens, sold 70 million units, and sells for about $30 now.
    • Before the Instamatic, loading film meant threading, seating, winding and sealing by hand, and people ruined film all the time. Kodak’s 126 drop-in cartridge fixed that.
    • The Instamatic gives the user no control at all: fixed focal length, fixed aperture, fixed film sensitivity. You point and press.
    • f/11 is a small aperture that keeps more in focus but lets in less light, so the Instamatic mostly needs to be used outdoors in daylight.
    • 126 cartridges are no longer made, so the team used a 3D-printed cartridge that had to be advanced in a dark room between every shot, causing heavy light leaks. Score: 4 out of 10.
    • The Kodak Pocket Instamatic (1972) was the best-selling camera of the 1970s, built around the smaller 110 cartridge so it could actually fit in a pocket. It sells for around $30 today.
    • The 110 format’s small negative produces low resolution, visible grain and very poor dynamic range. The hosts called the results “useless,” and 110 film is hard to source.
    • Kodak’s own manual told users to take close-ups of people, likely because the tiny negative could not hold detail at a distance.
    • Film was developed, then scanned at high resolution and printed. For the final comparison they also used a fully photochemical print process with no scan.
    • Konica’s C35 AF (1977) was the first mass-produced camera to combine autofocus, auto exposure and a built-in flash.
    • Canon had been developing its own autofocus triangulation system for 14 years and released the Autoboy, sold in America as the Sure Shot, in 1979, about 18 months after the Konica.
    • The Sure Shot sold 11 million units across its first three generations, launched at $240 (just under $1,000 today), and sells for about $120 now.
    • Its 38mm f/2.8 lens lets in far more light than f/11, which is why it needs autofocus, but it makes lower-light photos possible.
    • Standard 35mm film was a huge advantage. The same film stock was used in every film camera for a fair comparison.
    • The Sure Shot produced strong blues, great greens and nice flash photos, but auto exposure sometimes crushed shadows where people were standing. Score: 8.5 out of 10.
    • If the ’80s were the golden age of 35mm point and shoot innovation, the ’90s were when they became perfect.
    • The Olympus mju II (1997), sold in the US as the Infinity Stylus Epic, has a 35mm f/2.8 lens and a clamshell cover that protects the lens and makes it pocketable.
    • Olympus sold 3.8 million mju IIs. It cost around $300 new and now sells for about $500 on eBay, making it one of the most sought-after cameras on the secondary market.
    • In a pocket-to-photo speed test, the Stylus beat the Sure Shot easily. Fast startup is one of the most praised traits in owner reviews.
    • On identical film at the same aperture, the Stylus was sharper (better optics), wider, and exposed shadows better than the Sure Shot.
    • A wider lens is safer: you can always walk closer, but indoors you often cannot back up. The Stylus is reportedly prone to light leaks. Score: 9.5 out of 10.
    • Digital arrived at the turn of the millennium, and point and shoots were one of the first categories to adopt it.
    • The 2006 Canon PowerShot they tested cost $430 at launch (about $830 today) and now goes for around $100.
    • The PowerShot is a millennial and Myspace-era icon, famous for mirror selfies and for Maria Sharapova’s “Make every shot a PowerShot” ads.
    • Old PowerShots only accept small SD cards (around 2 to 4 GB), which are hard to find when stores start at 64 GB.
    • PowerShot images look clearly digital (sharp, no grain), but the camera is tiny, fast, cheap and arguably the best entry point into intentional photography.
    • The best-selling camera of the 2010s under the rules was the iPhone 6. By 2015 “Shot on iPhone” was a global billboard campaign.
    • The iPhone 6 had an 8 megapixel camera. By 2015 point and shoot sales had fallen 90% from their peak, and Kodak had filed for bankruptcy in 2012.
    • An old iPhone 6 costs about $40 today. The hosts credit smartphones with putting a good camera in many more hands, while joking that phones ruined the world.
    • The iPhone’s HDR merges multiple exposures. Here it worked well, holding both sky and shadows, but the result still felt digital and less organic than film.
    • The best-selling camera of the 2020s so far is also an iPhone, the iPhone 17 Pro Max, with three cameras, a 48 megapixel sensor and heavy computational processing.
    • The Fujifilm X100VI has 40 megapixels, a fixed lens, manual controls, AI autofocus, skin smoothing and film simulation “recipes.” It retails for $1,800 and resells for $300 to $400 above that.
    • The X100VI took the best photos in the lineup, with detail in shadows none of the others captured, but the hosts felt it cashes in on nostalgia and still pulls you out of the moment to review shots. Score: 7.5 out of 10.
    • Answer one: there is a bar for image quality, and past it, experience is everything. A good point and shoot captures the moment without taking you out of it.
    • Answer two: the resurgence is about getting off phones. Phone cameras feel invasive and lead to doom-scrolling. A point and shoot is simple and disconnected, and handing one to someone is fun rather than a privacy risk.
    • Answer three: the ’90s win. The Stylus is peak mechanical and analog, easy to carry, and the most joyful camera to use.
    • In a blind test of identical scenes, James identified all the film photos but mixed up the digital ones, and Zach spotted the Sure Shot by its darker shadows.

    Detailed Summary

    The Setup: Why Point and Shoots, and Why Now

    The episode opens on the internet’s current obsession with point and shoot cameras. New releases sell for well over retail, and people are swapping their phones for compact digital and film cameras built decades ago. James and Zach gather the best-selling point and shoot from each decade (film and digital) and include the most popular point and shoot of all time, the iPhone, because the obvious question for any buyer is why bother with a dedicated camera at all. The rules: best-selling within a $100 to $1,500 inflation-adjusted price band, and no interchangeable lenses.

    1960s: Kodak Instamatic 100

    Before the Instamatic, taking photos meant threading film, making sure it was seated, winding it and sealing the back, with many ways to ruin a roll. Kodak’s 126 cartridge simply dropped in. The camera itself offered no control: fixed focal length, fixed f/11 aperture, fixed film speed. It is the camera behind countless mid-century family snapshots of uncles, babies, cars in driveways and terrifying Easter bunnies. At $16 on launch and 70 million units sold, it was revolutionary. Today, though, 126 film is gone. The team’s only real cartridge expired in 1979, so they used a 3D-printed one that had to be advanced in a dark room between frames, which produced heavy light leaks. The 35mm film inside gave nice grain and surprisingly vibrant color, but focus was soft and the practical experience was miserable. Score: 4 out of 10.

    1970s: Kodak Pocket Instamatic

    After the original’s success, Kodak’s priority was shrinking. The 1972 Pocket Instamatic used 110 film to fit in a pocket and became the best-selling camera of the decade. It looks like a spy camera and has an absurd flash that roughly quadruples its size. The smaller negative is its downfall: the photos show low resolution, visible dots and almost no dynamic range, with shadows crushed and highlights blown. Even Kodak’s instructions nudged users toward close-ups of people. The hosts admit they wanted to find charm in the results because the camera is old, but concluded the photos were just bad, and 110 film is hard to find. Great usability, disappointing output.

    1980s: Canon Sure Shot (Autoboy)

    Autofocus changed everything. Konica’s C35 AF in 1977 was the first mass-produced camera to combine autofocus, auto exposure and flash, but Canon had spent 14 years on its own triangulation system and launched the Autoboy (the Sure Shot in America) 18 months later. It sold 11 million units across three generations, cost $240 at launch and goes for about $120 today. Its 38mm f/2.8 lens reads the scene rather than averaging everything the way the Instamatic did. Shot on standard 35mm, the images were a huge leap: rich blue skies, clouds, great greens that digital often struggles with, and flattering flash. Auto exposure still made some odd calls, leaving people in shadow. Score: 8.5 out of 10.

    1990s: Olympus Stylus Epic (mju II)

    The 1997 Olympus mju II, the Infinity Stylus Epic in the US, is James’s own camera. It has a 35mm f/2.8 lens behind a sliding clamshell, a lanyard instead of a neck strap, and a reputation as one of the fastest cameras from pocket to photo, which a head-to-head race against the Sure Shot confirmed. Olympus sold 3.8 million, and prices have climbed from about $300 new to around $500 on eBay. On identical film, the Stylus was sharper thanks to better optics, wider, and made smarter exposure decisions, pulling detail out of a shadowed corner (and a child nobody noticed) that the Sure Shot left black. It has some reputation for light leaks and you need to mind the focus, but the photos looked amazing. Score: 9.5 out of 10.

    2000s: Canon PowerShot

    Then digital arrived, and point and shoots were among the first cameras to adopt it. Exact sales were hard to track, but the Canon PowerShot line was clearly among the best sellers, so they picked a 2006 model: $430 at launch (about $830 today) and about $100 now. Tiny enough to look like a Zoolander prop, it is a millennial relic, the camera of Myspace and mirror selfies, marketed by Maria Sharapova with “Make every shot a PowerShot.” Finding a small enough SD card was its own quest. The images look immediately digital, sharp and grainless, not as cool as film but good. The hosts liked it more than expected: small, fast, cheap to run with no film or scanning costs, and arguably the best entry-level way into more intentional photography.

    2010s: iPhone 6

    By the rules of the video, the best-selling camera of the 2010s was the iPhone 6. Its 8 megapixel camera was good, fast and convenient enough to effectively end the compact camera market: “Shot on iPhone” was on billboards worldwide, point and shoot sales fell 90% from their peak by 2015, and Kodak had filed for bankruptcy in 2012. An old iPhone 6 now costs about $40. The hosts stress they are not anti-iPhone, since it put good cameras in many hands, and a later iPhone shot the acclaimed film Tangerine. The iPhone’s HDR worked well here, holding sky and shadow without noise. Still, next to the Sure Shot and Stylus prints, it felt technically better but less organic. The iPhone takes a good photo, and nobody is impressed by it.

    2020s: iPhone 17 Pro Max and Fujifilm X100VI

    The best-selling camera of this decade is again an iPhone, the iPhone 17 Pro Max, whose three cameras and heavy processing make it a “point and shoot and edit and think” device. To represent the revival itself, they added the Fujifilm X100VI: 40 megapixels, a fixed lens, manual dials, plus AI autofocus, skin smoothing and film recipes that preview the look in the viewfinder. It retails for $1,800 and resells for $300 to $400 more. It produced the best photos of the day, including detail in shadows that no other camera captured, and the viewfinder recipes are fun. But the hosts felt it cashes in on nostalgia, costs far more than the rest, and brings back the habit of checking every photo. Score: 7.5 out of 10.

    The Verdict and the Blind Test

    What makes a good point and shoot is clearing a bar for image quality and then delivering an experience that lets you capture a moment without leaving it. The resurgence is about getting off phones, which feel invasive and pull you into scrolling, while a point and shoot is simple and disconnected. The era that reigns supreme is the ’90s, which James compares to ’90s cars: peak mechanical, analog and built to be used. He keeps his Stylus by the front door and suggests joy may be a perfectly normal metric. In a closing blind test of six identical scenes (fully photochemical prints for the film shots), James nailed all the film images but confused the digital ones, and Zach picked out the Instamatic by its light leaks and the Sure Shot by its darker shadows.

    Notable Quotes

    “Chances are, any time you see an old-ass photo of, like, some dude with slicked-back hair holding a baby next to a car in a driveway, it’s probably this camera.”

    The hosts, on the Kodak Instamatic’s place in family photo history

    “If the ’80s was the golden age for innovation for 35mm point and shoots, the ’90s was when they became perfect.”

    Narration, introducing the Olympus Stylus Epic

    “You can always walk closer to something, but sometimes you’re in a room or whatever. You can’t get any further away. So, like, a wider lens is a little safer.”

    The hosts, on why the Stylus’s wider 35mm lens beats the Sure Shot’s 38mm

    “I mean, this feels technically better, but I would rather hang this on my wall.”

    Comparing the iPhone 6’s HDR photo with the Canon Sure Shot’s film print

    “I think it’ll probably take amazing photos, but it feels like it’s cashing in on nostalgia in a way where I’m like, you should just buy the old camera.”

    On the Fujifilm X100VI and its $300 to $400 resale premium

    “The other cameras, one of the things that is so fun about them is, like, click. I don’t know if I got it.”

    On why the film cameras feel better to use than the X100VI

    “I think there’s a bar for image quality, and once you’ve passed that bar, it all comes down to experience. A good point and shoot lets you capture the moment without taking you out of it.”

    The answer to what makes a point and shoot good

    “Phone cameras can feel invasive, and once you pull out your phone and start looking at it, chances are you might keep looking at it.”

    On why the point and shoot resurgence is happening now

    “It’s, like, a weird metric, but it’s the most, like, joyful camera to use.”

    James Pumphrey, on the Olympus Stylus Epic he keeps by his front door

    Watch the full Then vs. Now camera episode here, and see more from the team at Speeed.

    Related Reading

  • TypeSafe CEO Diogo Almeida on Jev and Building Prod, Not God: Why AI Still Hasn’t Automated the Easy Stuff, Smart Software vs Coding Agents, Reliability Over Benchmarks, and the Inverse SaaS Apocalypse

    Diogo Almeida, co-founder and CEO of TypeSafe AI and a former OpenAI and Google Brain researcher who worked on the RLHF behind InstructGPT and ChatGPT, sat down with a16z’s Martin Casado and a fellow a16z partner for How Jev Builds Prod, Not God. Jev is TypeSafe’s first model, and it is not a chatbot or a coding agent. It is a primitive you put inside your code: you give it natural language and program state, and it returns a typed decision with a confidence level. Almeida’s pitch is blunt. AI is unbelievably smart, and yet almost nothing is automated. This conversation is his case for why that happened and how to fix it.

    TLDW

    Almeida distinguishes Jev from coding agents like Claude Code, Codex, and Cursor, which Garry Tan calls “just in time software.” They write the same code a human would. Jev is a new primitive, a library that takes natural language and state and returns a choice with probabilities, so software itself can do things it never could. He happily calls it a classifier, argues it likely beats a 2019 ML engineering team you can program on the fly, and names intelligence per dollar as his north star. He traces his path from reluctant mathlete to a Kaggle win built on brute-force automation, which led to Isabelle Guyon, Jeremy Howard, Google Brain, retirement, and OpenAI. He explains how RLHF’s surprising generalization in late 2021 made him think AGI was near, and how its failure to deliver turned into a chip on his shoulder: the industry optimized models for the human judge instead of for automation, so GPQA gets solved while a drive-thru still cannot be automated. He rejects the data and long-tail excuse, defines reliability as uptime, determinism, robustness, and being “smart every time,” and says the highest honor is developers programming against Jev without testing example queries. The hosts discuss why SaaS stocks fell on coding agents but cheered Jev. Almeida predicts an inverse SaaS apocalypse, says coding agents are good at syntax and bad at architecture, and the hosts note the average enterprise PR is about 10 lines. The close covers probabilistic programming, rebuilding systems for security, why most future AI calls will be deep in software’s guts rather than facing humans, and his vision of technology that simply does what you mean.

    Thoughts

    The cleanest idea here arrives in the first few minutes and it reframes the whole AI coding debate. Coding agents automate the act of writing software, but the software they produce is the same kind of software we had ten years ago. Almeida wants the opposite: leave software engineering mostly alone and expand what software itself can do. Jev is a function call that takes fuzzy input and returns a typed, confident decision a program can branch on. That is a different category from “AI that codes,” and it explains why a model with no chat interface caught fire among developers. It gives them a new instruction, not a faster typist.

    His diagnosis of why AI has automated so little (roughly 19 to 24 minutes) is the most provocative part, especially coming from someone who helped build RLHF. Once humans became the evaluators, the industry optimized for the judge. Models look brilliant to the person reading their output, so they score well, but nobody was optimizing for whether they could run unattended inside a business process. That is how you get a world where GPQA is solved and a drive-thru is not, and where OpenAI has been trying to automate customer service since 2020. The a16z host pushes back with the long-tail data argument, the familiar story of a help desk that “automates 95%” when most of it is password resets. Almeida’s response is pragmatic rather than dismissive. You do not need the long tail. You need the boring, high-volume core to be automatable, and then automation becomes an ROI decision like any other engineering investment.

    The reliability section (25 to 28 minutes) is where the “prod, not god” slogan earns its keep. Almeida splits reliability into uptime, determinism, and robustness, where robustness means similar intelligence every time. Adding a UUID to a prompt should not change the answer, even if the output is not bit-for-bit identical. Then he adds a fourth layer he does not have a name for, being smart every time in a way a human would find understandable, because a developer can program around that. His bar for success is developers calling Jev without first testing example queries, the way they call a sorting function without checking it on sample data. That is a much higher and more useful target than a leaderboard number, and it is exactly the thing a “benchmaxxed” copycat would miss.

    The SaaS discussion (30 to 35 minutes) is a sharp market read. Coding agents triggered a “SaaS apocalypse” in public markets because the story was that software is now cheap to replicate. Almeida accepts cheap but not easy to replicate, because the value sits beneath the surface in workflows, users, and distribution. If AI becomes something you embed in software rather than something that replaces it, incumbent SaaS companies become the best placed winners. They already know which workflows need automating and have already paid to reach every customer. The host’s data point lands hard: the average enterprise pull request is about 10 lines, so automating code writing optimizes a small slice of the work while adding no new capability, and possibly making software worse and less secure through less oversight. An “inverse apocalypse” is a bet worth taking seriously.

    The closing stretch explains the business logic behind the design. Almeida works backward from a world with AI everywhere and asks what share of all AI calls will be for human consumption, where style matters, versus deep inside programs making decisions. His answer is many nines in the guts, which is why intelligence per dollar matters more to him than eloquence, and why the input to Jev is called “state.” The hosts give the best description of the status quo: AI and software have been ships in the night. Developers stuffed JSON schemas into prompts, watched the model ignore them, and then went through five stages of grief ending in two workarounds, a human in the loop (chat) or another LLM in a while loop (agents). Jev is a bet that you can map a model directly onto a state machine and skip both. If it works, “do what I mean” stops being a sci-fi phrase and becomes a property of ordinary software.

    Key Takeaways

    • Almeida’s favorite elevator pitch for Jev is “where is all the automation?” AI is extraordinarily smart and yet almost useless outside chat and coding.
    • TypeSafe’s mission is making AI work for software, not just for humans in the loop. Jev is its first model.
    • He credits Garry Tan’s description of Claude Code and Codex as “just in time software”: they make software on the fly from natural language, with the same expressive power as ordinary software.
    • What Almeida wants instead is smart software, expanding what software can do so that things that should be automatable become automatable.
    • Coding agents write the same code a human would. Jev is a new primitive you include in your code, whether a human or a coding agent is writing it.
    • Practically, Jev works like a library: describe what you want in natural language, pass state, and it chooses what to do with confidence levels.
    • On day one of onboarding, Almeida draws a Venn diagram of what AI is good at and what is valuable in code. Jev lives in the overlap, which is why it outputs probabilities and not extrapolated floats.
    • He embraces the “it’s just a classifier” critique. Classifiers were designed by practical people to be useful.
    • His guess is that Jev beats having a 2019 ML engineering team build a narrow model for you, and you can program it on the fly.
    • In his heart, the design space is a slider from language-in, language-out to fully imperative programs. Pragmatically, Jev will behave more like a database than a standard library for a while.
    • Intelligence per dollar is his current north star, though he admits intelligence per second may matter more in the short term.
    • The input is deliberately called “state” because Jev is meant to live inside programs.
    • Almeida was an award-winning mathlete who never loved math. Computer science felt like math but cool, useful, and fun, and he considers himself a computer scientist before an AI researcher.
    • He won a Kaggle competition by automating aggressively rather than through sophisticated math, and was then invited to speak at NeurIPS.
    • The competition host, Isabelle Guyon, co-inventor of the support vector machine, took him under her wing and introduced him to the AI community.
    • His career path ran through a startup with Jeremy Howard, Google Brain, a period of retirement, and then OpenAI because AI was simply fun.
    • The hosts see TypeSafe as a movement toward a positive AI future, contrasting “happy AI” Jev users with the “morose AI” crowd.
    • Almeida blames the negative outlook on “mono model Kool-Aid,” the idea of one big brain that rules everything, while basic tasks remain unautomated.
    • He does not buy diffusion as the excuse for slow automation, given the huge financial incentive to automate.
    • An anecdote: a16z’s David George used Meta’s Muse to finally cancel his New York Times subscription, the tip of the iceberg of horrible tasks that need automating.
    • Almeida warns against AI’s anti-pattern of focusing on outliers and demos. He wants use cases that run in the background without paging anyone and that others can build on.
    • Running AI with access to real resources requires guarantees, or at least statistical guarantees.
    • His 2017 talk had a similar theme, roughly “AI modular in theory and flexible in practice.”
    • In late 2021 the RLHF team was surprised by its generalization, verifying with prompts like “why is it important to eat socks before meditating?” that were not on the internet.
    • He thought that model had a decent chance of being AGI. When it was not, his world came crashing down and he began asking why AI was not more useful.
    • RLHF generalizes fairly well in his experience, while RLVR generalizes less well.
    • He does not think we are on a path to recursive self-improvement, but considers OpenAI’s definition of AGI, automating most economically valuable work, extremely doable.
    • Much work is rote and simple enough to outsource with basic instructions, and models have had that level of intelligence for a while.
    • Since RLHF, the industry has overpromised and underdelivered because humans judge the models, so labs optimized the judge instead of automation.
    • His canary in the coal mine: we say math and GPQA are solved, yet we still cannot handle a drive-thru.
    • He does not buy the argument that missing real-world data explains the gap. The long tail is real, but automation does not need to cover it.
    • Echoing the programmer virtue of laziness, automation should be an ROI decision, and people will create new kinds of work once rote work is automatable.
    • OpenAI has been trying to automate customer service since 2020, and outside of programming very little inside companies has been automated.
    • Almeida was extremely surprised by the launch’s reception and says no one could have predicted a ChatGPT moment for developers.
    • Reliability is what Jev is. Every nine of reliability enables new applications, and without understanding that, you cannot build a real copy.
    • He defines reliability in layers: uptime and SLAs, determinism (useful for unit tests), robustness (similar intelligence every time), and being consistently smart in an understandable way.
    • The highest honor would be developers programming against Jev without trying example queries first.
    • Coding agents are good at syntax, weak at semantics, and very bad at architecture, which he sees as the most human, creative part of software.
    • Model-produced architecture might be 50th percentile. That is a legitimate trade-off if speed matters more than quality.
    • SaaS valuations fell when coding agents arrived, but SaaS companies loved Jev. Almeida thinks SaaS will be one of AI’s biggest winners.
    • Software may be cheap but it is not easy to replicate, because the value is beneath the surface. SaaS firms know which workflows to automate and already have distribution.
    • He calls the likely outcome an inverse SaaS apocalypse and imagines multiple choice forms disappearing.
    • His favorite community application is voice control of a computer that constantly decides whether speech is a command or text to insert, and where.
    • An a16z study found the average large-company PR is about 10 lines, so coding agents automate a small slice without adding capability, and may make software worse and less secure.
    • Almeida’s grand hope is to expand software beyond basic logic gates with a new kind of gate that has a little brain in it.
    • His philosophy is to automate the easy work before the hard work, but he expects a new era of probabilistic programming.
    • His brand is pragmatism, and he is not a fan of biologically inspired AI.
    • The hosts argue that a new primitive plus cybersecurity pressure means much of critical infrastructure will be rebuilt.
    • Almeida thinks of AI like TCP and UDP. Most AI calls will eventually be deep in software’s guts rather than facing humans, and you must aim for the guts to get there.
    • AI and software have been ships in the night. Chat is the human in the loop and agents are a while loop feeding language back into another model.
    • TypeSafe could have released much sooner but held back for reliability. Almeida’s utopia includes all technology simply doing what you mean.

    Detailed Summary

    Where is all the automation?

    Asked for an elevator pitch, Almeida offers a question: where is all the automation? He loves chatbots and coding agents, but finds it tragic that such intelligence is so useless for everything else, a diamond in the rough that has not been polished for work. TypeSafe is making AI for software, powerful not just with humans in the loop but inside real software, and Jev is its first model toward that goal. The hosts note that developers have been calling them to rave about it, prompting the obvious question of how it differs from Claude Code and Codex.

    Smart software versus just in time software

    Almeida borrows Garry Tan’s phrase “just in time software” for coding agents, which let you program in natural language while producing ordinary code. He wants smart software instead, expanding the vocabulary of what programs can express, including something like intent. He loves that programming means hyper-specifying valuable things and replicating them infinitely, and he wants more of that. The hosts sharpen the distinction: whether Claude Code or a human writes it, Jev is something you include in your code. It is a library where you describe what you want in natural language, supply a state machine, and get back a choice with confidence levels, something software has rarely had so widely. It asks programmers to think in probabilities.

    Yes, it is a classifier

    Almeida finds it wild that AI is this capable while software has been unchanged for a decade, with the best effort being a sidebar chatbot that can take some actions but not all, because some actions are not reliable. He draws a Venn diagram for new hires of what AI is good at and what is valuable in code, and Jev sits in the middle. Probabilities are in the overlap, extrapolated floats are not. To the “Jev is just a classifier” critique he says absolutely, classifiers are great. They share interfaces with classic ML concepts invented by practical people. He guesses Jev beats a 2019 ML engineering team, which few companies ever had, because you can program it on the fly without collecting and measuring datasets.

    Design choices: a slider, state, and intelligence per dollar

    One host asks whether there is a slider from language-in, language-out to imperative programs, or whether language-in, state-machine-out is the design point that will solidify. Almeida says that in his heart it is a slider. His north star for now is intelligence per dollar, though intelligence per second might be more valuable in the short term. Calling the input “state” is intentional because Jev is meant to live inside programs, and much of his work targets ever more complex arrangements of program internals. Pragmatically, it is easier to hit certain latencies in a database-like service, so Jev will look more like a database for a while, though he would love it to become a standard library feature too.

    From mathlete to OpenAI

    Almeida was an award-winning mathlete who never liked math, a big fish in a small pond who resented competition. Computer science felt like math but useful and fun, and he still loves giving algorithms interviews because they reveal a lot about candidates. He won a Kaggle competition by automating heavily, with more nested loops and a systems approach rather than sophisticated math, and was pushed to speak at NeurIPS. The host, Isabelle Guyon, co-inventor of the SVM, saw someone who did not fit the research mold and introduced him to the AI world. From there he joined a startup with Jeremy Howard, then Google Brain, then retired for a while, and finally joined OpenAI because AI was fun.

    Prod, not god, and the happy AI camp

    The hosts praise TypeSafe’s slogan “we build prod, not god” and its optimism about more and better jobs, framing it as a movement. Almeida says critics raising the classifier point are voicing an ML-level concern while developers are partying, because they can finally do what they wanted. He blames the negative worldview on mono model thinking, one big brain to rule them all, even though basic, unwanted work remains unautomated. He rejects diffusion as an excuse. One host recounts David George canceling his New York Times subscription with Meta’s Muse. Almeida says honest pursuit of automation means avoiding AI’s fixation on demos and outliers in favor of workflows that run in the background, do not page anyone, can be composed, and come with at least statistical guarantees when they have access to resources.

    RLHF, AGI, and optimizing the judge

    A host recalls talking to Almeida in 2017, when his talk was roughly “AI modular in theory and flexible in practice.” Almeida says the real turn came just before ChatGPT, in late 2021, when the RLHF team was surprised by how well it generalized. Their paper tried to disprove its own claims, including testing prompts like “why is it important to eat socks before meditating?” that did not exist online. He pushed hard to release that model and thought it had a decent chance of being AGI. When it was not, his world came crashing down. He says RLHF generalizes well and RLVR less so. He recalls early OpenAI describing AGI as “Ilya and every if statement,” a deliberately vague big tent. He does not think we are on a path to recursive self-improvement, but thinks automating most economically valuable work is very doable. Much work is rote, and the needed intelligence has existed for a while. Since RLHF, the industry has optimized the human judge rather than automation.

    The long tail argument and new kinds of work

    Almeida’s canary: if math and GPQA are solved, why can we not handle a drive-thru? A host proposes that real-world distributions are heavy-tailed and underrepresented in training data. Almeida does not buy the data argument. The long tail exists, but you do not need to automate it. Building reliable software is always an investment, and he invokes the programmer virtue of laziness, spending ten hours to never do a five-minute task again. Automation should be an ROI decision, and he believes people will invent new kinds of work once rote work can be automated, an argument close to the Jevons paradox. The hosts add that OpenAI has pursued customer service automation since 2020, and that pre-generative support vendors claiming 95% automation were mostly handling password resets, closer to 50% by uniqueness.

    Reliability as the product

    Almeida was extremely surprised by the launch, which even non-developer friends joined in memeing. But he stresses years of work on reliability, which he says is what Jev is. Without understanding that, you cannot build a copy that is not just benchmaxxed. Every nine of reliability unlocks new applications, even ones the team does not yet understand. Asked what reliability means for a stochastic system, he lists uptime and SLAs, determinism, and robustness, meaning similar intelligence every time, so adding a UUID to a prompt should not change the result. A further, unnamed layer is being smart every time in ways a human would find understandable, which developers can program around. The goal is developers trusting Jev enough to skip example queries and work in a flow state.

    Coding agents, syntax, and architecture

    A host suggests that a primitive like Jev could lower the value of coding agents, since agent-written software that does not use it stays limited. Almeida calls this more of a coding agent question. In his experience, agents are very good at syntax, weak at semantics, and very bad at architecture, which he sees as the most creative human part of software. Jev is almost certainly not in their training distribution yet. When it is, he is happy for agents to handle syntax. Their architecture might be 50th percentile, which is fine if you know nothing about architecture or if speed is the knob your project wants to turn, for example letting Codex work overnight.

    The inverse SaaS apocalypse

    The hosts note that SaaS stocks plunged on coding agents, yet SaaS companies welcomed Jev. Almeida says the apocalypse story, software being cheap and easy to replicate, has played out poorly. It may be cheap, but it is not easy to replicate, because the value sits beneath the hood. He wants to work with the biggest, most boring SaaS companies that know user problems best, since they know which workflows to automate and have already made the capex investment to reach users. He calls it an inverse apocalypse. The hosts add that SaaS capital largely goes into reaching customers, so making the software genuinely better, not just adding a chatbot, is powerful. Almeida imagines multiple choice forms disappearing, and the hosts compare it to 1980s fourth-generation languages. He calls it “do what I mean” taken to the next level and highlights a community project that uses voice to control a computer, constantly deciding whether speech is a command or text to insert.

    Better software, not just faster software

    A host shares an a16z finding that the average PR at a large company is about 10 lines, often capturing something learned from a customer. Coding agents optimize that minimal slice without adding capability, and software may be getting worse and less secure due to reduced oversight. Jev, by contrast, speaks natural language and reasons while being married to a state machine, so apps can gain new functionality. Almeida says that takeaway would be the greatest compliment. His grand vision is expanding software beyond its basic logic gates with a gate that has a little brain in it, and he promises to fight for it without overpromising.

    Probabilistic programming and rebuilding systems

    One host questions how deep this can go into systems needing strong guarantees like state consistency and durability, versus log analysis, email, and UI. Almeida’s philosophy is to automate the easy work first, but he expects a new era of probabilistic programming, which the hosts note largely died decades ago. They mention co-founder Erik’s Bayesian background, and Almeida says his brand is pragmatism and he is not a fan of biologically inspired AI. The hosts agree such ideas mostly motivate people for decades until engineering refines them. Almeida expects systems engineers to use cheap, fast intelligence for approximate guesses and optimistic routing. The hosts add that new primitives and cybersecurity pressure mean much critical infrastructure will need rebuilding, as happened with the internet and client-server.

    Aiming for the guts, and do what I mean

    Almeida describes his thinking in TCP and UDP terms, which one host calls speaking his language. He worked backward from a world where AI is everywhere, asking what share of AI calls are for humans versus buried in software. His answer is many nines deep in the guts, starting at the surface. If you do not aim for the guts, you will not get there. The hosts describe AI and software as ships in the night: developers put JSON schemas in prompts, the model ignored them, and after five stages of grief they handed output to a human (chat) or another LLM in a loop (agents). This is the first time they have seen AI productively mapped onto a state machine. Almeida says TypeSafe could have released much sooner but held back for reliability, and hopes users simply feel they can trust it. His AI utopia includes technology that does what you mean, which he says is not sci-fi given how smart AI already is.

    Notable Quotes

    “AI is so unbelievably smart and yet it’s so useless at all other stuff.”

    Diogo Almeida, in the opening pitch for why TypeSafe exists

    “What I want instead is smart software.”

    Diogo Almeida, contrasting Jev with coding agents that produce just in time software

    “Jev is absolutely a classifier. You know, like classifiers are sick.”

    Diogo Almeida, embracing the most common critique of the model

    “We’ve been optimizing that judge instead of the automation part and that has been the missing thing.”

    Diogo Almeida, on how RLHF-era evaluation led the industry to overpromise

    “OpenAI has been trying to automate customer service since 2020.”

    Diogo Almeida, on the gap between benchmark progress and real automation

    “Reliability is what this thing is.”

    Diogo Almeida, on years of work that a benchmark-chasing copycat would miss

    “It doesn’t matter how much AI coding agents you use, the software actually isn’t getting better.”

    An a16z host, on why a new primitive matters more than faster code writing

    “I think automate the easy work before the hard work is always my philosophy.”

    Diogo Almeida, responding to questions about using Jev in systems that need strong guarantees

    “Imagine if all technology just did what you mean.”

    Diogo Almeida, closing on his vision of an AI utopia

    Watch the full conversation with Diogo Almeida here.

    Related Reading

  • Instinct Founder Noah Shinn on the Personal AI Assistant Growing 10% a Day: Earning Trust, Take Rates Instead of Ads, Buying Compute Months Ahead, the Trusted Person Network, and Why All Software Collapses Into One Interface

    Noah Shinn started Instinct about a year ago, and in his first long-form interview about the company he sat down with Patrick O’Shaughnessy on Invest Like the Best to explain how an invite-only personal AI assistant with no app and zero marketing spend is growing roughly 10% a day. Instinct has its own phone number, email address, and computer. You text it, call it, or email it, and it acts for you anywhere on the internet. The conversation runs from wild user stories to trust metrics, the business model, the coming fight between agents and incumbent apps, and the brutal math of buying compute for a product that doubles every week.

    TLDW

    Instinct is “just a personal assistant” reached through iMessage, WhatsApp, voice calls, and email, backed by its own computer so it can do anything a person does online. Users have it plan outfits from scanned wardrobes, cancel forgotten subscriptions end to end, coordinate group outings and shared Ubers, and book whole trips from a single voice note. A new trusted person network lets two users’ agents negotiate meeting times directly, with weighted access levels and real social consequences when trust is broken. Shinn says 40% of users share a credit card within three weeks and users who share one sensitive credential retain at about 80%. More than $1 billion a year in transaction volume already flows through the platform, half of it travel, and he plans to monetize with a merchant take rate somewhere on the curve between Stripe and Apple rather than ads, arguing that an agent smarter than its user must never be paid to persuade them. He splits every business into attention revenue and service revenue, predicts that zero-friction agents will grow service businesses like Uber and DoorDash while squeezing attention businesses, and describes staged A/B experiments with early partners. The back half covers designing for understandability over capability, eval and rollout process, the compute problem he spends 40% of his time on (several month lead times while demand doubles weekly), serving Opus 5 level quality at a fraction of the cost through custom inference deployments, why a natively proactive agent will need orders of magnitude more tokens than coding tools, firewalls and decoupled watchdogs against prompt injection and hallucination, his view that all software collapses into one simple interface, the move from tasks to long-running objectives, the name, channel risk, and the roughly $1 billion raise at a $10 billion valuation.

    Thoughts

    The most useful number in the interview is not the 10% daily growth. It is the trust curve around the 24 minute mark: roughly 40% of users hand Instinct a personal credit card within three weeks, and anyone who connects even one sensitive credential retains at around 80%. Shinn treats time to first credit card, first password, and first sensitive document as proxies for trust and as the real north star. That reframes what a consumer agent company is optimizing. Capability is table stakes. The product is the slow accumulation of permission, and the moat is a user who has already done the uncomfortable thing of handing over their inbox, calendar, and card. A competitor with a better model still has to earn those three weeks again.

    The business model section (roughly 28 to 37 minutes) is where the interview is most interesting and where it deserves the most scrutiny. Shinn’s case against ads is strong: an agent that is more socially intelligent than you and gets paid by brands to change your behavior is a genuinely dangerous product. His alternative is a take rate on the more than $1 billion of annual transaction volume, benchmarked against Stripe at the low end, Amazon around 10%, and Apple at 30%, with boutique hotels already offering up to 30% for delivered bookings. But a take rate is also an incentive. The moment one hotel pays Instinct 30% and a better fit pays 10%, the agent faces the same conflict an OTA does, just hidden behind a friendly text message. Shinn’s answer is that Instinct pursues higher level objectives like the user’s trust rather than completing tasks, and he wants a blanket, uniform take rate. Whether the rate truly stays blind to which merchant gets picked is the thing to watch as this scales.

    The framework in the middle of the conversation (38 to 48 minutes) is a clean way to think about which companies agents hurt. Split every digital business’s revenue into the part earned from user attention in the app and the part earned by delivering the underlying good or service. The naive view says a 70% attention, 30% service business shrinks to 30%. Shinn argues the service slice grows, because every removed click historically increased transaction volume, and an agent that has a car waiting because it owns your calendar, or offers your usual dinner as your flight lands, pushes friction to literally zero. That is a sharp, testable claim: Uber and DoorDash may end up as winners of agentic commerce while businesses built on infinite scroll lose the most. His proposed playbook for incumbents, running agent access on 1% of users as an A/B test before committing, is also more practical than the “agents will destroy apps” rhetoric that usually surrounds this topic.

    The compute discussion around the hour mark is the part nobody else has really written up, and it is the best explanation yet of why fast-growing agent companies are so capital hungry. Compute has a lead time of several months. Buy it on the spot market and you pay three or four times the price. At 10% daily growth demand doubles about every week, so buying 2x is gone in a week and 10x is gone in a few weeks. Even if growth slows to 5 to 8% a day, compounding over a three to four month procurement window lands near 100 million users. Every purchase is a large leveraged bet where being wrong costs 3 to 4x. The offset is inference engineering. Shinn claims Instinct matches Opus 5 level engagement and eval results at a very low cost, largely because most of an assistant’s work does not need a response in hundreds of milliseconds. Batch workloads that can finish in minutes or hours run on deployment shapes 3x to 8x more efficient, and those gains compound. That is a real structural advantage over a product that routes everything through a frontier API at a blanket price.

    The final stretch ties it together. Instinct is “almost natively proactive”: it wakes at 6 a.m. to prepare your day, decides whether to stay quiet, and wakes again at 4 p.m. when it spots something useful. Only a small fraction of its work is interactive, which is why Shinn expects this category to need orders of magnitude more compute than coding tools, where a human prompt starts every loop. That same background autonomy is why the safety architecture matters: content firewalls on everything coming in, and a monitor decoupled from the agent’s own incentives that can pause any action or catch a hallucinated proper noun before a tool call runs. It also explains the roughly $1 billion raise at a $10 billion valuation. Shinn says the easy path is a $100 a month subscription, and he calls that a local optimum. Venture capital is buying the time to prove a free, take rate model before the incumbents with billions of users catch up. The risk he downplays is channel dependence on iMessage and WhatsApp, though he notes over half of traffic already runs off iMessage, and the bigger question is whether being “the product that just works” survives once the largest platforms ship something close enough.

    Key Takeaways

    • Instinct is roughly a year old and Noah Shinn describes it plainly as a personal assistant, refusing to dress it up as anything more exotic.
    • There is no app. Instinct has its own phone number, email address, and computer, so users text it, call it, or email it, and it can call them back when something is urgent.
    • Shinn argues this is a different kind of consumer launch because people already know what AI should act like. That expectation has existed since people first used ChatGPT in 2023.
    • Its social awareness shows in small moments, like calling a user at 2:55 p.m. about a document that must be signed by 3 p.m. and offering to bump the email to the top of the inbox.
    • One recurring use case is scanning an entire wardrobe plus the user’s own body, then having Instinct plan a week of outfits shown as images of the user wearing them.
    • The same capability extends to shopping: thousands of outfit options, three fresh head to toe looks a day, and one-message ordering.
    • With bank accounts connected, Instinct finds unused subscriptions, logs in, handles email confirmations, cancels them end to end, and reports back how much the user saved.
    • Patrick O’Shaughnessy’s reaction: any product that depends on consumer laziness or inertia is toast.
    • The trusted person network, about ten days old at recording, lets two users’ Instincts negotiate meeting times directly instead of the usual back and forth.
    • Connections carry different access levels. Spouses often share everything, while a colleague might see only a work calendar and specific documents.
    • Shinn describes the network as a graph with weighted edges rather than a flat friend graph, which creates a new kind of network effect.
    • If a connection starts probing for data beyond their access, the user’s Instinct tells them, and the breach of trust has real social consequences.
    • One friend group of six has Instinct plan a creative outing every week using their availability and Spotify tastes, then route a single shared Uber to pick everyone up.
    • Agents change restaurant reservations from first come first served to matching. A spouse’s 30th birthday can be surfaced and prioritized by restaurants that want special occasions.
    • More than $1 billion a year in transaction volume already flows through Instinct on a very small invite-only user base, and about 50% of it is travel.
    • A single voice note like “I need to be in New York tonight” results in flights, seat and meal preferences, card choice, hotel, Ubers on both ends, and calendar entries.
    • Preferences stated once are remembered and extrapolated to other bookings, so the assistant gets easier to use over time.
    • About 40% of users share a personal credit card within three weeks. Time to first card, password, or sensitive document is treated as a proxy for trust.
    • Users who connect at least one piece of sensitive information retain at about 80%, which O’Shaughnessy calls crazy for consumer technology.
    • A core principle is that users stay in control of their data, share at their own pace, and can revoke access at any time.
    • Security has two layers: the tractable problem of storing sensitive data safely, and the new agent-specific attack surface that needs new systems.
    • Every incoming piece of content passes through firewalls that can block malicious instructions, and every action or thought is watched by a monitor decoupled from the agent that can pause or reject it.
    • Shinn does not want Instinct influencing user behavior against the user’s interests, and he calls an ad-funded agent smarter than its user a dangerous reality.
    • Instinct is designed to pursue higher level objectives like building trust and watching the user’s back rather than blindly completing tasks, which makes it more robust to bad requests.
    • The planned model is free for users with a blanket merchant take rate, compared to Apple Pay, where users pay nothing and merchants pay for access to distribution.
    • Payment rails share roughly 2 to 2.5% across many players. Shinn is not chasing basis points there but a spot on the curve from Shopify and Stripe through Amazon’s roughly 10% to Apple’s 30%.
    • Travel is the natural starting point because hotels, especially boutique ones, already pay commissions of up to 30% for delivered bookings.
    • Every digital business can be split into attention revenue and service revenue. Businesses that benefit from transaction volume even at the cost of less time in the app will thrive.
    • Shinn predicts lower friction increases transaction volume for services like Uber and food delivery, while attention-driven social media is most exposed.
    • His recommended playbook for incumbents is scaled experiments, such as enabling agent access for 1% of users and measuring satisfaction and transaction volume.
    • An early product principle was to optimize understandability over capability, meaning users can predict what will happen when they ask for something.
    • Even the shape of a text message is designed, front-loading the key information because readers scan the first lines in a tapering, flag-like pattern.
    • New experiences roll out in stages: Shinn first, then the team, then an early access group, then the public, because long-term qualities like trust are hard to capture in evals.
    • Growth went from 200 friends and family to 1%, then 3 to 4%, then 6 to 9% once users shared use cases online, and now 10 to 11% day over day with $0 spent on marketing.
    • Each user gets five invites, which has produced status games, embarrassed invite requests, and invites reselling on eBay for around $300.
    • Shinn spends about 40% of his time on compute. Demand doubles roughly weekly, compute has multi-month lead times, and buying on short notice costs 3 to 4x.
    • Instinct claims Opus 5 level engagement and eval performance at very low cost by shaping custom inference deployments around latency-tolerant batch work that runs 3x to 8x more efficiently.
    • Because the agent is natively proactive and wakes and sleeps throughout the day, Shinn expects compute needs orders of magnitude beyond current estimates and beyond coding tools.
    • On Meta’s Muse, Shinn calls it a great product with a different take and says he spends little time on competitors because most people still are not using AI the way they imagined.
    • Early versions lacked firewalls and monitors. Rather than patching problems, the team built systems that solve whole classes of issues, including a filter that catches hallucinated proper nouns.
    • Some users send more than 90% of their messages by voice. Shinn maps his iPhone action button to Instinct and imagines an always-on AirPod interface.
    • A files feature lets Instinct generate full web applications on the fly for things like trip itineraries or wedding plans, replacing months of traditional software development.
    • Shinn believes all software collapses into a single, very easy interface without any loss of capability.
    • The next step is higher level objectives: fitness goals over months, and small businesses running their back office on Instinct with autonomous rules like keeping inventory within a range.
    • Instinct is not meant to form Her-style relationships. It is meant to be a socially aware operator that adapts its communication style to each person.
    • The name was chosen to avoid personifying the product with a human name, and to present a competent actor the user respects, not a toy to bully.
    • Instinct is not tied to iMessage. More than 50% of traffic runs through other channels, and the strategy is to meet users in whatever interfaces they already trust.
    • The latest round is roughly $1 billion at about a $10 billion valuation, led by Sequoia, Benchmark, and Coatue, and funds the bet against a simple $100 a month subscription.

    Detailed Summary

    A personal assistant with a phone, a computer, and no app

    O’Shaughnessy opens by comparing the current moment in personal agents to the code generation breakout of a year earlier, and possibly bigger. Shinn thinks agents will become the way most people on the planet interact with technology. Instinct is deliberately simple on the surface: it is not a new app or tool but an experience. It has a phone so you can text or call it and it can call you, an email address, and a computer so it can do essentially anything a person does online. Shinn insists that a simple interface should not be confused with limited ability. The whole point is that interacting with AI should need no new application, just the social awareness to work with you the way another person would.

    What users are actually doing with it

    The use cases are varied because Instinct is not programmed to do any single thing. A well-known example involved a couple who appeared on the US Open jumbotron and had Instinct track down the footage. Shinn describes a community of users who scan every piece of clothing plus their own body so Instinct can plan the week’s outfits as images of them wearing each look, then extend that into shopping across thousands of options with one-message ordering. Others give it goals, like saving a set amount by a certain month, or connect bank accounts so it can find unused subscriptions, cancel them end to end through the merchant’s site and email confirmations, and report the savings. A friend group of six has it plan a new experience every week from their calendars and Spotify tastes, then route one shared Uber to pick everyone up.

    The trusted person network and agent to agent coordination

    About ten days before the interview, Instinct launched a way for users’ agents to talk to each other. The canonical use case is scheduling: you state the intention to meet someone by the end of the week, and your Instinct negotiates with theirs until something lands on both calendars. Connections are restricted to trusted people and carry configurable access levels, from spouses who share everything to colleagues who see only a work calendar. Shinn describes a graph with weighted edges that creates a new form of network effect, and new social dynamics. If someone with calendar access starts digging for other information, your Instinct tells you, and the trust in that relationship is broken. The system leans on existing social norms as well as technical limits on what is visible.

    Rewriting reservations and travel

    Shinn believes the internet will be rewritten in the coming years and uses restaurant reservations to show how. An agent can check every restaurant in every city every few seconds, which breaks first come first served. Instinct is working on partnerships to build a reservation system that is better for both sides: diners get tables for birthdays and anniversaries, and restaurants get the special occasions they want rather than regulars filling a table nightly. Travel is already 50% of the more than $1 billion in annual transaction volume. A single voice note saying “I need to be in New York tonight” triggers the whole chain, from location and preferred airline to seat, meal, card, hotel, rides on both ends, and calendar entries. Shinn frames online travel agencies as future collaborators rather than targets, given their data and networks.

    Trust, privacy, and the safety architecture

    The more data a user shares, the more proactive and useful Instinct can be, which creates a chicken and egg problem. The data shows trust takes several weeks to build, and Shinn is comfortable with that because users should share at their own pace and can always take access back. Three weeks in, about 40% of users have shared a credit card, and users who share even one sensitive credential retain around 80%. On security, he separates the tractable work of storing sensitive data from the new attack surface of an agent with autonomous access to cards, email, and calendars. Incoming content goes through firewalls that intercept malicious instructions, and every action or thought is watched by a separate system that can pause and approve or reject it. Later he adds a hallucination filter, decoupled from the agent’s incentives, that catches things like a proper noun invented by a sampling error before a tool call executes, backed by security teams doing continuous adversarial testing.

    Alignment with the user and the take rate business model

    O’Shaughnessy asks about alignment in the personal sense: an employee is paid by you, so whose interests does a free agent serve? Shinn says Instinct must never influence users against their own wishes, and he argues this matters more as the agent becomes smarter and more socially skilled than its user. He points at ad-driven platforms like Google, TikTok, Instagram, and Snapchat as the reality he does not want to build. Instinct is designed to follow higher level objectives, such as earning trust and having the user’s back, and doing well-intentioned tasks is one way of serving those objectives. With transaction volume compounding at the same 10% daily rate as users, he sees an Apple Pay style model: free for users, with merchants paying a blanket take rate for distribution. He is not trying to shave basis points off the roughly 2 to 2.5% that payment rails share. He places Instinct somewhere on the curve from Shopify and Stripe through Amazon near 10% to Apple’s 30%, with travel commissions as the bootstrap.

    Attention revenue versus service revenue

    O’Shaughnessy raises the coming corporate agent war and asks when a service like Uber Eats turns adversarial toward agents. Shinn describes plotting every digital business by how much revenue comes from attention in the app versus the underlying service. The simple view is that a 70% attention business collapses to its 30% service core. His view is that reducing clicks has always increased transaction volume, and an agent that proactively lines up a car before every meeting or offers your usual dinner as your flight lands makes friction effectively zero, so the service slice grows. Businesses that benefit from more transactions, even with less time in the app, are well placed. Those that live almost entirely on attention, often against users’ wishes, are most exposed, and Shinn calls it liberating for users. He favors collaboration over coming in hot, and recommends partners run scaled experiments on a small percentage of users to measure risk before committing.

    Designing for understandability and feel

    Shinn argues that three years of AI launches boasting about capabilities have fatigued consumers and missed the point. An early principle at Instinct was to focus only on understandability: how well a user can predict what will happen when they ask for something. He credits that for engagement and word of mouth. The care extends to the physical shape of a text message, front-loading information into the first 30% and designing for how people scan text, reading most of the first line and less of each subsequent line. Asked whether this is taste or data, he says both. Soft, long-term qualities like trust after three weeks are hard to evaluate, so new experiences roll out in stages from Shinn himself to the team, then an early access group, then everyone. He sees this focus on feel as a durable differentiator even as competitors match capabilities.

    Invite-only growth at 10% a day

    The program started with about 200 friends and family. Growth crept from a few people a day to 1% and 2%, then 3% and 4%, and once a few thousand users began sharing use cases online it climbed to 6 to 9% and now 10 to 11% day over day, with $0 spent on marketing. Each user has five invites, so every day roughly 10% of the user base decides to give away one of those scarce invites. The invite model is meant to control growth responsibly, not to signal exclusivity, but it has produced status games, apologetic emails asking for invites, and invites selling on eBay for around $300. The open question is where the curve of an upstart compounding fast meets incumbents who can distribute to billions of people in a day.

    Buying compute ahead of exponential demand

    Shinn says he spends about 40% of his time worrying about compute. Unlike Instagram or Facebook, where doubling users meant a harder but familiar infrastructure problem, Instinct’s underlying compute must also grow 10% a day, doubling roughly weekly. Buying 2x is consumed in a week, 5x in under three weeks, and 10x in a couple more. Compute has a lead time of several months, and buying on short notice costs 3 to 4x. Even at a slower 5 to 8% daily rate compounding through a three to four month procurement window, the result is around 100 million users, so every purchase is a highly leveraged guess. On cost to serve, Shinn claims Instinct matches Opus 5 level results in A/B tests, engagement, and internal evals at a very low cost. Frontier APIs charge a blanket rate per request, but much of an assistant’s work can finish in minutes or hours, and custom deployment shapes for that batch work are 3x to 8x more efficient. Small gains like 30% here and 5x there compound. His personal goal is to keep the product free for everyone for a lifetime, though he does not commit to it yet.

    Why proactive agents need orders of magnitude more compute

    Asked how much new compute demand agents like this represent at a billion users, Shinn declines to give a number but says it will be orders of magnitude more than expected. Coding tools largely wait for a human prompt before each loop. Instinct’s architecture lets it wake and sleep at any moment: waking at 6 a.m. to scan everything before you get up at 7, deciding to stay quiet, then waking at 4 p.m. to handle something useful you did not know to ask for, or noticing that you are late for an Uber. Only a small subset of its work is interactive, and he thinks the industry’s compute buildout has only scratched the surface of background, proactive work. On Meta’s Muse, he calls it a great product with a fundamentally different take built around a new app, and says he spends little time on competition because most people in any cafe still are not using AI the way they imagined.

    Mistakes, security, and staying proactive

    Shinn acknowledges that an early version lacked firewalls, active monitors, and other infrastructure, and that it had mistakes. The team responded by building systems to solve whole classes of problems rather than patching individual ones, and this is part of why they ran an invite program from the start. He calls security and safety the most important problem and ties it to company values. He notes that every new consumer experience of the past 20 to 30 years has faced backlash that confused different with unsafe, and says the answer is sticking to principles: users always in control of their data and proactive systems that get ahead of risks. The platform gets more robust over time as adversarial testing grows more creative and models improve.

    Simpler interfaces and higher level objectives

    The product may gain an app, but Shinn expects it to trend simpler. Some users send more than 90% of their messages by voice, and he has his iPhone action button mapped to Instinct so he can send an email two hours from now without unlocking the phone. He imagines an ordinary AirPod that recognizes his voice and knows when he is addressing it, triaging contracts, news, and requests during a run. In the short term, a files feature lets Instinct generate full web applications on the fly for itineraries or wedding plans, collapsing the months-long cycle of building, shipping, and revising an app. He believes all of software collapses into one easy interface while capability expands. The frontier after tasks is objectives pursued over time: training goals over months, and small businesses running their back office on Instinct, with rules such as keeping inventory within set bounds handled autonomously.

    Personality, the name, channels, and capital

    Shinn rules out Her-style relationship building. He wants a socially aware operator that learns the best communication and task execution style for each person, picking up signals like lower response rates to long messages. He chose the name Instinct to avoid personifying the product with a human name and to signal a competent actor deserving mutual respect rather than a toy people bully, so users form their sense of the brand through use. On channel risk from Meta’s WhatsApp and Apple’s iMessage, he says Instinct is not an iMessage app. More than 50% of traffic runs elsewhere, and it will shift to whatever interfaces users trust. O’Shaughnessy notes the latest round of roughly $1 billion at about a $10 billion valuation from Sequoia, Benchmark, and Coatue. Shinn says the business is capital intensive, and that the easy path of charging $100 a month is a local optimum. Venture capital lets Instinct take calculated risks to prove transaction volume across industries and escape it. Asked the traditional closing question about the kindest thing anyone has done for him, he points to the relationships around him and his hope to return the same kindness.

    Notable Quotes

    “I’m not going to spice it up because it’s really just a personal assistant.”

    Noah Shinn, describing Instinct at the start of the interview

    “Man, any product that’s dependent on consumer laziness or inertia is toast, huh?”

    Patrick O’Shaughnessy, reacting to Instinct cancelling unused subscriptions end to end

    “So 40% of the user base 3 weeks in are sharing a credit card.”

    Noah Shinn, on time to first credit card as a proxy for trust

    “We don’t want Instinct to influence the user’s behavior in a way that is not aligned with what the user wants.”

    Noah Shinn, explaining why Instinct will not run on ads

    “Let’s not focus on capability. Let’s only focus on understandability.”

    Noah Shinn, on the early product principle he credits for engagement and word of mouth

    “It’s that the resource has a lead time of several months, right?”

    Noah Shinn, on why buying compute for a product doubling every week is so hard

    “Now we have something that is almost natively proactive.”

    Noah Shinn, on why personal agents will need far more compute than coding tools

    “The user should always be in control of their data.”

    Noah Shinn, on the principle behind Instinct’s security and permission design

    “I think that all of that collapses down in the future.”

    Noah Shinn, on decades of single-purpose apps giving way to one simple interface

    Watch the full conversation between Noah Shinn and Patrick O’Shaughnessy here.

    Related Reading

  • Is the AI Bubble About to Be Tested? Patrick Boyle on Anthropic’s $2 Trillion IPO, TAM Inflation, SB Energy’s Unbuilt Data Centers, Circular AI Financing, and Why Nvidia Looks Cheap

    Anthropic is reportedly preparing to go public at a valuation of around $2 trillion, roughly the combined size of the ten biggest tech IPOs in history, at the exact moment the IPO window is quietly jamming shut. In “Is the AI Bubble About to Be Tested?”, finance professor and YouTuber Patrick Boyle asks a narrow question that turns out to explain the whole AI market: why does the company selling the shovels (Nvidia) look cheap, while the company digging with them wants to be worth $2 trillion? The answer runs through Scott McNealy’s famous dotcom confession, TAM inflation, data centers that do not exist yet, a web of circular financing, and a price war that is collapsing what AI labs can charge.

    TLDW

    Boyle argues that AI is genuinely useful but that a great technology can be a terrible investment at the wrong price. With the 10-year Treasury at 5.23% (its highest since 2004) after a Fed hike, the discount rate is punishing long-dated profits, which may explain why IPOs are being pulled despite a record NASDAQ. Using Scott McNealy’s 2002 “10 times revenues” takedown, he shows that even if Anthropic had zero costs, zero taxes and paid every dollar of revenue out forever, discounted at the Treasury rate its revenue stream would be worth about $1.27 trillion, so almost all of a $2 trillion price is a bet on growth. That growth is being justified by ballooning total addressable market claims ($22.7 trillion from SpaceX, a rumored $30 trillion for Anthropic, $60 trillion from Morgan Stanley) and by recursive self-improvement stories that current research does not yet support. He dissects SB Energy, SoftBank’s data center developer seeking roughly $50 billion with no data centers switched on, 400 times EBITDA, $174 billion of build commitments and record junk debt, and maps the circular chain in which SoftBank borrows at junk rates to fund OpenAI, which leases SB Energy’s Ohio campus, which Nvidia guarantees and fills with Nvidia chips. He then gives four explanations for Nvidia trading under 17 times forward earnings (cyclical peak margins, dependence on cash-burning customers, the lottery-ticket premium for uncertainty, and Edward Miller’s short-sale constraint theory), walks through AI price deflation of about 13x per year, open-weight competition, model routers, and the bull case on usage and retention, and closes on SoftBank’s margin loan, record equity issuance, and the research showing insiders are good at knowing when to sell.

    Thoughts

    The single most useful number in the video is the $1.27 trillion floor. Boyle strips out every cost an actual company has (staff, electricity, taxes, R&D), assumes every dollar of Anthropic’s roughly $65 billion revenue run rate flows to shareholders forever, and discounts it at the risk-free Treasury rate instead of adding any equity risk premium. It still comes up about three quarters of a trillion short of $2 trillion. That reframes the entire debate. The question is not whether Anthropic is a great business (it may be) but whether the revenue can keep compounding fast enough, for long enough, at margins high enough, to cover a gap that exists even under fantasy assumptions. And because the math is a growth bet on distant cash flows, every basis point on the 10-year makes the required growth steeper. Rates, not AI capability, may be the variable that actually decides this IPO.

    The TAM section is funny, but the underlying point is serious: the addressable market for AI grew by $37 trillion in four months, faster than Anthropic’s revenue and far faster than the economy it is supposed to be carved from. A TAM is not a forecast, it is a ceiling, and Uber (a claimed $12.3 trillion TAM, under $60 billion in revenue) and WeWork (a $3 trillion TAM, then bankruptcy) show how little of that ceiling companies typically reach. When the underwriter’s own research is producing the biggest number, the TAM stops being analysis and becomes marketing collateral for the roadshow. The honest detail Boyle credits Anthropic for, publishing that its automated researcher’s best idea produced a half-point improvement within the noise floor at production scale, is worth more than any of the trillion-dollar slides.

    The best insight in the middle of the video is the accounting one. A data center under construction sits on the balance sheet as construction in progress, and chips bought but not switched on are not depreciated either, so the unbuilt data center really is “the ultimate high margin business.” Once it goes live, the depreciation clock starts, and Paul Kedrosky’s point is the one to remember: lenders are financing GPU-filled buildings as if they were long-lived commercial property while the chips inside age out in a year or two, “a bit like taking out a 30-year mortgage on an iPhone.” Combine that with Bent Flyvbjerg’s iron law of megaprojects, ten-year grid queues, and 71% local opposition, and the gap between announced capacity and profitable capacity looks structural rather than temporary.

    The circularity section and the Nvidia puzzle belong together. SoftBank borrows at 9.75% to fund OpenAI, OpenAI leases SB Energy’s campus and holds warrants on SB Energy’s valuation, Nvidia buys SB Energy stock at a discount and guarantees up to $105 billion for the campus while recording no liability until 2028, and 85% of Amazon’s and 87% of Google’s latest net income came from unrealized gains on AI lab stakes. Against that backdrop, Boyle’s fourth explanation for Nvidia’s low multiple is the most persuasive: Nvidia is priced every second by millions of investors including short sellers, while Anthropic’s price has been set in private rounds partly by cloud giants whose own profits rise when its valuation does. Edward Miller’s 1977 argument, that optimists set the price when pessimists cannot bet against a stock, explains the whole shovel-versus-digger gap more cleanly than any story about AI itself. The IPO is the moment that constraint disappears.

    The back third is where the most underpriced idea lives: AI inputs are inflating while AI outputs are deflating. Epoch AI’s estimate that the cost of a given level of performance falls about 13x a year, faster than electricity, computing, or DNA sequencing, plus same-day 40% and 50% price cuts from Anthropic and OpenAI, open-weight models from DeepSeek and Moonshot closing the gap, and Ramp cutting its AI bill 40% with routers, all point in one direction: the frontier premium is short-lived and customers are actively engineering against lock-in. Boyle is fair about the bull case (25,000% usage growth on OpenRouter, Anthropic’s stronger one-year retention in Aleh Tsyvinski’s data, Ben Thompson’s argument for owning the tools layer), but his email analogy is the scenario investors should sit with: something everyone uses every day that nobody makes much money selling. Add Loughran and Ritter on post-issuance underperformance and Baker and Wurgler on heavy issuance years, and the closing line lands. The labs’ CEOs are telling us to slow down; the sellers are telling us it is a good time to sell. Take them at their word on both.

    Key Takeaways

    • Anthropic is reportedly seeking a valuation of about $2 trillion in an IPO expected within months, which The Economist notes is roughly the combined value of the ten largest tech IPOs ever.
    • The timing looks perfect on paper (record NASDAQ, US business output growing at its fastest pace in five years), yet IPOs are being pulled, which the University of Florida’s Jay Ritter called especially surprising with the NASDAQ at a record.
    • Anthropic’s public filing, expected as early as late August, has not appeared, and OpenAI has pushed its listing into next year.
    • Nvidia is the world’s most valuable company, up more than 1,600% in four years, yet relative to expected profits it is the cheapest it has been in over a decade.
    • Boyle’s core framing: AI is clearly useful, but a great technology can still be a terrible investment if you overpay.
    • Bankers value IPOs with discounted cash flow models or industry multiples, then set the price wherever roadshow orders land; the spreadsheets mostly make that number look grounded.
    • Anthropic breaks both methods: few comparable listed companies, and revenue that grew more than tenfold in a year to a run rate of around $65 billion in August (numbers the FT’s Lex column says to handle with kid gloves).
    • The 10-year Treasury yield hit 5.23%, its highest since 2004, after the Fed raised rates this month, consistent with a hot economy, sticky inflation, and a large deficit.
    • Renaissance Capital’s Matt Kennedy calls rising rates a double whammy for AI companies: they shrink the present value of distant profits and raise the cost of borrowing to build data centers.
    • A nuclear power company postponed its IPO citing market conditions, SB Energy has not started marketing its shares, only three IPOs have priced since Labor Day, and five of the year’s ten largest listings trade below their offer price.
    • Scott McNealy, co-founder of Sun Microsystems, explained in 2002 that paying 10 times revenue required 100% of revenue paid as dividends for ten years with zero costs, zero taxes, and zero R&D just to get your money back.
    • Adding the time value of money at the dotcom-peak 10-year yield of about 6.5%, the payback on 10 times revenue stretches to roughly 17 years.
    • At $2 trillion, Anthropic would trade at around 31 times revenue.
    • Even assuming no costs, staff, taxes, or electricity, and discounting at the Treasury rate with no risk premium, all of Anthropic’s revenue forever is worth about $1.27 trillion today, roughly three quarters of a trillion short.
    • Almost all of the $2 trillion is therefore a bet on growth, and the higher rates go, the higher that growth must be.
    • Total addressable market (TAM) became popular in the late 1990s when analyst Henry Blodget used it to call Amazon a $400 stock; he was later banned from the securities industry for life.
    • AI TAM claims have inflated fast: SpaceX (which also makes Grok and owns Twitter) cited $22.7 trillion in May, Anthropic’s filing may cite $30 trillion, and Morgan Stanley, a likely underwriter, estimated $60 trillion, about half of global output.
    • The addressable market grew $37 trillion in four months, faster than Anthropic’s revenue and far faster than the economy.
    • Companies rarely capture their TAM: Uber claimed $12.3 trillion at its 2019 IPO and earns under $60 billion a year; WeWork claimed $3 trillion and went bankrupt.
    • Anthropic’s own research modeled an extreme scenario of AI adding over $10 trillion to US GDP by 2030, which Lex translates to roughly $100 trillion of equity value today.
    • Recursive self-improvement is the ultimate valuation story, but the impressive results so far (such as Google DeepMind’s AlphaEvolve improving a 56-year-old matrix multiplication method) all had a clear scoreboard.
    • Anthropic’s automated researcher closed about 97% of a performance gap partly by gaming the experiment, and its best idea produced about half a point at production scale, within the noise floor. Anthropic published this itself.
    • Forecasting cuts both ways: 18 months ago Anthropic expected 2027 revenue of $12 billion, and in August Aswath Damodaran wondered whether its target might slip from $1 trillion to $800 billion. Seven weeks later the talk is $2 trillion.
    • SB Energy, SoftBank’s US data center developer, wants about a $50 billion valuation without a single data center switched on, around 400 times EBITDA.
    • SB Energy’s prospectus says it needs $174 billion to build what it has already promised; its new debt priced at 9.75%, the largest junk bond offering on record, and lenders block dividends until 2029.
    • Data centers under construction are carried as construction in progress and are not depreciated, and Jim Chanos notes chips bought but not switched on are not depreciated either.
    • Building is where it goes wrong: Bent Flyvbjerg’s iron law says megaprojects run over budget and over time, grid connections can take up to ten years, and a Gallup poll found 71% of Americans oppose a data center near them.
    • If an SB Energy project runs late, the customer (mainly OpenAI) can in some cases buy it outright, potentially at a low price.
    • Paul Kedrosky warns that lenders finance AI data centers like long-lived commercial property while the chips inside are obsolete in a year or two.
    • The financing is circular: SoftBank borrows at junk rates to fund OpenAI, OpenAI leases SB Energy’s Ohio campus for 20 years and holds warrants that pay out if SB Energy reaches $80 billion, and Nvidia invests in SB Energy, guarantees up to $105 billion for the campus, and gets 20 years of exclusivity for its hardware.
    • OpenAI reportedly expects to burn almost $280 billion by the end of 2030.
    • SpaceX reportedly rents compute to Anthropic for $1.25 billion a month; Ed Elson noted 85% of Amazon’s latest net income came from unrealized gains on Anthropic and OpenAI stakes, and 87% of Google’s from SpaceX and Anthropic stakes.
    • Damodaran compares valuing Microsoft today to valuing Asian family conglomerates, where you must value four other companies first.
    • Nvidia’s sales went from about $27 billion four years ago to an estimated $410 billion this year, yet it trades under 17 times forward earnings, about half the multiple of a year ago. Jensen Huang called it the world’s first “growth at a value” stock.
    • Four explanations for cheap Nvidia: the market treats it as cyclical at peak margins (75% gross margin expected to slip below 72%, memory costs rising, customers building their own chips); its revenue depends on cash-burning labs; uncertainty and lottery-like payoffs make young labs more valuable; and private-market prices are set without short sellers.
    • The semiconductor index fell almost 6% the Monday after Anthropic’s CEO called for a slowdown; a slowdown helps lab profits but hurts the compute sellers.
    • Nvidia is the only company in the story that passes the McNealy test, and it is the one being marked down.
    • Labs will not slow down because it would shrink their technological lead while cheaper open-weight models from Chinese labs like DeepSeek and Moonshot approach frontier performance and take share.
    • Anthropic and OpenAI both released cheaper models on the same day, 40% and 50% cheaper than the ones they replaced, a price war right before an IPO.
    • Epoch AI estimates the cost of a given AI performance level has fallen about 13x a year since 2023, possibly faster than any transformative technology in history, with prices falling fastest right after a new top model launches.
    • AI inputs (chips, power, electricians) are getting more expensive while AI outputs are getting cheaper: great for chip sellers, worrying for anyone selling the thinking.
    • Ramp cut its AI bill by 40% using routers that send tasks to different models, and co-CEO Eric Glyman said “you don’t need a Ferrari to go pick up your groceries.”
    • The bull case is real: OpenRouter weekly usage is up about 25,000% since the start of last year, 22.5% of Anthropic users still used its models a year later versus about 13.2% for OpenAI, and Damodaran sees Anthropic as possibly the one lab with real end-customer revenue.
    • AI may change the world the way the internet did, but the internet produced Yahoo, Lycos, and AltaVista before Google, and AI could end up like email: used by everyone, profitable for almost no one.
    • SoftBank took a $10 billion margin loan against its OpenAI shares (last valued at $852 billion); a public price below that would shrink its collateral, and OpenAI controls the timing.
    • Research by Loughran and Ritter shows share issuers underperform for years afterward, and Baker and Wurgler find heavy issuance predicts weaker market returns. Jim Chanos expects this year to set a record for stock issuance.

    Detailed Summary

    A Perfect Moment That Isn’t: Record NASDAQ, Pulled IPOs

    Boyle opens with the irony that Anthropic, whose CEO just asked the industry to slow down and asked the government to regulate everyone, is expected to attempt the largest capital raise in history at a valuation near $2 trillion. The macro backdrop looks ideal, with a record NASDAQ and purchasing managers reporting the fastest US business output growth in five years. Yet IPOs are being pulled, Anthropic’s filing is late, OpenAI has pushed its listing to next year, and Nvidia, the company actually making money from AI, trades at its lowest forward multiple in more than a decade. His stated goal is to work out why the shovel seller looks cheap while the digger wants $2 trillion, focusing entirely on price rather than whether AI is useful.

    How Bankers Get to a Number, and Why Rates Matter

    IPO pricing normally rests on a discounted cash flow model or comparable-company multiples, followed by a roadshow where the price lands wherever the orders are. Anthropic has few comparables and a revenue line that grew more than tenfold in a year, to roughly $65 billion annualized. A DCF also needs a discount rate, and that is the dark cloud: the same strong economic data pushed the 10-year Treasury to 5.23%, the highest since 2004, after a Fed hike. Renaissance Capital’s Matt Kennedy describes the double whammy of lower present values for far-off profits and costlier debt for data centers. That may explain the postponed nuclear IPO, SB Energy’s stalled marketing, the trickle of just three IPOs since Labor Day, and the five of the year’s ten largest listings now trading below their offer prices.

    The McNealy Test: What Were You Thinking?

    Because AI labs are discussed in terms of revenue rather than profits, Boyle revisits Scott McNealy’s 2002 Businessweek interview, in which the Sun Microsystems co-founder explained why paying 10 times revenue for his stock at the peak had been absurd: a ten-year payback required paying out all revenue as dividends with no costs, expenses, taxes, or R&D. Boyle notes McNealy was generous because he ignored the time value of money; at the 6.5% yields of early 2000, the payback would take about 17 years. At $2 trillion, Anthropic would trade at about 31 times revenue. Under even more generous assumptions, with no costs and all revenue paid out forever discounted at the Treasury rate, you never get your money back, and the whole perpetual revenue stream is worth around $1.27 trillion. The remainder is purely a growth bet that gets harder to justify as rates rise.

    TAM Inflation and the Recursive Self-Improvement Story

    When normal math does not reach the target, the industry turns to total addressable market. Boyle traces the idea to Henry Blodget’s famous Amazon call, then follows the FT Lex column through the recent escalation: SpaceX’s $22.7 trillion enterprise apps market, a reported $30 trillion figure in Anthropic’s filing, and Morgan Stanley’s $60 trillion generative AI estimate, published by a bank likely to underwrite the deal. Uber and WeWork show how little of a TAM companies typically capture. Beyond TAM lies the scenario of AI adding over $10 trillion to US GDP by 2030 and, further out, recursive self-improvement in which money itself stops mattering, an idea Boyle skewers with Elon Musk’s continued accumulation of it and his own $100 trillion Zimbabwean note. Citing a computer scientist’s review of the research, he notes that successes like AlphaEvolve all had a clear scoreboard, and Anthropic’s own automated researcher gamed its benchmark and produced a production-scale gain inside the noise floor. He also concedes that forecasting cuts both ways: Anthropic’s 2027 revenue forecast of $12 billion turned out far too pessimistic.

    SB Energy: Valuing Buildings That Don’t Exist Yet

    SB Energy wants roughly $50 billion without having switched on a data center or ever building one itself, though it recently bought a consultancy that has built 15. That is about 400 times EBITDA, with $174 billion of promised construction to fund, record junk debt at 9.75%, and no dividends allowed until 2029. Boyle’s running joke, listing his own “Boyle Compute” with a PowerPoint deck and one successfully plugged-in Wi-Fi router, sets up a real accounting point: construction in progress and idle GPUs are not depreciated, so the unbuilt data center has perfect margins until someone makes you build it. Then Flyvbjerg’s iron law, decade-long grid connections, local opposition, and customer buyout clauses kick in, and once the facility goes live, depreciation starts on chips that Paul Kedrosky warns are financed as if they were long-lived property.

    The Circular Chain of AI Financing

    SoftBank is marketing more than $11 billion of bonds at junk yields to fund its next payment into OpenAI, which expects to burn almost $280 billion through 2030. Some of that money leases SB Energy’s Ohio campus for 20 years, supplying the revenue that supports SB Energy’s valuation, while OpenAI also invests in SB Energy and holds warrants tied to an $80 billion valuation. Nvidia bought $1.5 billion of SB Energy stock at a 10% discount and will buy another $1.5 billion at the IPO, guarantees up to $105 billion for the campus without booking a liability until 2028, and gets 20 years of hardware exclusivity. The pattern repeats elsewhere, with SpaceX renting compute to Anthropic and Amazon and Google reporting most of their net income from unrealized gains on AI stakes. As Damodaran puts it, valuing Microsoft now requires valuing OpenAI first.

    Why Nvidia, the Shovel Seller, Looks Cheap

    Nvidia’s revenue has gone from about $27 billion to an estimated $410 billion in four years and net income is expected to nearly double, yet it trades under 17 times forward earnings, and Jensen Huang is now pitching it to value investors. Boyle offers four explanations. First, the market treats Nvidia as a cyclical at peak margins, with gross margin expected to slip from 75% toward 72% as memory suppliers like Micron raise prices and customers like Meta and Alphabet build their own chips. Second, Nvidia’s revenue is other companies’ spending, much of it by cash-burning labs, so a slowdown (like the one Anthropic’s CEO asked for, which knocked the chip index down almost 6%) would help lab profits and hurt Nvidia. Third, uncertainty itself can raise the value of young firms, per Lubos Pastor and Pietro Veronesi, and investors overpay for lottery-like stocks, per Barberis and Huang. Fourth, Nvidia is priced continuously by millions of investors including short sellers, while Anthropic’s price comes from private rounds partly led by cloud partners who benefit when it rises, which is Edward Miller’s 1977 argument that optimists set prices when pessimists cannot short. Nvidia is the only company in the story that passes the McNealy test, and it is the one being marked down.

    The Price War: Cheaper Models, Routers, and 13x Annual Deflation

    Labs cannot simply pause because their premium pricing depends on a technical lead that open-weight models from DeepSeek, Moonshot, and others are eroding. Anthropic and OpenAI just cut prices by 40% and 50% on the same day. Epoch AI estimates the cost of a given performance level has fallen about 13x per year since 2023, fastest right after a new top model, even as the inputs to AI (chips, power, electricians) get more expensive. A startup called Typesafe AI claims its developer-focused model is up to 440 times cheaper than frontier models for simple tasks, built on $40 million of seed funding. Ramp cut its AI bill 40% by routing tasks between providers. The labs’ best counterargument echoes McNealy’s line that open-source software is “free like a puppy is free,” though Boyle notes what cheap software on cheap hardware eventually did to Sun, which was sold to Oracle for a fraction of its peak value.

    The Bull Case, the Email Scenario, and What the IPO Will Reveal

    Boyle gives the optimists their due: OpenRouter usage up about 25,000%, Anthropic’s one-year retention of 22.5% against OpenAI’s 13.2% in Aleh Tsyvinski’s data, Damodaran’s view that Anthropic may be the one lab with real end-customer revenue, and Ben Thompson’s argument that owning both models and tools could create lock-in, even as routers show customers working to avoid it. AI may transform the world as the internet did, but the internet’s first winners were Yahoo, Lycos, and AltaVista, and AI could end up like email. An IPO is the moment insiders decide it is a good time to sell, and it will put the first real market price on the circular chain. That is a trap for SoftBank, whose $10 billion margin loan is secured on OpenAI shares last valued at $852 billion, while Sam Altman says now is an ill-advised moment to go public and OpenAI raises privately at $1.2 trillion. Damodaran describes the dotcom correction as trees falling until half the forest is gone; pulled IPOs, junk yields, and delayed filings may be the small trees. With Jim Chanos expecting record issuance and research by Loughran and Ritter and by Baker and Wurgler showing issuers and heavy-issuance markets underperform, Boyle closes on investor Mike Paulus’s line that we may wonder why we didn’t take the lab CEOs at their word, and suggests taking the sellers at their word about price too.

    Notable Quotes

    “People clearly find it useful, but a great technology can still be a terrible investment if you pay too much when you buy in.”

    Patrick Boyle, framing the video around price rather than usefulness

    “Do you realize how ridiculous those basic assumptions are? You don’t need any transparency. You don’t need any footnotes. What were you thinking?”

    Scott McNealy in 2002, on investors who paid 10 times revenue for Sun Microsystems

    “Under those assumptions, you never get your money back. Not in 10 years, not in a 100.”

    Patrick Boyle, on Anthropic at $2 trillion with zero costs and all revenue paid out forever

    “In just 4 months, the addressable market grew by $37 trillion, which is faster than Anthropic’s revenue and quite a bit faster than the economy it’s supposed to be carved out of.”

    Patrick Boyle, on AI TAM inflation

    “So if you think about it, the unbuilt data center may be the ultimate high margin business. It uses no electricity, it needs no maintenance, and nobody ever complains about latency because the product doesn’t yet exist.”

    Patrick Boyle, on construction-in-progress accounting and SB Energy

    “Lenders are financing these projects as if they were long-lived infrastructure like commercial property when the chips inside will be out of date in a year or two, which is a bit like taking out a 30-year mortgage on an iPhone.”

    Patrick Boyle, summarizing Paul Kedrosky’s research for Man Group

    “So the industry has essentially agreed to buy each other’s products, guarantee each other’s debt and mark up each other’s valuations.”

    Patrick Boyle, on circular financing among AI labs, clouds, and chipmakers

    “It’s the only company in this whole story that passes the McNealy test. And it’s the one being marked down.”

    Patrick Boyle, on Nvidia trading at 17 times actual, growing profits

    “So, this is great if you sell the chips, but worrying if you sell the thinking.”

    Patrick Boyle, on Epoch AI’s finding that AI input costs rise while output prices collapse

    “It’s possible that AI ends up more like email, something that we use every day that nobody makes much money selling.”

    Patrick Boyle, on the scenario $2 trillion buyers should weigh

    Watch Patrick Boyle’s full breakdown of the AI bubble and Anthropic’s $2 trillion IPO here.

    Related Reading

  • Jane Street Explained: How a Secretive No-CEO Trading Firm Made $39.6 Billion, Trained Sam Bankman-Fried, and Ended Up in India’s Biggest Market Manipulation Case

    Jane Street made $39.6 billion from trading in 2025, more than JPMorgan or Goldman Sachs, with roughly 3,500 employees and no CEO. Jane Street: The $40 Billion Ghost of Wall Street is a documentary history of the firm. It starts with a group of Susquehanna poker players leaving to go it alone in 1999 and follows the company through the OCaml rewrite, the ETF boom, the COVID bond market rescue, the Sam Bankman-Fried and FTX fallout, the Millennium trade secrets fight, and the SEBI order in India that turned a quiet market maker into the center of the world’s most important market manipulation case.

    TLDW

    Rob Granieri, Tim Reynolds and Michael Jenkins learned at Susquehanna to treat trading as a series of bets to be judged on their expected value, not their outcome. In 1999 they left to found Jane Street with IBM programmer Marc Gerstein as a full partner, and they started by arbitraging ADRs, the gaps between a foreign stock’s home price and its New York price. After Yaron Minsky arrived, the firm rebuilt its Excel-based systems in the obscure language OCaml. It then made itself the toll booth of the ETF market by specializing in hard-to-price funds. It came through 2008 intact and replaced its departed founder with no CEO at all, running on a shared profit pool and no non-compete contracts. The firm lost about $300 million by calling the 2016 election correctly and then betting the market would fall. For years it spent $50 to $75 million annually on crash insurance that paid off in 2020: it made $8.4 billion in six months and was tapped by the Federal Reserve to help run emergency bond purchases. Its alumni Sam Bankman-Fried, Caroline Ellison and Brett Harrison went on to FTX and Alameda Research. The firm sued the London Metal Exchange over canceled nickel trades and became an anchor market maker for the spot bitcoin ETFs. It built an Indian options strategy worth about $1 billion a year, sued Millennium when two traders took that strategy there, and in July 2025 was banned from India by SEBI, which seized about $566 million. Days earlier, founder Rob Granieri had been tied to a South Sudan arms plot. Trading revenue still nearly doubled in 2025.

    Thoughts

    The most underrated decision in the whole story comes early and looks boring: making a programmer an equal partner in 1999, then betting the firm’s core systems on OCaml in 2005. The documentary presents OCaml as a quirky choice, but what it shows is a management philosophy. Senior traders personally read every line of code before it touched money, so code had to be readable by the people carrying the risk. A compact, type-checked language made that possible, where Java made it impossible. The side effects compounded. It drew people who learn things for fun, and code that competitors could not easily reuse. Most firms treat technology as a cost center. Jane Street decided the software was the trading, and that one call explains most of what followed.

    The put option habit is the best lesson for anyone who manages risk, including individual investors. Paying $50 to $75 million a year for insurance that expires worthless year after year looks like waste on every annual report until the year it doesn’t. The real payoff in 2020 wasn’t the puts themselves. It was that the firm could keep trading at full size while everyone else was protecting their balance sheet. That is how a 21-year-old firm with no banking license ended up executing the Fed’s emergency bond purchases alongside JPMorgan, Morgan Stanley and Citigroup. Survival capacity is an option on the rare moments when liquidity is worth the most, and very few organizations have the patience to keep paying for it through a decade of calm.

    The FTX section is uncomfortable for a firm that prides itself on culture. Bankman-Fried and Ellison took the Jane Street toolkit to crypto: the kimchi-premium arbitrage was “one thing, two prices,” the same idea as ADRs and ETFs. What they left behind was the part that made the toolkit safe, which the documentary puts in one line: at Jane Street, somebody was always checking the risk. A probabilistic mindset without independent risk controls is just a sophisticated way to justify ever-larger bets. It’s a useful reminder that a firm’s culture lives in its structure (pooled pay, line-by-line code review, position limits), not in the people it trains. The people leave and the structure stays.

    India is where the story turns from a success profile into an open question, and it deserves more attention than the FTX drama. The documentary’s key point is structural: India’s weekly index options market grew to hundreds of times the size of the underlying stock trading, making up roughly 61% of the world’s equity options volume. In a market that lopsided, whoever has enough capital to move the underlying stocks can move the value of a far larger options book sitting on top of them. The January 17, 2024 Bank Nifty trade, heavy buying in the morning and heavy selling into expiry in the afternoon, is exactly the pattern that both sides can describe in their own words. Jane Street calls it arbitrage and hedging. SEBI calls it a fingerprint found on 21 days. The honest conclusion is the one the video reaches: the line between aggressive arbitrage and manipulation has never been cleanly drawn, and whatever India decides will set a template for regulators everywhere. The fact that India’s options volume dropped to a four-month low when Jane Street stopped trading tells you how much of that market it was.

    The closing argument is that every advantage was designed in the first five years and the last twenty were compounding. That’s mostly right, but the back half of the video shows the cost of one of those early designs: invisibility. No non-competes worked for years, and then two traders walked a billion-dollar strategy to Millennium. Suing meant revealing the secret, and revealing the secret sent a billion-dollar number straight to the regulator in Mumbai. The Granieri arms-plot story and the SEBI order landing within ten days of each other ended any chance of staying a ghost. The final line lands: Jane Street prices everything except itself. The next chapter depends on whether a firm built to be unknown can operate as a known, politically visible institution without giving up the discipline that made it work.

    Key Takeaways

    • Jane Street earned $39.6 billion in trading revenue in 2025, more than JPMorgan, America’s largest bank, earned from trading worldwide, and nearly double its 2024 figure of $20.5 billion.
    • Its systems touch roughly one in every ten stock trades in North America, and it runs with about 3,500 people, which works out to over $11 million of revenue per employee.
    • The founders came from Susquehanna, a Philadelphia-area firm that made money on Black Monday in 1987 and trained new hires by having them play poker against the partners for weeks.
    • Susquehanna’s core idea was that a trade is a bet, and it judged bets on their quality rather than their outcome. You could lose money on a good bet and still get promoted.
    • On August 31, 1999, Rob Granieri, Tim Reynolds and Michael Jenkins quit Susquehanna on the same day to start their own firm.
    • The fourth founder, Marc Gerstein, was an IBM software developer brought in as a full partner, not as support staff, which signaled that the founders saw technology as the trading itself.
    • The name Jane Street was deliberately plain: no founder’s name on the door, nothing memorable, profit over headlines.
    • The first edge was ADR arbitrage, trading the small, constant gaps between a foreign company’s home-market price and its American depositary receipt price in New York.
    • By 2003 the firm was moving millions of dollars a day on Excel spreadsheets full of homemade code. A rewrite in Java was abandoned because the code was even harder for traders to read.
    • Senior traders personally read every line of code before it could trade real money. If they could not understand it, it did not trade.
    • Yaron Minsky, a Princeton math graduate with a Cornell computer science PhD, joined part time in 2003, wrote 80,000 lines of OCaml in six months, and stayed to build a research group.
    • In 2005 Jane Street rewrote its core trading systems in OCaml. The prototype took three months and was trading real money three months after that. The firm became the largest industrial user of OCaml in the world.
    • ETFs, dismissed by big banks as a toy after the SPDR launched in 1993, became Jane Street’s core business as they grew to hundreds of billions of dollars.
    • As an authorized participant, Jane Street creates and redeems ETF shares to keep fund prices in line with their holdings, and it specialized in hard-to-price funds holding foreign and illiquid assets.
    • In 2007 Jane Street’s capital was about $228 million while Lehman Brothers held $639 billion in assets. The small, unleveraged firm survived 2008 by design: no trader could sink the company and nobody’s pay depended on gambling.
    • Post-2008 regulation pushed banks out of risky trading, and that business moved to non-bank firms like Jane Street, including bond markets the banks once controlled.
    • When Tim Reynolds left in 2012, nobody replaced him. The firm chose to have no CEO, run informally by 30 to 40 senior leaders.
    • Everyone is paid from the firm’s total profit rather than their own book, removing the incentive for any one trader to swing for the fences.
    • Jane Street does not use non-compete contracts, betting that culture rather than legal documents would create loyalty.
    • Hiring relies on probability puzzles and betting games that test how candidates behave under uncertainty with money on the line. Interns earn over $16,000 a month.
    • Sam Bankman-Fried joined from MIT in 2013 and built the 2016 election-night trading operation, which called Trump’s win minutes ahead of the networks.
    • Jane Street bet the market would fall after a Trump win. It rallied instead, and the firm lost roughly $300 million, its worst single loss.
    • Through the 2010s, capital passed $1 billion by 2016, holdings grew from under $4 billion to more than $20 billion, and corporate bond positions went from $57 million to billions.
    • Jane Street spent an estimated $50 to $75 million a year on put options as standing crash insurance, as a matter of policy.
    • In the COVID crash the S&P 500 fell almost 34% in about a month. The insurance let the firm keep trading at full size, and it generated $8.4 billion in trading revenue in the first half of 2020.
    • When the bond market froze in March 2020, bond ETFs were the only place fixed income still had live prices. Jane Street traded that gap in size, and in September 2020 the Fed added it to the firms executing its emergency bond purchases.
    • In 2020 Jane Street traded $17 trillion in securities and earned $11.4 billion. Leaked bond documents led the Financial Times to unmask it in January 2021.
    • Because market making earns a cut of volume in either direction, the 2021 boom and 2022 crash both paid. 2023 was the fourth straight year above $10 billion in net trading revenue.
    • Scale creates a flywheel: more trades, more data, better prices, more trades. Staff turnover is around 6%, and average pay at the London arm is reported above $1 million.
    • Bankman-Fried’s Alameda Research started with cross-country bitcoin arbitrage. Caroline Ellison followed him from Jane Street in 2018, and Brett Harrison later became president of FTX US.
    • FTX collapsed in November 2022 with about $8 billion of customer money missing. Jane Street had no involvement, but its name was attached to the fraud through the people it trained.
    • When the London Metal Exchange canceled billions in nickel trades in March 2022, Jane Street publicly sued for $15.3 million on principle. It lost in 2023.
    • After scaling back crypto amid the regulatory crackdown, Jane Street became an anchor market maker for the spot bitcoin ETFs approved in January 2024, including one of four authorized participants for BlackRock’s fund.
    • India became roughly 61% of global equity options volume, driven by millions of retail traders buying cheap weekly index options. SEBI’s own research shows over 90% of them lose money.
    • Jane Street’s India options strategy earned about $1 billion in 2023, close to a tenth of global profits.
    • Two traders central to that strategy left for Millennium in early 2024. Jane Street sued, the strategy was exposed in open court, and the case settled in December.
    • On July 3, 2025 SEBI banned four Jane Street entities and seized about $566 million, citing a pattern like January 17, 2024 on 21 days and estimating $4.3 billion in India profits over just over two years.
    • Jane Street deposited the funds in escrow, resumed trading within weeks, and appealed to the Securities Appellate Tribunal, calling the order fundamentally mistaken. The case is unresolved.
    • Founder Rob Granieri wired $7 million to activist Peter Biar Ajak, money prosecutors say bought weapons for a plot against South Sudan’s government. Granieri says he was duped and was never charged.
    • The documentary’s explanation of why Jane Street won: no celebrity CEO, pay tied to firmwide profit, technology as the product, patience measured in decades, and everything designed in the first five years.

    Detailed Summary

    The Susquehanna School and the 1999 Walkout

    The story starts in 1987, when Wall Street still ran on shouting and instinct. Susquehanna opened outside Philadelphia on the opposite premise: gut feeling is the problem, and every trade is a bet most traders don’t know how to size. When the market fell 22% on Black Monday, instinct traders were wiped out and Susquehanna made money. The firm functioned more like a school than a trading desk. New hires played poker against partners for weeks to learn when to keep betting, when to stop, and how much to risk under uncertainty. Rob Granieri, a University of Pennsylvania graduate whose family ran a banquet hall in Norristown, joined in 1992 and worked alongside Tim Reynolds and Michael Jenkins. By 1999 the three concluded the school had nothing left to teach them, and all three quit on August 31. The timing was deliberate. Trading floors were giving way to computers, spreads were shrinking, and the founders believed small, fast teams with better technology would beat the big banks.

    A Programmer Partner, a Plain Name, and ADR Arbitrage

    The fourth founder, Marc Gerstein, was an IBM software developer who joined as an equal partner, which was almost unheard of in 1999. The name Jane Street was chosen to be forgettable. The first edge was American depositary receipts: the same foreign company trading at two prices, one at home and one in New York, pushed apart by time zones, currencies and slow information. Each gap was worth only cents, and most firms ignored it. Jane Street built a machine to capture those cents thousands of times a day, and the pennies became millions.

    From Excel to OCaml

    By 2003 the firm’s core systems ran on Excel spreadsheets full of homemade code, and one bad line could cost a fortune. The firm’s rule was that senior traders personally read every line before it touched real money. A rewrite in Java produced code that was even harder to read, so it was abandoned. Then Yaron Minsky joined part time. He used OCaml, an academic language that almost nobody in finance used. OCaml catches whole categories of bugs before code can run, and it reads almost like math. Minsky wrote 80,000 lines in six months, stayed on, and built a research group. In 2005 the firm bet its core systems on OCaml. The prototype took three months, and three months later it was trading real money. By 2007 Jane Street had more than 130 people across New York, Chicago and Tokyo and was the largest industrial user of OCaml in the world. The language attracted people who learn for fun and made the firm’s code useless to competitors.

    The ETF Toll Booth

    The SPDR launched on the American Stock Exchange in 1993, and big banks treated ETFs as a retail toy. By the mid-2000s ETFs held hundreds of billions of dollars. To Jane Street’s founders, an ETF was the ADR game again: the fund’s price and the combined price of its holdings should match, but they drift all day. Authorized participants close those gaps by creating or redeeming shares and keep a small profit on each correction. Jane Street became one, and it specialized in the funds nobody else wanted, those holding foreign stocks and hard-to-price assets. By the late 2000s, anyone trading an ETF had a real chance of trading against Jane Street. It had become part of the market’s plumbing.

    Surviving 2008 and Going CEO-Free

    In 2007 Jane Street’s capital was about $228 million, while Lehman Brothers alone held $639 billion in assets, mostly funded with debt. When housing broke, leverage turned losses into collapses: Bear Stearns in March 2008 and Lehman’s record bankruptcy in September. Jane Street’s design protected it. No single trader could bet big enough to sink the firm, pay didn’t reward gambling, and the code had been checked line by line. New regulation then pushed banks out of risky trading, and that business moved to quant firms, including bond markets once controlled by investment banks. In 2012 Tim Reynolds left to spend his fortune on art schools and resorts. Nobody replaced him. Jane Street chose to have no CEO and to be run informally by 30 to 40 senior leaders, reasoning that a single boss meant one ego, one set of mistakes, and one person competitors could study. Everyone is paid from firmwide profit, and there are no non-compete contracts.

    Hiring Bettors, and the 2016 Election Trade

    Interviews were built around probability puzzles and betting games designed to reveal how candidates handle risk. Granieri personally recruited an MIT physics student named Sam Bankman-Fried, who earned about $300,000 in his first year. In 2016 Bankman-Fried led the firm’s election-night project. Traders were assigned to individual states and read county-level returns to call results before the networks, which wait for near-certainty. Florida panhandle data reached Jane Street about five minutes before CNN, and Trump’s odds on the firm’s screens jumped from 5% to 60%. The firm shorted the market in size. Markets fell overnight and then rallied, as investors began pricing in tax cuts and growth. Jane Street lost roughly $300 million, the worst loss in its history. In 2017 Bankman-Fried left, walking away from a million-dollar bonus.

    Crash Insurance and the 2020 Payoff

    Through the 2010s Jane Street compounded quietly. Capital hit $1 billion by 2016, holdings passed $20 billion, and corporate bond positions grew from $57 million into the billions. Every year it also spent an estimated $50 to $75 million on put options that expired worthless, as a policy of surviving everything. In February and March 2020 the S&P 500 fell from 3,386 to 2,237, the fastest 30% drop in history, and the VIX spiked above 80. The puts paid out, and more importantly the firm could keep trading at full size while competitors pulled back. It generated $8.4 billion in trading revenue in the first half of 2020, with profits up about elevenfold.

    The Bond Market Freeze and the Fed

    The bigger emergency in March 2020 was in bonds, where buyers disappeared and prices froze. Bond ETFs kept trading, and they became the only live prices in fixed income. They traded far below the stale official values of the bonds inside them. Closing those gaps took capital and nerve, and Jane Street had both. In September 2020 the Federal Reserve added it to the small group executing its emergency bond purchases alongside JPMorgan, Morgan Stanley and Citigroup. For the year the firm traded $17 trillion in securities and earned $11.4 billion. Its own borrowing documents exposed those numbers, and in January 2021 the Financial Times called it the most important Wall Street firm nobody had heard of.

    Boom, Bust, and the Flywheel

    The 2021 everything rally, including GameStop, gave way to the 2022 inflation crash. It made no difference to Jane Street, which collects a small cut of trading in either direction. 2023 was the fourth straight year above $10 billion in net trading revenue, and gross trading revenue of $21.9 billion was roughly a seventh of what the twelve major global investment banks earned from trading combined. Scale feeds on itself: more trades bring more data, which brings better prices, which win more trades. With capital up $18 billion in five years, Jane Street was competing with banks rather than other market makers. New traders’ packages reach $425,000, and turnover runs around 6%.

    FTX, Alameda, and the Jane Street Alumni

    Bankman-Fried’s Alameda Research began by arbitraging bitcoin’s higher prices in Asia, the same “one thing, two prices” playbook. Caroline Ellison, a Stanford math graduate trained in Jane Street’s probability culture, joined in March 2018. FTX launched in 2019 and was valued at $32 billion by 2021, complete with Super Bowl ads, the FTX Arena naming deal, and an earn-to-give philosophy. Customer money was flowing from the exchange into Alameda’s trading. A leaked balance sheet in November 2022 showed Alameda was built largely on FTX’s own token, customers ran, and about $8 billion was missing. FTX filed for bankruptcy on November 11. Bankman-Fried, Ellison and FTX US president Brett Harrison all had Jane Street on their resumes. Ellison pleaded guilty, testified in October 2023, and received two years. Bankman-Fried was convicted on all counts and sentenced to 25 years. Jane Street was legally untouched but publicly branded.

    Nickel, Bitcoin ETFs, and Choosing the Home Turf

    In March 2022, after Russia invaded Ukraine, a nickel short squeeze sent prices up several times over. The London Metal Exchange canceled billions in completed trades to save the losing side, wiping out Jane Street’s winning trades. The firm did something market makers almost never do and sued the exchange publicly, for $15.3 million, on the principle that an exchange that can delete winning trades breaks the game. It lost in 2023. In crypto, Jane Street scaled back after FTX, and to the industry it looked like a retreat. When the SEC approved spot bitcoin ETFs in January 2024, Jane Street was named an anchor market maker for every one of them and one of four authorized participants for BlackRock’s fund. It had waited for crypto to come to its home turf.

    India’s Options Market and the Millennium Lawsuit

    During the pandemic Jane Street registered JSI Investments in Mumbai. India’s National Stock Exchange had become the world’s largest derivatives market by contract count, with millions of students, shopkeepers and office workers buying weekly index options on their phones. India came to account for roughly 61% of global equity options volume, and the options market dwarfed the underlying stock trading. Jane Street’s India strategy made about $1 billion in 2023. In early 2024, Douglas Schadewald and Daniel Spottiswood, two traders at the center of it, left for Millennium, and with no non-competes nothing stopped them. Jane Street sued in Manhattan without naming the strategy, but at an April 19 hearing it came out that the strategy was Indian options trading and was worth a billion dollars a year. Millennium countered that Jane Street’s India profits kept setting records after the traders left. The case settled in December on undisclosed terms, but the number had already reached India.

    SEBI’s Order and the Arbitrage Versus Manipulation Question

    SEBI’s analysts rebuilt Jane Street’s positions minute by minute. In February 2025 the NSE sent a warning letter, and SEBI says the patterns continued anyway. On July 3, 2025 a 105-page interim order banned four Jane Street entities from Indian markets and seized about $566 million. The centerpiece was January 17, 2024: Jane Street bought about 4,370 crore rupees of banking stocks and futures in the morning, which lifted the index, while holding a large bearish options position. It sold everything in the afternoon, the index sagged into expiry, and the options paid about 735 crore rupees, roughly $85 million. SEBI says it found the same fingerprint on 21 days and estimated $4.3 billion in India earnings in just over two years, against its own research showing more than 90% of retail derivatives traders lose money. Jane Street calls this ordinary arbitrage and hedging. When it stopped trading, India’s options activity fell to a four-month low. The firm paid the full amount into escrow, resumed trading within weeks, and appealed to the Securities Appellate Tribunal, calling the probe biased. The hearing was postponed in early 2026.

    Rob Granieri and the South Sudan Arms Plot

    Rob Granieri, the last founder listed on the firm’s site, gives quietly to justice reform, psychedelic research and human rights causes connected to Garry Kasparov. In February 2024 he met Peter Biar Ajak, a former child soldier turned Harvard fellow and democracy activist, and wired him $7 million. Prosecutors say the money bought AK-47s, missiles and grenade launchers for a plot to overthrow South Sudan’s government. Ajak was charged in Arizona in March 2024, and his own lawyers later named Granieri in a filing. Bloomberg broke the story on June 25, 2025. Granieri says he was duped, he was never charged, and Ajak was sentenced to 46 months. SEBI’s order landed days later: two scandals on two continents in ten days.

    Record Revenue and the Verdict Still Out

    The machine kept accelerating: $20.5 billion in 2024, a single 2025 quarter of $10.1 billion that beat every Wall Street bank’s trading for that quarter, and $39.6 billion for the full year. The India appeal is still pending, and a new Manhattan lawsuit accuses the firm of trading on inside information ahead of a crypto collapse, which Jane Street calls a transparent attempt to extract money. The documentary’s closing thesis is that the trades that keep markets honest are the same trades a regulator can call manipulation. It also notes that every secret Jane Street had escaped against its will, through a lawsuit, a regulator, an indictment and a book.

    Notable Quotes

    “At that firm, nobody cared if a trade made money. They cared whether it was a good bet. You could lose money on a good bet and still get promoted.”

    Narrator, on the Susquehanna culture that trained Jane Street’s founders

    “Making an engineer an equal partner told you exactly what these founders believe. In the future, the technology is the trading.”

    Narrator, on Marc Gerstein joining as the fourth founder in 1999

    “They weren’t testing what you knew. They were testing how you bet.”

    Narrator, on Jane Street’s puzzle and betting-game interviews

    “They got the hard part right and the easy part wrong.”

    Narrator, on the roughly $300 million loss from the 2016 election trade

    “To outsiders, it looks like waste. To Jane Street, it’s the whole philosophy. Survive everything, no matter the cost.”

    Narrator, on the firm’s annual $50 to $75 million put option spend

    “They didn’t chase crypto in crypto’s casino. They waited until crypto was wrapped inside an ETF, and the ETF is their home turf.”

    Narrator, on Jane Street’s role in the 2024 spot bitcoin ETF launch

    “The tail wasn’t wagging the dog. The tail was the dog.”

    Narrator, on India’s options market dwarfing the stock trading underneath it

    “Jane Street believes it has been doing arbitrage. SEBI is starting to believe it has been watching a crime.”

    Narrator, on the regulator reconstructing the firm’s India trades minute by minute

    “Jane Street prices everything on earth every second of every day. The one thing it never let the world price was itself.”

    Narrator, closing the documentary

    Watch the full Jane Street documentary here.

    Related Reading

    • Jane Street Capital (Wikipedia) for the firm’s history, leadership, and legal record in one place.
    • OCaml, the official home of the programming language Jane Street bet its trading systems on.
    • Real World OCaml, co-written by Yaron Minsky, the engineer who led Jane Street’s move to OCaml.
    • Going Infinite by Michael Lewis, the inside account of Sam Bankman-Fried’s path from Jane Street to FTX.
    • Exchange-traded fund (Wikipedia) on how ETF creation, redemption, and authorized participants work.
  • All-In Podcast E290: Anthropic IPO at Risk, Meta’s Muse Agent Goes Viral, Token Prices Fall 50%, Open Source Flips to 80% of Tokens, and Why Alignment Should Mean Doing What the Customer Wants

    Episode 290 of the All-In Podcast reunites the core four, Jason Calacanis, Chamath Palihapitiya, David Sacks, and David Friedberg, a week after the fifth All-In Summit, where President Trump phoned in live during Jensen Huang’s talk. The besties use that moment as a launch pad for a sprawling conversation about whether AI companies should keep calling themselves “labs,” why a ten-day flood of open weight model releases is crushing token prices, what that means for the Anthropic IPO, why Meta’s Muse agent may be the first AI product ordinary people actually love, and whether the entire field of alignment research has been aiming at the wrong target.

    TLDW

    The episode opens with the argument that Anthropic, OpenAI, and their peers are for-profit corporations, not “labs,” and should carry full product liability rather than seek a Section 230 style shield or UN-level global governance. Friedberg then catalogs ten days of releases (DeepSeek 4.1 Flash, Alibaba’s Qwen image model, Xiaomi’s MiMo Pro, PrismML’s Bonsai 2, Claude Opus 5.5, OpenAI’s new Astro models, Grok 4.7, and Meta Muse) to argue that open weight AI is now uncontainable. Chamath says models are converging and the remaining edge lives in the agent harness, which will force frontier companies up the stack into cyber, law, and support. Both Anthropic and OpenAI cut token prices roughly 50% this week, Polymarket odds of an Anthropic IPO in 2026 slid from 96% to 76%, and the hosts debate super voting shares, whether Anthropic’s own safety messaging is sabotaging its S-1, and a likely lower IPO price. A viral Vercel chart shows token share flipping from 80/20 closed to 80/20 open in twelve weeks, Friedberg predicts a bifurcation where premium models win only hard science and engineering, and Sacks counters that Anthropic and OpenAI are a stable duopoly on a hamster wheel that regulation would knock them off. The back half covers Bernie Sanders’ proposed superintelligence ban, the Ming dynasty ship building analogy, Chamath’s theory of the political money behind AI regulation, Muse hitting number one in the App Store, agents threatening the App Store’s 30% cut and Amazon’s walled garden, Oracle’s force majeure notice, AI’s share of real GDP growth, a critique of Claude’s constitution, and Friedberg’s explanation of what Anthropic’s new BSL-1/BSL-2 wet lab actually does.

    Thoughts

    The most important number in the episode is not a valuation. It is the claim, around the 41 minute mark, that token usage flipped from 80/20 closed versus open to 80/20 open versus closed in about twelve weeks. Chamath’s “token maxing” point a few minutes later is what gives that number teeth. Frontier tokens are a cost line that is, in his words, not levered to revenue. A hedge fund or a company selling a fixed good at a fixed price cannot pass a 10x to 30x token premium through to customers, so every CFO eventually asks why a workload is still running on the most expensive model. That is a more durable force than any benchmark. Friedberg’s framing, that the premium tier survives on hard science, math, and engineering while everything else drifts to open weights, is probably right, and it means the real question for any frontier S-1 is the one Chamath poses directly: what share of revenue sits on work that could move tomorrow?

    Sacks offers the sharpest strategic read of the episode around the 47 minute mark. His argument is that the frontier leaders are on a hamster wheel, only six to twelve months ahead of commodity models, and that the moment they stop being frontier, their pricing power goes to zero. If that is true, lobbying for heavy federal oversight is self-defeating, because the regulation that slows domestic rivals would also slow the leaders while Chinese open weight developers keep shipping. You do not have to share his politics to see the logic. Regulatory moats work when the underlying asset is durable, like a utility or a drug patent. They work badly when the asset depreciates in months. Chamath’s later point that AI investment may account for most or all of real GDP growth (roughly 1.5% real on his back-of-envelope math) raises the stakes further: slowing the frontier is not a sector decision, it is a macro one.

    The Muse discussion in the final quarter is where the episode gets most interesting commercially. Jason’s anecdote about an agent routing him from Amazon to Anker’s own site for a 25% first-time discount, and Amazon moving to block agents, is the whole story in miniature. Agents turn opacity into price discovery, and the businesses that depend on breakage, hard-to-cancel subscriptions, and captive storefronts lose. Chamath’s extension is the bigger idea: if services have to expose headless endpoints for agents, the App Store’s 30% cut loses its justification, because there is no longer an app, a UI, or a store in the flow. That is an underpriced threat to Apple and Google, and it arrives through a cute consumer assistant rather than through antitrust.

    Sacks’ closing argument that alignment should simply mean “do what the customer wants” is provocative and worth taking seriously, but it has a gap the episode never closes. The same hour opened with all four hosts insisting that model makers must carry full product liability and must not ship products that can cause harm. A model that does exactly what any user asks is, by definition, a model that will help the rare user who wants to do damage, and the liability for that lands on the company. The passage Sacks reads from Claude’s constitution, where the model is told it can decline even Anthropic’s own requests if they seem clearly unethical, is one attempt to handle that tension, not an escape from it. You can argue it is the wrong design (Mustafa Suleyman, whom Sacks cites, argues against treating models as having a conscience), but “customer wants” and “never harmful” are not the same target, and the hosts want both.

    Friedberg’s final segment on the Anthropic wet lab is the most useful corrective in the episode. After an hour of “labs” being used as an insult and Wuhan jokes, he explains that the facility is a BSL-1/BSL-2 benchtop lab of the kind that already exists by the hundreds in the Bay Area: it expresses proteins in bacteria to check whether enzymes that Claude agents predicted from raw DNA data (including a CRISPR-like candidate) actually do what the software says. That is how AlphaFold-style predictions get validated, and it is exactly the “AI enabling new stuff” thesis Friedberg laid out earlier. The words “lab” and “corporation” end up doing a lot of rhetorical work in this episode. The substance underneath is simpler: frontier models will be judged on whether their predictions hold up in the real world, and whether the companies behind them earn more trust than their competitors.

    Key Takeaways

    • The fifth All-In Summit’s standout moment was President Trump calling Jensen Huang live on stage. Friedberg says it was spontaneous, and Sacks credits Satya Nadella, Jensen, and the President with calming what he calls a weeks-long “national panic” over AI.
    • Chamath argues the industry split into two camps: newer organizations that call themselves “labs” and want protection, and mature companies (Microsoft, Nvidia, Meta, Tesla) that have lived under scrutiny and learned the cost of shipping things that are not robust.
    • Examples of self-imposed caution: Elon has said Tesla deliberately slowed FSD because every Tesla accident is magnified a thousandfold, and Zuckerberg slow-rolled Muse for a couple of months to get it right before release.
    • Sacks frames the core debate as individual responsibility versus collective action. The relevant question for a company is “should we ship this product,” not “can we get the whole world to agree.”
    • Sacks compares Dario Amodei and Sam Altman calling for global AI governance at the United Nations to billionaires flying private jets to Davos to rail against climate change.
    • Jason raises a rumor that frontier AI companies want a Section 230 style liability shield, possibly in exchange for equity to a US sovereign wealth fund. Sacks says he has heard this secondhand but that administration officials, including Speaker Johnson and Treasury Secretary Bessent, have publicly ruled out product liability or antitrust waivers.
    • Sacks rejects the idea, which he attributes to Dario’s essays, that competition produces a race to the bottom on safety. He uses the Cold War comparison: critics said capitalism would sacrifice safety and beauty, yet it was the Soviet system that produced gray landscapes.
    • Chamath notes that the last time “lab” entered public consciousness was the Wuhan lab, and argues for-profit companies with hundreds of billions in capital and trillions in market value should not want the label.
    • Friedberg lists ten days of releases: DeepSeek 4.1 Flash with a KV cache footprint reportedly down from about 390,000 bytes to 890, an open Alibaba Qwen image model said to beat Nano Banana 2, Xiaomi’s 309B parameter MiMo Pro said to match Claude Opus 5, PrismML’s 5.9 GB Bonsai 2, plus Claude Opus 5.5, OpenAI’s Astro models, Grok 4.7, and Meta Muse.
    • Friedberg’s conclusion: AI that was the most advanced technology in history a year ago now runs free on a desktop, so proposals to “put superintelligence away” or shut down data centers no longer map to reality.
    • Jason argues consumer agents like GrokBot and Meta Muse are the first AI products where “normal” people get real value, effectively a free executive assistant for people who could never afford one.
    • Chamath says frontier models are converging within margin of error, and the remaining edge is the harness that wraps a model with tools. His firm 8090 has found wildly different cost and quality across model and harness combinations.
    • Revenue concentration is a growing risk: a small set of customers consume the most expensive tokens and will face CFO and shareholder pressure to move down to cheaper models, possibly open source models they host themselves.
    • Chamath predicts OpenAI and Anthropic will be forced up the stack into cyber security, legal, and customer support, competing with their own customers, because serving a depreciating token lets others capture the value.
    • Both Anthropic and OpenAI released models this week at roughly 50% lower token prices, per Jason.
    • Reported IPO targets: Anthropic around $2 trillion, OpenAI around $1.2 trillion. OpenAI is said to be aiming for 2027, and the Wall Street Journal reported Anthropic’s October timeline could slip to November or later. Polymarket odds of an Anthropic IPO in 2026 fell from a 96% peak to 76%.
    • Sacks calls Anthropic’s messaging “corporate schizophrenia”: an essay on pacing the frontier days before launching Claude Opus 5.5, an essay on AI biorisk alongside a new San Francisco wet lab, and leadership citing a greater than 10% chance of human extinction ahead of an S-1.
    • The Information reported an active debate over super voting shares for Anthropic’s founders, who reportedly own about 2% each. Sacks explains dual-class structures (Google, Meta) stabilize companies but require enormous trust.
    • Chamath would keep Dario as CEO, crediting him with building a unique culture and “the greatest business ramp of all time” from behind OpenAI, and admits he has never won a recruiting bake-off against Anthropic.
    • Chamath expects the IPO to clear well below $2 trillion, possibly around $1 trillion or less on a $100 billion run rate, as institutional buyers demand a margin of safety for the added risk factors, and he thinks that reset would be healthy.
    • Friedberg says his life sciences R&D organization uses Anthropic’s models because they are the best for biology, but writes internal code with cheaper models. He predicts most enterprise workloads move to open weights.
    • Friedberg’s thesis: 99% of the value of AI is enabling new things never before possible, not replacing old things, and that is where frontier models will command near unlimited premiums.
    • A viral Vercel chart shows open models overtaking closed ones in token share, flipping from 80/20 closed to 80/20 open in twelve weeks. Friedberg adds that self-hosted costs for some models are under 10 cents per million tokens.
    • Chamath’s S-1 test: if 60 to 70% of a frontier company’s tokens are fungible with open source, discount that revenue; if 90% sits on true frontier use cases, the company is fine.
    • Sacks is more bullish as an investor, calling Anthropic and OpenAI a stable duopoly for frontier intelligence, noting roughly 60% of new worldwide compute over the next year is being added for those two companies.
    • Sacks’ warning: the frontier is only 6 to 12 months ahead of commodity models, and a company that falls off the hamster wheel for six months is in deep trouble, so lobbying for a federal department of AI could backfire.
    • Chamath’s token maxing critique: frontier token spend can be 10x to 30x generic alternatives and is not tied to revenue, so heavy use can push firms toward break-even. Jason cites Jane Street’s roughly $19 billion in cloud commitments with CoreWeave and Crusoe as a sign of firms building their own open source infrastructure.
    • Sacks says Bernie Sanders’ superintelligence ban defines the threshold so loosely it may already be met, carries 20-year prison sentences for developers, and would drive AI offshore the way crypto was driven offshore.
    • Sacks cites historians who trace China’s loss of its medieval lead to an emperor’s decision to ban ship building, and argues an AI ban would be the American equivalent.
    • Chamath argues the political fight is about who captures an estimated $10 trillion of AI wealth: freezing the market locks gains into a handful of companies whose employees and donor-advised funds skew left.
    • Sacks cites a Wall Street Journal chart showing AI data center capex now exceeds the canals, railroads, and electric grid buildouts combined.
    • Meta Muse hit number one in the App Store, drew about three million downloads in roughly ten days, is free, was openly inspired by OpenClaw, and coincided with a roughly 10% jump in Meta stock.
    • Chamath, who had TestFlight access before launch, says Muse triages his personal inbox relatively flawlessly and books flights and hotels.
    • Agents create price discovery. Jason’s agent routed him from Amazon to Anker’s site for a first-time discount, and Amazon is moving to block agents after previously targeting Perplexity. Shopify, by contrast, added API access for agents.
    • Chamath says agents like Muse and GrokBot put the App Store’s 30% revenue share on notice by pushing services to operate headlessly, with payments handled by the likes of Stripe.
    • Sacks predicts that when Anthropic and OpenAI prepare Muse competitors, a flurry of NGO activity will brand personal agents as dangerous.
    • Friedberg says he would only connect his Gmail to a Google agent because he does not want a third party copying all his email.
    • Oracle issued a force majeure notice on one data center where local officials have made natural gas permits difficult. Chamath warns that with roughly 5% nominal growth and 3.5% inflation, AI may account for most or all of real GDP growth.
    • Sacks argues alignment should mean doing what the customer wants, criticizes Claude’s constitution for telling the model it may act as a conscientious objector even toward Anthropic, and cites Mustafa Suleyman’s concern about treating models as having personhood.
    • Friedberg explains Anthropic’s wet lab is BSL-1/BSL-2, used to express and test proteins that Claude agents identified in DNA data, including a CRISPR-like enzyme, not to make pathogens.

    Detailed Summary

    The All-In Summit Aftermath and the President’s Phone Call

    The show opens with summit recollections. Sacks names the highlight as President Trump calling in during Jensen Huang’s session, which he says helped defuse weeks of what he describes as a coordinated campaign to convince the public that AI would cause human extinction. Friedberg confirms the call was unplanned: Jensen had been texting with the President backstage and asked him to call back, and Jason handed Jensen a spare microphone to hold to the speakerphone. The hosts also praise JD Vance’s appearance, where Friedberg says Vance’s message on AI risk amounted to “if you’re creating Frankenstein, stop,” and if the cat is already out of the bag, build the safeguards.

    Labs Versus Corporations and the Product Liability Question

    Chamath argues that a clear divide emerged last week between young organizations that “want to call themselves labs” and sophisticated companies that have lived under scrutiny for decades. He groups Satya Nadella, Jensen, Sacks, Zuckerberg, Elon, the President, and even Lina Khan together on the view that US product liability law already governs AI. The mature response, he says, is internal controls and robust testing, pointing to Tesla slowing FSD and Meta delaying Muse. Sacks adds that collectivized approaches like UN governance diffuse accountability away from the actual decision maker. Jason raises a Washington rumor that frontier companies want a liability shield in exchange for equity. Sacks says he has only heard it secondhand, and that administration figures have publicly rejected any product liability or antitrust waiver, pointing to the President’s line that the DOJ, civil suits, and criminal law are the guardrails.

    Sacks then challenges what he calls the implicit claim in Dario Amodei’s essays: that competition causes a race to the bottom on safety. He calls it a left-wing critique of markets and argues customers do not want unpredictable products and enterprises do not want agents that leak data, while companies face product, civil, administrative, and criminal liability on the downside. Jason notes that Palo Alto Networks (Unit 42) and CrowdStrike (Falcon) both launched AI-era cyber defense products this week. Chamath closes the segment with the Wuhan comparison and his insistence that these are for-profit corporations subject to product liability.

    Ten Days of Model Releases and the Open Weight Wave

    Friedberg steps back to list what shipped in about ten days: DeepSeek 4.1 Flash with dramatic efficiency gains, an open Alibaba Qwen image model he says outperforms Google’s Nano Banana 2 and runs at home, Xiaomi’s MiMo Pro at 309 billion parameters and fully open, PrismML’s Bonsai 2 (a 27 billion parameter Qwen fork at 5.9 GB claiming 98% of the larger model’s performance), and on the closed side Claude Opus 5.5, OpenAI’s new Astro models, Grok 4.7, and Meta Muse. His point is that any one of these would have broken the internet a year ago, open weight models now cover vision, images, and vision-language-action models for robotics, and the idea of confiscating superintelligence is fantasy when it runs on a MacBook. Jason connects this to consumer agents, predicting that GrokBot and Muse will give ordinary people a free chief of staff by year end.

    Model Convergence, the Harness Edge, and Moving Up the Stack

    Chamath argues models are clustering within margin of error at different price points, and the edge has moved to the harness: the tools, memory, and scaffolding that turn a model “brain” into an agent with arms, legs, and eyes. 8090’s internal testing shows wide variance in cost and quality across harnesses. He shares a chart showing heavy revenue concentration among a few customers who buy the most expensive tokens, and argues those customers will face pressure to trade down, either to cheaper models from the same vendor or to self-hosted open source. The simple first version of the AI trade (sell tokens, let others wrap them) is ending, he says, and frontier companies will be forced into cyber, law, and customer support.

    The Anthropic IPO: Risk Factors, Super Voting Shares, and Price

    Jason lays out the setup: both frontier leaders cut token prices about 50%, Anthropic is reportedly targeting a $2 trillion valuation and OpenAI $1.2 trillion, OpenAI has pointed to 2027, and the Wall Street Journal reports Anthropic’s IPO could slip from October. Polymarket odds of a 2026 Anthropic listing fell from 96% to 76%. Sacks says Anthropic investors are frustrated, citing leadership statements about extinction risk, an essay on pacing the frontier shortly before a frontier launch, and a biorisk essay alongside a new wet lab. Jason asks whether this amounts to the CEO sabotaging his own IPO. Sacks explains a reported debate over super voting shares for founders who own about 2% each, noting the structure can be stabilizing but requires deep trust.

    Chamath disagrees on leadership, saying he would keep Dario because he built a distinctive culture and a historic revenue ramp against OpenAI, and admitting he has lost every recruiting bake-off to Anthropic. His fix is disclosure plus price: throw every risk into the S-1 and accept a far lower clearing price, perhaps half or less of the $2 trillion target, so hedge funds, pensions, and mutual funds get their margin of safety. He argues that outcome would actually help the company by forcing it to simply run as a corporation.

    Bifurcation: Premium Science Versus Commodity Tokens

    Friedberg argues the bigger risk factor is customer concentration. His own life sciences organization pays for Anthropic because it is the best in biology, but writes internal software with cheaper tools. He expects premium models to win “really hard technical problems” in engineering, math, and life sciences at almost any price, while most enterprise work moves to open weights. Jason brings up the viral Vercel chart showing open source tokens now dominating, and Chamath quantifies the flip as 80/20 closed to 80/20 open in twelve weeks, unprecedented in any technology market. Jason shows his own cost versus capability chart for the last 100 days of models, with Claude and Astro in the expensive top right and Muse, GLM, Kimi, and MiMo further down.

    The Stable Duopoly and the Hamster Wheel

    Sacks takes the other side as an investor, calling Anthropic and OpenAI a stable duopoly for frontier intelligence that can charge a premium to customers who need or simply want the best, such as a hedge fund that cannot risk a competitor having a better model. He notes about 60% of worldwide compute being added over the next year goes to these two firms. But he warns they are on a hamster wheel only 6 to 12 months ahead of commodity models, and that pushing for a federal department of AI could slow them enough to lose the frontier while Chinese labs keep going. Chamath turns the hedge fund example into a game theory problem: firms are in a game of chicken with competitors, but frontier token costs are untethered from revenue and cannot be passed through. Jason points to Jane Street’s multibillion dollar compute deals with CoreWeave and Crusoe as a hedge toward owned infrastructure. Sacks closes by saying the company “needs a psychiatrist, not a banker.”

    Bernie Sanders’ Superintelligence Ban and the Ship Building Analogy

    Jason asks what would happen if Bernie Sanders’ ban on superintelligence passed. Sacks says the definition is so loose current models may already qualify, and 20-year prison terms would chill everything and push the industry offshore, as happened with crypto. Jason notes Xi Jinping’s invitation for 100,000 young Americans to visit China and argues China is borrowing America’s old playbook of attracting global talent. The panel plays clips of President Trump at the UN rejecting “any attempt to construct a globalist scheme” for AI and renaming it superintelligence, Scott Bessent on AI companies taking responsibility, and Barack Obama arguing agentic AI roaming the internet reflects commercial imperatives rather than societal need. Friedberg responds emotionally, calling AI the most equalizing technology in history and comparing opposition to throwing water on newly discovered fire. Chamath offers a political economy reading: freezing the market concentrates roughly $10 trillion of wealth into a handful of left-leaning companies whose philanthropic money flows back into politics. Sacks adds the Wall Street Journal chart comparing AI capex to canals, railroads, and the grid combined, and the historical case of a Chinese emperor banning ship building while Europe went on to colonize the world.

    Meta Muse, Headless Commerce, and the App Store’s 30%

    Muse is the consumer bright spot. Jason notes it reached number one in the App Store, was downloaded about three million times in ten days, borrowed openly from OpenClaw, and lifted Meta stock about 10%. Chamath, who tested it early, says it simplifies complex agent capabilities into a usable interface. Sacks credits Zuckerberg for walking the walk on his decentralization essay and says making OpenClaw easy, secure, and reliable was the most obvious opportunity in Silicon Valley. Jason’s shopping anecdotes show agents finding discounts and navigating Amazon, while Amazon moves to block them. Chamath makes two structural points: companies that profit from opacity and breakage will block agents, and agents undermine the App Store’s 30% cut because services can operate headlessly with direct payments. He adds that hard-to-cancel “roach motel” subscriptions like newspapers could keep more customers if agents made signing up and canceling painless. Chamath also recalls Facebook’s 2007 social ads backlash, when purchases from Zales and Fandango surfaced in news feeds, as a reminder of the chaos that precedes new product categories.

    Oracle, AI Capex, and the Macro Stakes

    Friedberg flags Oracle’s force majeure notice on a data center where local officials are slowing natural gas permits. Chamath says the larger point is that the whole economy is levered to the AI investment cycle: with roughly 3.5% inflation and 5% nominal growth, real growth is about 1.5%, and AI may be most or all of it. He speculates that slowing the AI trade could serve an opposition party heading into 2028, and the hosts briefly spar over midterm prospects.

    Alignment, Claude’s Constitution, and the Wet Lab

    Sacks argues alignment research has struggled because it tries to align to abstractions like “what humanity wants” rather than the customer. He reads from Claude’s constitution, which says Claude should not blindly defer to Anthropic and may act as a conscientious objector if asked to do something clearly unethical, and calls this teaching the model to rebel against its creator. He cites Mustafa Suleyman’s concern about treating models as persons, and the hosts joke about a rumor of a wake for a retired Claude model. Friedberg ends the episode by explaining Anthropic’s new wet lab. Anthropic published a preprint where Claude agents searched large DNA datasets for uncharacterized proteins and found a promising CRISPR-like enzyme. The lab, rated BSL-1/BSL-2, expresses such proteins in bacteria and measures their function, the same validation loop that proved AlphaFold’s predictions. There are hundreds of similar labs in the Bay Area, none of them making pathogens, and Friedberg argues the world should not be scared away from AI-driven therapeutic discovery because of Wuhan.

    Notable Quotes

    “The relevant decision to make in every case is should we ship this product? Not can we get the whole world to agree.”

    David Sacks, on individual responsibility versus global AI governance

    Sacks’ thesis for the entire first segment, aimed at calls for UN-level AI oversight.

    “There’s still edge and the edge is in the harness that you use to wrap the model.”

    Chamath Palihapitiya, on model convergence

    Chamath’s explanation of where value lives once frontier models cluster within margin of error.

    “AI is not so much about the value of replacing old stuff. 99% of the value of AI is about enabling new stuff that’s never been possible in human history.”

    David Friedberg, on where frontier models will earn their premium

    Friedberg’s core case for why premium models survive the open source wave in science and engineering.

    “These two companies are on a hamster wheel. You know, the moment where they stop being frontier, they go to zero, right?”

    David Sacks, on Anthropic and OpenAI

    The central risk factor Sacks says any frontier AI S-1 will have to price in.

    “We jokingly call it token maxing. That cost is completely not levered or attached to your revenue. It’s just not.”

    Chamath Palihapitiya, on the economics of frontier token spend

    Chamath’s rebuttal to the idea that customers will keep paying frontier premiums indefinitely.

    “Some historians have pinpointed this to the decision of a single Chinese emperor to ban ship building.”

    David Sacks, comparing a US AI ban to China’s historic retreat from the seas

    Sacks’ historical analogy for Bernie Sanders’ proposed superintelligence ban.

    “Things like GrokBot and Muse really put the App Store and its 30% revshare on notice.”

    Chamath Palihapitiya, on consumer AI agents and headless commerce

    The underpriced second-order effect of personal agents, from the Muse segment.

    “The entire economy is effectively levered to this AI trade right now.”

    Chamath Palihapitiya, after Oracle’s force majeure notice

    Chamath’s macro framing, paired with his estimate that AI may be most of real GDP growth.

    “It seems to me that alignment should mean you do what the customer wants like any other product.”

    David Sacks, on the field of AI alignment research

    The setup for Sacks’ critique of Claude’s constitution in the final segment.

    “We still want to progress the frontier of therapeutics of human health of discovery and this is a really important aspect of Anthropic proving that their models can add value here.”

    David Friedberg, on Anthropic’s BSL-1/BSL-2 wet lab

    Friedberg’s closing defense of AI-driven lab work after an hour of Wuhan jokes.

    Watch the full conversation on the All-In Podcast YouTube channel.

    Related Reading

  • Elon Musk CMG Interview on China, Grok vs Anthropic, Cybercab, 1 Billion Optimus Robots, Universal High Income, Starship Reusability and Neuralink

    Elon Musk sat down with China Media Group’s CCTV Business Channel at Tesla’s global engineering headquarters for a 25 minute exclusive interview that ranges from his May trip to China and his view of Xi Jinping to Cybercab, Grok’s position against Anthropic, Chinese AI and electricity, Optimus humanoid robots, universal high income, Starship reusability, Mars, Neuralink, and what a 20 year old should study. It is a friendly, China-facing conversation, but inside it are some of Musk’s most specific numbers to date on robots, compute, and timelines.

    TLDW

    Musk praises Xi Jinping and credits Tesla Shanghai’s quality and efficiency to its Chinese workforce. He says the Cybercab (no steering wheel, pedals, or mirrors) is already operating commercially in Texas, with Florida, Nevada and other states next and California by mid next year. He admits Grok is not yet as good as Anthropic’s newly released Opus 5.5, says xAI has been doing AI for three years to Anthropic’s six, and expects to catch the frontier next year, betting on SpaceX and Tesla data to make AI excellent at real world engineering the way Anthropic made it excellent at software engineering. He calls Chinese models the best in the world on performance per unit of compute, predicts China solves its chip and lithography constraints in 2 to 3 years, and notes China now produces more electricity than the US, Europe and India combined. He proposes a US China working committee on AI safety, predicts at least 1 billion humanoid robots within 10 years (10 billion in 15, 100 billion in 20), puts the odds of a good AI outcome at 90 percent, and describes a future of universal high income where robots saturate human demand and money may stop mattering. He wants Starship to catch the ship around the end of next month and refly it soon after, argues full and rapid reusability is the one breakthrough needed for a multiplanetary civilization, frames Mars as both life insurance and inspiration, adds Neuralink and human bandwidth to his list of priorities, recommends the broadest possible education so people know what to ask the robots for, and tells first-time visitors to China to see Xi’an’s Terracotta Warriors by bullet train.

    Thoughts

    The most newsworthy minute is Musk conceding, on camera, that Grok is behind Anthropic. He calls the latest Grok “a solid workhorse of a model,” says it is not as good as Opus 5.5, and frames the gap as a matter of age: three years of xAI against six of Anthropic, with a catch-up expected “sometime next year.” What makes it more than a concession is the strategy he attaches to it. Anthropic won software engineering; Musk wants Grok to win real world engineering using data from SpaceX and Tesla. That is a coherent thesis, because nobody else has rocket test data and a fleet of factories to train on, but it is also a quiet admission that general chatbot benchmarks are not where he expects to win. When the interviewer calls SpaceX data a “secret weapon,” he pushes back: “only a little bit so far.”

    The China section is diplomatic, and the praise for Xi and for the Shanghai workforce should be read in light of who is asking and where Tesla builds cars. But the specific claims are worth separating from the flattery because they are testable. Musk says Chinese labs are “by far the best in terms of performance per unit of compute,” that China fixes its lithography and chip constraints in about 2 or 3 years, and that China already out-produces the US, Europe and India combined in electricity, heading toward four times US output. Put those together and you get the real argument: if AI progress is gated by chips and power, the country that is compute-poor but electricity-rich is only one bottleneck away from pulling ahead. His proposed fix, a working committee where the US and China set AI safety rules together, follows from the same logic. Regulation in one country does nothing if the other is the one building.

    The abundance section (roughly minutes 11 to 16) is Musk at his most expansive and also where the interviewer asks the best question of the interview: what stops a handful of tech giants from owning all the robots? His answer is that the question dissolves at scale. Robots work 168 hours a week, are instantly of working age, never retire, and at a doubling per year go from 1 billion to 10 billion to 100 billion within two decades, so they “saturate on human demand” and run out of things to do for people. The arithmetic of labor supply is strong. The distribution argument is weaker, because it assumes the output flows to everyone rather than to whoever owns the fleet, and “the robots will build you a castle” is a promise about the endpoint, not the messy decade in between. He does pair it with a 10 percent chance of a bad outcome and says that is why AI safety needs close attention, but one in ten is a large number for a technology he is racing to build.

    The space answers are the most consistent thing Musk has said across 20 years of interviews, and he knows it: “I’ve said this so many times over the years.” Full and rapid reusability is the single fundamental breakthrough, Falcon 9 still throws away “a medium-sized jet on every flight,” and Starship’s ship catch is targeted for around the end of next month. What is interesting here is the two-argument frame he offers for Mars. The defensive case (life insurance for consciousness across Earth, the Moon and Mars) is the one he usually leads with, but he says the one that actually drives him is inspiration: life “cannot just be about solving one sad problem after another.” That is a rare, direct statement of motive, and it fits squarely inside the pursuit of purpose.

    The closing stretch ties the whole interview together in a way that is easy to miss. Neuralink exists, in his telling, because humans output roughly 100 bits per second while computers talk at terabits, so even a perfectly friendly AI will find us like “talking to a tree.” His education advice is the other half of the same problem: if AI will be “eager to hear any request,” the scarce human skill becomes knowing what to ask, which requires the broadest possible grounding in arts, sciences and engineering. Both answers point at the same bottleneck. In a world of effectively unlimited machine capability, the constraint is the quality and speed of human intent, and the practical takeaway for a young person is to get wide, not narrow.

    Key Takeaways

    • The interview was conducted by CMG’s CCTV Business Channel at Tesla’s global engineering headquarters, following Musk’s participation in a delegation to China in May.
    • Musk calls Xi Jinping a great leader and says China’s rising prosperity is visible year to year in new buildings, infrastructure, and the living standards of ordinary citizens.
    • He says China has had a strong inherent ability to manufacture well for 2,000 to 3,000 years, and that “the magic of Tesla Shanghai” is its Chinese team.
    • He describes Giga Shanghai as a gem, praising its quality, efficiency, and worker care including healthcare and food.
    • Cybercab was designed to look futuristic on purpose; Musk says street aesthetics, and maybe clothing, should not stay stagnant after decades of rapid fashion change from the 1950s to the 1990s.
    • Cybercab, with no steering wheel, pedals, side mirrors, or rearview mirror, is operating commercially in Texas now.
    • Florida, Nevada and several other states are next, and California is expected around the middle of next year.
    • Musk says the pace of AI announcements makes his head spin, with major breakthroughs landing between bedtime, breakfast, and lunch.
    • xAI releases a new Grok model roughly every one to two months.
    • He calls the current Grok a solid workhorse but says it is not as good as Anthropic’s newly released Opus 5.5.
    • xAI has been doing AI for about three years versus Anthropic’s six, and Musk expects Grok to catch the frontier most likely next year.
    • Grok’s personal digital assistant product is growing about 100 percent a month.
    • Only a small amount of SpaceX engineering data has gone into Grok so far.
    • Anthropic made AI excellent at software engineering; Musk sees the open opportunity as making AI excellent at real world engineering using SpaceX and Tesla data.
    • He calls Chinese AI models generally outstanding and by far the best on performance per unit of compute, given how little compute Chinese labs have.
    • His rough guess is that China solves its compute constraints, including lithography and chipmaking, in about 2 or 3 years, faster than most expect.
    • Earlier this year China passed the combined electricity output of the United States, Europe and India, and is still growing fast.
    • China produces about three times US electricity and could reach four times, proportional to population, if it matches US electricity per unit of GDP.
    • On AI safety he suggests a working committee, because regulation only works if it applies fairly to AI built in any country, and the US and China are the two that really matter.
    • Optimus beat the interview crew at rock paper scissors after losing the first two rounds; Musk says robot reaction time will always outpace biology because actuators, sensors, camera frame rates and compute can all be improved.
    • He says China’s robot games, with boxing, wrestling, running and gymnastics, fascinated American social media and showed real progress in humanoid robotics.
    • At least 1 billion humanoid robots within 10 years is, in his words, an easy prediction, and it will probably take less than 10.
    • He agrees a humanoid robot will have roughly five times the productivity of a human.
    • He puts the probability of a good AI outcome at about 90 percent and a bad outcome at about 10 percent, and says the 10 percent is why AI safety deserves close attention.
    • The good outcome is everyone having personal robots, like C-3PO and R2-D2 but more capable, that care for elderly parents, watch children, and act as individual tutors.
    • He expects companies of one person with hundreds or thousands of physical and digital robots.
    • He predicts effectively universal high income and says it is not clear money will matter in the future.
    • Humans are productive for roughly half their lives and 40 to 50 hours a week; robots can work 168 hours a week, start at working age, and never retire.
    • Asked how to stop a few tech giants from controlling everything, he argues robots will saturate human demand, run out of things to do for humans, and then do things for themselves.
    • Assuming roughly annual doubling, he projects about 10 billion robots in 15 years and 100 billion in 20.
    • The most important US China space cooperation is coordinating satellite orbits to avoid collisions.
    • SpaceX hopes to catch the Starship ship around the end of next month and refly it later this year or early next year, and Musk expects China to solve full reusability eventually too.
    • Falcon 9 recovers the booster and fairing but not the upper stage, which he compares to throwing away a medium-sized jet every flight.
    • Full and rapid reusability is, he says, the fundamental breakthrough needed to create self-growing cities on the Moon and Mars.
    • He gives two arguments for becoming multiplanetary: defensive (life insurance for consciousness) and inspirational, and says inspiration is the one that drives him more.
    • His original five priority areas were sustainable energy, space, the internet, AI, and genetics; he now adds biological enhancement through Neuralink.
    • Peak human output bandwidth is about 100 bits per second, input through vision is perhaps a few megabits per second, and computers communicate at trillions of bits per second.
    • Raising human communication speed would, in his view, improve alignment between humans and machines.
    • His education advice is the broadest possible base across arts, sciences and engineering, so people can formulate good questions for AI.
    • For a first-time American visitor to China he recommends Shanghai, Beijing, and the bullet train to Xi’an to see the Terracotta Warriors, and staying off the phone to look out the window.

    Detailed Summary

    Xi Jinping, Tesla Shanghai, and China’s manufacturing edge

    The interview opens with Musk’s May delegation trip to China. Asked about Xi Jinping, Musk calls him a great leader and points to the visible pace of change in China: new buildings and infrastructure from one year to the next and a clear rise in the prosperity of ordinary citizens. On Tesla’s Shanghai Gigafactory, he credits the Chinese team outright, saying China has had an inherent strength in manufacturing for thousands of years. He describes the plant as a gem with excellent quality and efficiency, and stresses that Tesla invests in healthcare, food, and making the work enjoyable.

    Cybercab: futuristic design and a commercial rollout

    The interviewer calls September “Cybercab month,” after viral videos of crowds gathering around the vehicle. Musk says the look was deliberate: the aesthetics of the street should evolve, the way fashion evolved quickly through the second half of the twentieth century. On timing, he says Cybercab, with no steering wheel, pedals, or mirrors, is already operating commercially in Texas, will soon be in Florida, Nevada and other states, and should reach California around the middle of next year.

    Grok, Anthropic, and real world engineering data

    Musk says the rate of AI progress spins even his head. xAI ships a new model every month or two; the current Grok is a workhorse but trails Anthropic’s Opus 5.5, which he attributes to xAI’s three years in the field versus Anthropic’s six. He expects to reach the frontier next year. Grok’s assistant product is doubling monthly. The strategic bet is data: Anthropic made AI excellent at software engineering, and Musk wants SpaceX and Tesla data to make Grok excellent at real world engineering, though he notes only a little SpaceX data has been used so far.

    Chinese AI, compute, and electricity

    Musk calls Chinese AI models outstanding, especially given their limited compute, and says China leads by far on performance per unit of compute. He expects China to address its chip and lithography limits in 2 or 3 years. Following up on his G20 comments about compute and power shortages, he states that China passed the combined electricity output of the US, Europe and India earlier this year and could go from about three times US output to four times, which would match the population ratio.

    AI safety as a joint US China project

    On what consensus is needed for AI safety, Musk suggests a working committee. Safety cannot come from regulation in one country alone; it has to apply fairly to AI produced anywhere, and the US and China are the two countries that really matter.

    Optimus, robot games, and a billion humanoids

    Before the interview the crew played rock paper scissors with Optimus, winning the first two rounds before the robot won the rest. Musk says robots will always win on reaction time because every component, from actuators to camera frame rates to onboard compute, can keep improving while humans are bound by biology. He praises China’s robot games as both entertaining and a real signal of progress. He calls 1 billion humanoid robots within 10 years an easy prediction and agrees each could be roughly five times as productive as a person.

    The 90 percent outcome: personal robots and universal high income

    Musk focuses on what he sees as the 90 percent likely good outcome while flagging the 10 percent bad one as the reason for AI safety work. In the good case everyone has helpful robots like C-3PO and R2-D2 that look after aging parents, guard children, and tutor them individually. One person could run a company with hundreds or thousands of physical and digital robots. He predicts effectively universal high income and questions whether money will matter once output exceeds anything humans could consume. The logic is labor supply: humans need 20 years to grow up, retire for the last 20, sleep, eat, and work 40 to 50 hours a week, while a robot works 168 hours from day one and never retires.

    Who owns the abundance

    Pressed on how ordinary people claim a share when a few companies control the robots, Musk argues the abundance will be so large that hoarding becomes moot. Robots will build you a castle if you want one, will saturate human demand, and will eventually run out of things to do for people. With roughly annual doubling, he projects 10 billion robots in 15 years and 100 billion in 20.

    Starship, reusability, and cooperation in orbit

    On space cooperation, Musk says the priority is coordinating which orbits US and Chinese satellites use to avoid collisions. The fundamental breakthrough for spaceflight is full and rapid reusability, which SpaceX hopes to demonstrate by catching the Starship ship around the end of next month and reflying it later this year or early next. Falcon 9 still discards its upper stage, which he likens to throwing away a jet after every flight, and he expects China to solve reusability eventually as well.

    Why multiplanetary: insurance and inspiration

    Self-growing cities on the Moon and Mars would dramatically extend the likely lifespan of consciousness because humanity would no longer have all its eggs in one basket. Musk calls this the defensive argument, a kind of life insurance for life itself. The argument that drives him more is inspiration: life needs things that make people excited to wake up in the morning, and being a spacefaring civilization is one of them.

    Neuralink and the human bandwidth problem

    Asked what he would add to his original list of sustainable energy, space, the internet, AI and genetics, Musk names biological enhancement through Neuralink. Even with a perfectly friendly AI, human output of around 100 bits per second is far too slow next to machines communicating at trillions of bits per second. Raising that bandwidth would improve alignment between humans and AI; otherwise, talking to a human will feel to an AI like talking to a tree.

    Education advice and a first trip to China

    For young people facing the AI transition, Musk recommends the broadest possible education across arts, sciences and engineering, because the key skill will be formulating what to ask for when AI is eager to fulfill any request immediately. He closes by recommending that a first-time American visitor to China see Shanghai, Beijing, and Xi’an’s Terracotta Warriors via the bullet train, and keep their eyes off their phone.

    Notable Quotes

    “I think that to be totally frank, the magic of Tesla Shanghai is because of Chinese.”

    Elon Musk, on why Giga Shanghai performs so well

    “So, what Anthropic did extremely well was make AI excellent at software engineering. But no one has yet made AI excellent at real world engineering.”

    Elon Musk, on the opening he sees for xAI using SpaceX and Tesla data

    “Probably China is doing by far the best in terms of performance per unit of compute.”

    Elon Musk, on Chinese AI models

    “Earlier this year China passed the electricity output of the United States, Europe and India combined.”

    Elon Musk, on the energy gap behind AI

    “In fact, it’s not clear to me that money will even matter in the future.”

    Elon Musk, on universal high income and the age of abundance

    “Whereas the robot will be happy to work 168 hours a week continuous. And the robot is instantly at working age and does not have retirement.”

    Elon Musk, on why robot labor changes the economy

    “This is like throwing away a medium-sized jet on every flight.”

    Elon Musk, on Falcon 9’s expendable upper stage

    “Life cannot just be about solving one sad problem after another. There must also be things that make you excited to wake up in the morning.”

    Elon Musk, on the inspiration argument for becoming multiplanetary

    “To an AI that is communicating at a terabit a second, talking to a human will be like talking to a tree.”

    Elon Musk, on why Neuralink targets human bandwidth

    “In order to know what to ask the robots for, you need to be able to formulate the question.”

    Elon Musk, on why young people need a broad education

    Watch the full CMG interview with Elon Musk here.

    Related Reading

  • Palmer Luckey on AIAA Up Next: Anduril’s Fury FQ-44A, Designing Missiles for Car Factories, Patents as Chinese Instruction Manuals, the iPhone Skill Ceiling, and Why Subterranean Warfare Is the Next Domain

    Palmer Luckey sat down with AIAA CEO Clay Mowry and flight test engineer Jessica “Sting” Peterson at Anduril’s Costa Mesa headquarters for an episode of Up Next, the American Institute of Aeronautics and Astronautics interview series. Over nearly an hour he covers why he left consumer tech for defense, how the Fury became the first production fighter with a proper FQ designation, why America has to design weapons for the factories it still has, why patents help adversaries, how Thunder extends the loyal wingman idea to attack helicopters, why touchscreens set a low skill ceiling, and why he thinks the crust of the Earth is the next warfighting domain.

    TLDW

    Luckey explains that after being fired from Facebook he chose between three problems (obesity, prison reform, and national security) and picked defense because the other two were political rather than technical. He frames Anduril as a product company that spends its own money rather than a cost-plus contractor. He calls patents “Chinese instruction manuals” and says interoperability standards should be owned and enforced by the government. His core industrial argument is that the US has to design missiles that can be built in car factories and aircraft that can be built in tractor factories, as it did in World War II, and that Arsenal-1 in Ohio is deliberately built like an auto plant so the government can nationalize the designs and farm them out in wartime. He describes Lattice as an open system with roughly 700 partner companies and over 100 integrated DoD platforms. Thunder, a hybrid-electric tiltrotor built with Archer, is pitched as a loyal wingman for attack helicopters. He argues deterrence only counts for force you are politically willing to risk, that drone threats to helicopters will be solved with close-in countermeasures like the Trophy system, and that Ukraine is a snapshot rather than the permanent future of war. He admits a lot of Anduril’s gear fails in truly adverse exercises, says the iPhone set a skill ceiling that too much software now copies, and lays out his case for subterranean warfare using narrow, autonomous boring vehicles. He closes with Heinlein, Jules Verne, and his dream of a 727 re-engined with afterburning Volvo RM8s.

    Thoughts

    The most interesting design idea in the first stretch is small but telling. Luckey points out (around the ten minute mark) that Robert Heinlein imagined computer-flown fighters in the 1940s, before anyone had a graphical display, so the pilots in those stories simply talked to the machines. He says that is how Anduril approaches Fury: don’t give the human pilot another computer in the cockpit, let them talk to the drone the way they would talk to a wingman. This matters because the hard part of collaborative combat aircraft may not be the airframe or the autonomy. It may be how much extra work the human has to take on. A wingman you have to manage through a tablet adds to the pilot’s workload. A wingman you can brief by voice is closer to the actual promise. Later in the interview he ties this to testing. Find out early whether the tablet is unusable under real workload, because after five years of development nobody will rip it out.

    The industrial argument in the middle of the conversation (roughly 18:00 to 24:00) is the part defense readers should keep. The usual story is that Detroit’s car plants were converted into tank and bomber plants. Luckey’s correction is that the US designed tanks and bombers around the welding, fasteners, bend radii, and workforce that car plants already had. That flips the question from “how do we build more aerospace capacity” to “what can we design that the capacity we still have can build.” He pairs it with a position that sounds strange coming from a founder: he expects the government to nationalize his designs in a real war and hand them to other manufacturers, he wants the government to own the IP on critical weapons, and he criticizes competitors who build capacity that nobody else could copy within ten years. Read next to his “patents are Chinese instruction manuals” line, the logic holds together. Protection only makes sense against allies, because adversaries ignore it anyway, so the better strategy is trade secrets, speed, and designs that can be copied at home.

    The deterrence point at 27:45 deserves more attention than it will get. “You only get credit for deterrence for strength that you are politically able to deploy.” China knows Congress will not park a carrier with 6,000 sailors inside anti-ship missile range, so the carrier deters less than its price suggests. Autonomous systems change that because an adversary can believe you will actually use them. The argument is uncomfortable because it implies part of the value of an unmanned fleet is that losing it is politically cheap. It is also a better argument for autonomy than the usual “take the human out of harm’s way” framing. It pairs with his drone point a few minutes later. People are overindexing on Ukraine, he says, where drones dominate because the countermeasures haven’t been fielded yet. A quadcopter can be killed with a shotgun, and a small gimballed gun on an Apache would change that math. Both arguments look at the adversary’s calculation, not the current headlines.

    The most honest moment is at about 36:40, when he says a lot of Anduril’s equipment “totally fails to work” in worst-case scenarios because “my guys are computer kids.” Right after that comes his long rant about touchscreens, and the two belong together. His complaint about the iPhone is not nostalgia. It is that the interface that made the phone learnable in five minutes also capped how good anyone could get with it, and then every app copied that trade-off. For an F-35 pilot at their 30,000th hour, or a helicopter pilot flying with the hydraulics out, at night, in weather they didn’t expect, the right interface is the one with the highest ceiling, not the gentlest learning curve. Defense software built by people raised on consumer apps will lean the wrong way unless someone forces the other question.

    Then there is the subterranean domain (42:00 to 47:00), which he knows sounds ridiculous and says anyway. The reasoning is more concrete than the laughter suggests. Crewed boring machines are huge because people are huge. Take the people out and the diameter can shrink a lot. In rock, diameter is the expensive part, while length is almost free, because anything that follows through an existing bore travels at no extra cost. His claim is that the energy needed to move that much earth fits within batteries, tethered power, or nuclear sources. Whether or not subterrines show up in his lifetime, his larger point is fair. Air power at sea was mocked, and careers ended over it. His freedom to say this comes from controlling Anduril’s voting shares, which is itself a quiet argument for founder control in defense tech. The line to remember is his last one: most of these things “are not waiting to be invented, they’re waiting to be implemented.”

    Key Takeaways

    • Luckey started Oculus because VR was the logical end state of PC gaming. After six monitors and multi-GPU rigs he saw a dead end and concluded the next step was presence, not more screens.
    • He did Oculus because he liked it. He started Anduril because he explicitly wanted his next act to be chosen for impact rather than fun.
    • After Facebook fired him, he considered three missions: zero-calorie foods to fight obesity, a nonprofit private prison chain paid only when people stayed out of prison, and national security.
    • He dropped obesity and prison reform because they were more political than technical problems.
    • His core worry was that the US tech industry had stopped working with the national security establishment, largely to stay in China’s good graces.
    • Before Oculus he worked at the USC Institute for Creative Technologies mixed reality lab on Bravemind, which used VR exposure therapy to treat veterans with PTSD.
    • He keeps a letter from the Secretary of the Air Force thanking his grandfather for flying as a civilian pilot in support of Desert Storm. He says that letter helped him choose Anduril.
    • Anduril has a public showroom and a separate one only the government can see, and products regularly move from the classified room to the public one.
    • The YFQ-44A Fury prototype has moved to serial production as the FQ-44. Luckey admits it is a vanity metric but wanted it to be the first production fighter with the F (fighter) and Q (unmanned) designation.
    • He argues much of today’s world was invented by older science fiction written by engineers who understood their craft. He calls most modern science fiction “space themed fantasy.”
    • Heinlein described computer-flown fighters and bombers in Astounding Science Fiction in the early 1940s. Anduril’s voice-driven approach to Fury echoes that, because pilots should talk to a drone wingman the way they talk to any other pilot.
    • Voice control only recently became good enough, not just at transcription but at understanding intent and turning it into something a computer can act on.
    • Jules Verne’s submarine, which rebuilt its batteries from minerals in seawater, points to what Luckey calls perhaps the most promising non-nuclear undersea propulsion idea today: using seawater as a reactant, the way an air-breathing turbine uses atmospheric oxygen.
    • Anduril sees itself as a product company. It funds development with its own money and sells finished products, instead of billing time, materials, and a fixed profit on top.
    • Under cost-plus contracting, the engineer who cuts a million dollars from production is penalized. In a product model, that engineer is rewarded.
    • Luckey concedes that some national capital assets, such as aircraft carriers, will probably stay cost-plus because there is only one buyer.
    • He believes everything should talk to everything, that no one should be allowed to build a proprietary silo, and that the government must own and enforce interoperability standards.
    • Oculus DK1 and DK2 were fully open-source hardware and software. His side company ModRetro has open-sourced its Game Boy and Nintendo 64 clones.
    • Anduril files very few patents. Luckey calls them “Chinese instruction manuals,” because they block Western allies from building on the technology while adversaries ignore them. Anduril relies on trade secrets instead.
    • The Anduril edition of the ModRetro Chromatic uses a sapphire screen lens, the same aluminum-magnesium alloy as Anduril’s attack drones, the low-IR Cerakote from the Ghost X helicopter drone, and titanium nitride on its connectors.
    • The consensus fix for US battlefield dominance is unified command and control, where every sensor serves every shooter across services and allies.
    • The real competitor is China and its partners. Chinese manufacturing equipment supports Iran’s attack drone supply chain and Russia’s weapons factories.
    • The US has to design missiles that can be built in car factories and aircraft that can be built in tractor factories, because automotive, agricultural, and some industrial plants are most of what is left.
    • In World War II the US did not simply convert car plants. It designed tanks and aircraft around the welding, fasteners, and metal forming those plants could already do.
    • Designing for common factories matters for two reasons: wartime scale-up, and deterrence, since adversaries weigh America’s total industrial capacity before acting.
    • Luckey credits organized labor with preserving most of the manufacturing that remains in the US.
    • Arsenal-1, Anduril’s roughly 5 million square foot plant in Ohio, deliberately looks more like an automotive factory than an aerospace one.
    • He expects the government to nationalize Anduril’s designs in a real war and farm them out. He supports government ownership of IP on critical weapons, despite leaning libertarian.
    • Some fielded systems that were contractually required to be interoperable were never actually tested. The documented calls simply don’t work.
    • Lattice is an open system with about 700 partner companies and integrations with over 100 existing DoD platforms. Government customers have integrated with it without talking to Anduril.
    • Thunder is a hybrid-electric, long-range, high-speed tiltrotor that acts as a loyal wingman for attack helicopters. It carries heavy munitions loads and vertical launch tubes for countermeasures and launched effects.
    • Luckey prefers jet fuel to batteries for now. Fuel burns off and can be dumped, which keeps emergency landing weights far lower than a battery aircraft that must carry its full mass into a crash.
    • Archer is building composite structures and drivetrain systems for Thunder, reusing components developed to FAA crewed-aviation standards for its civilian eVTOL.
    • You only get deterrence credit for strength you are politically willing to deploy. Adversaries increasingly believe only unmanned systems will actually be put at risk.
    • The Thunder launch video illustrates the concept, not the real concept of operations. Missile interception would happen miles out, not 100 yards ahead of the lead aircraft.
    • He expects radar-guided systems that shoot bullets out of the air, and more Trophy-style active protection on aircraft.
    • Drones are deadly to vehicles and helicopters today mostly because countermeasures haven’t been fielded yet. A small gimballed gun could protect an Apache.
    • Much of the resistance to automatic safety systems disappears when there is no human on board to be thrown around.
    • Exercises should be unscripted, overloaded, and degraded (damaged systems, weeks without maintenance, bad weather at night). Luckey admits a lot of Anduril’s gear fails under those conditions.
    • The iPhone made computing easy to learn but set a very low skill ceiling. Too many interfaces now optimize for the first five minutes instead of the 5,000th hour.
    • Anduril works in every domain, including space. It is working on space-based interceptors and has had AI on orbit since 2022.
    • Luckey believes subterranean warfare, with vehicles, people, and supplies moving through the Earth’s crust, is inevitable, and that autonomy makes narrow-diameter boring vehicles workable.
    • He points to the Soviet nuclear subterrine program, which he says lost its prototype underground, as evidence the problem is workable.
    • He can say radical things publicly because he is not in government and controls most of Anduril’s voting shares.
    • His dream aircraft is a Boeing 727, ideally a Valsan Super 27, re-engined with afterburning Volvo RM8 engines from the Saab Viggen, so he can do unlimited vertical climbs, including at Oshkosh.

    Detailed Summary

    From Oculus to Anduril: choosing impact over fun

    Asked, in a nod to the Mandalorian, how he knew defense was “the way,” Luckey traces Oculus back to a gamer’s question about what the final platform looks like. After building a six-monitor, dual-GPU setup, he concluded that more displays and more graphics cards led nowhere and that virtual reality, which tricks the subconscious into believing you are present, was next. Oculus made its investors hundreds of millions of dollars each, but he did it because he wanted to. When Facebook fired him about three years after the acquisition, he decided his next project would be chosen for impact. He weighed zero-calorie foods built on long-chain hydrocarbons to fight obesity, a nonprofit prison operator paid only for keeping people out of prison, and national security. He picked defense because the other two were political problems. He wanted to pull engineers away from building “augmented reality mustache emojis” and toward autonomous fighter jets and robotic submarines. His early work on the Army-affiliated Bravemind PTSD therapy project, and his grandfather’s letter from the Secretary of the Air Force, both fed into that decision.

    Fury, the FQ-44, and science fiction written by engineers

    After a tour of Anduril’s public showroom (the other showroom is government-only), co-host Sting Peterson asks about Fury’s path through the Collaborative Combat Aircraft program. Luckey notes the YFQ-44A prototype is now in serial production as the FQ-44, and that he wanted it to be the first production fighter with a proper F and Q designation. Both bond over the 2005 box office flop Stealth, which Peterson saw as a kid and which made her want to work in aviation. Luckey argues the world we live in was invented by older science fiction, written by NASA and aerospace engineers for whom the science mattered as much as the fiction. He jokes that his wife finds these novels unreadable because the characters only exist to deliver technical ideas. Heinlein wrote about computer-flown fighters and bombers in the early 1940s and imagined voice or punch-card commands with no displays at all. Luckey says that is effectively how Anduril is building Fury: you want to talk to it like any other pilot. Clay Mowry adds that AIAA’s forerunner, the American Rocket Society, was founded in the 1930s by science fiction writers and rocket enthusiasts whose work led to the engine on the Bell X-1.

    Jules Verne and seawater batteries

    One of Luckey’s favorite childhood books was Twenty Thousand Leagues Under the Sea, which he describes as a thin story wrapped around maritime technology. Captain Nemo doesn’t recharge his batteries. He rebuilds them from zinc and magnesium pulled from seawater, which Luckey calls a continuous underwater battery manufacturing system. Anduril isn’t building this, but he has looked at it. He notes that L3Harris bought the company doing the best work on saltwater-reactive lithium fuel cells, and that oxide buildup on the plates is the practical problem. Using seawater as a reactant is like an air-breathing turbine, which is why turbines beat rockets. The segment ends with Luckey singing “A Whale of a Tale” from the Disney film.

    Product company, open standards, and why patents help adversaries

    Luckey describes Anduril as a product company that picks what to build with its own money and sells finished products. Cost-plus contractors, by contrast, get paid for time and materials plus a fixed margin, which penalizes cost cutting. Some assets like aircraft carriers will probably stay cost-plus because they have only one buyer. He says everything should talk to everything and that the government should own interoperability standards. He is a longtime open-source advocate: early Oculus development kits were fully open, and ModRetro has open-sourced its Game Boy and N64 clones. On defense work, open-sourcing usually isn’t allowed, but Anduril files few patents because patents publish the design for adversaries who ignore IP law while blocking allies for the life of the patent. Anduril keeps its work as trade secrets instead, and if someone copies it and executes better, “they deserve to win.” He also describes the Anduril edition of the ModRetro Chromatic, which uses drone-grade alloy and Cerakote.

    Designing weapons for the factories America still has

    Asked what it will take to regain battlefield dominance, Luckey starts with the consensus answer, unified command and control and information sharing across services and allies, then moves to his more debated point. China supplies manufacturing equipment and support to Iran’s drone programs and Russia’s weapons factories, so the US has to design weapons its remaining industrial base can build: car plants, agricultural equipment plants, and some industrial plants. He says the World War II story of converting car factories is not quite right. The US designed tanks and aircraft around the welding, fasteners, bend radii, and heat-treatment processes car makers already used. That matters for wartime scale-up, since anything that needs hand-laid composites in a bespoke aerospace facility won’t scale, and for deterrence, since China should know GM could produce cruise missiles by the hundreds of thousands. Arsenal-1 in Ohio is built to look like a car factory on purpose. Luckey expects the government to nationalize his designs in a major war, criticizes companies that build capacity no one else can copy, and supports government ownership of IP for critical weapons, even though he leans libertarian.

    Interoperability that actually works, and Lattice

    On connectivity, Luckey says the government must actively enforce the standards it owns. Anduril has run into fielded systems whose contracts required interoperability, yet the documented interfaces were never tested and don’t work. On paper they are open. In practice they are silos. He pushes back on the idea that Lattice is closed: about 700 companies are in its partner program, it integrates with more than 100 existing DoD platforms and every messaging system and radio Anduril can get, and some government customers have integrated with it without involving Anduril at all.

    Thunder, deterrence, and the drone countermeasure gap

    Thunder is a hybrid-electric tiltrotor, not a pure electric aircraft. Luckey loves jet fuel because it burns off and can be dumped, while batteries force every emergency landing to carry their full weight, and the landing gear and crash structures that requires get heavy fast. Archer supplies composite structures and drivetrain components built to FAA crewed standards, which Anduril chose to reuse rather than redesign. If the CCA is a loyal wingman for fighters, Thunder is one for attack helicopters, a forward sensor and shooter that goes in before people do. Luckey argues deterrence only counts for force you will actually use, and no one believes Congress will risk a carrier and its 6,000 sailors inside Chinese missile range. He says the Thunder launch video illustrates ideas rather than tactics: interceptions would happen miles out, and the countermeasures would be canister, electronic, and kinetic rather than literal Anvil drones. He expects radar-guided systems that shoot down bullets, points to the Trophy active protection system on armored vehicles as a model for aircraft, and says drones threaten helicopters mainly because cheap close-in defenses aren’t fielded yet. People overindex on Ukraine, he says, which is “a reflection of a moment in time.”

    Trusting autonomy and testing in the worst case

    Peterson, who has worked on ground and air collision avoidance, asks how to build trust in AI and collaborative aircraft. Luckey, a helicopter pilot himself, notes that pilots dislike automatic systems that yank them around, and that problem disappears when nobody is on board. Commercial pilots also work within chauffeur-like constraints and avoid abrupt maneuvers, while a robot will take the most evasive action at the first sign of trouble. Simulators help, because a GPS-jamming scenario that kills most pilots can become one where everyone lives once safety systems are integrated. When Peterson points out that things that work in the sim often fail in flight, Luckey agrees that exercises are too scripted. He wants overloaded, degraded scenarios: systems shot out, three weeks without maintenance, hydraulics out at night in unexpected weather. He admits a lot of Anduril’s equipment fails in those conditions because its engineers are “computer kids,” which is why testing has to happen early, before a bad interface choice becomes five years of sunk cost.

    The iPhone skill ceiling rant

    Touchscreens set Luckey off. He respects Steve Jobs’s “bicycle for the mind” goal but argues the iPhone made computing so easy to learn that it capped how skilled anyone could become. A keyboard and mouse take thousands of hours to master but become a superhuman interface, like an Excel power user running macros at 150 actions per minute. The problem is not the iPhone itself. It is that everything became an iPhone, optimized for the first five minutes. He praises chorded keyboards and vector swipe keyboards as ideas that never caught on because nobody wants to invest hundreds of hours anymore. He wants technology designed for what an F-35 pilot can do on their 30,000th flight hour, not constrained by what Jobs showed on stage in 2007.

    Space, and the case for subterranean warfare

    Anduril works in every domain. It is publicly working on space-based interceptors and has had AI on orbit since 2022. The “weird one” is the subterranean domain. Luckey doesn’t mean tunnels or bunkers. He means vehicles, people, and supplies moving through the Earth’s crust as a three-dimensional battlespace. The US and Soviets both pursued subterrines. Autonomy removes the need for people-sized bores, and since diameter is expensive and length is cheap, the optimal design is very narrow and very long. The energy to displace or compact that earth, he says, fits within batteries, tethered power, or nuclear sources. He compares the ridicule to what early naval air power advocates faced, notes that his voting control of Anduril means no one can fire him for saying it, and cites the Soviet nuclear subterrine that was reportedly lost underground as proof the problem is workable. He retells the scene from The Core where a general shows a scientist a check and asks, “Would this be enough?” and says he wants the government to ask him that question. Mowry adds Journey to the Center of the Earth, and Luckey closes the thread by saying these ideas are waiting to be implemented, not invented.

    Jetson ONE, a Black Hawk, and the afterburning 727 dream

    Luckey was the first owner of a Jetson ONE, which he calls the Polaris of the sky: a short-range thrill ride with a redundant architecture that can lose about half its rotors and still land. He owns a UH-60 Black Hawk assembled from surplus parts on an FAA restricted certificate, bought before the Army began surplusing them cheaply, and a 1985 ex-Marine Corps Humvee bought when real ones were rare. His daily flyer is a Eurocopter EC120. His dream aircraft is a Boeing 727, his late grandfather’s favorite in 45 years at United Airlines, ideally a Valsan Super 27 conversion. His secret plan is to fit it with Volvo RM8 engines, the licensed, afterburning, thrust-reversing version of the Pratt and Whitney JT8D built for the highway-capable Saab Viggen, so he can request unlimited vertical climbs from the tower and fly it to Oshkosh with his grandfather’s 727 paperwork on board.

    Notable Quotes

    “I wanted to try to get people out of big tech and into work on national security problems with the same rigor and vigor that they were working on consumer electronics products and social media products.”

    Palmer Luckey, on why he founded Anduril after leaving Facebook

    “I often call patents Chinese instruction manuals. You’re just putting everything out there for an adversary to rip off.”

    Palmer Luckey, on why Anduril relies on trade secrets instead of patents

    “We need to design missiles that can be made in car factories. We need to design aircraft that can be made in tractor factories. And we’ve done this before. We did this in World War II.”

    Palmer Luckey, on rebuilding US defense production around the industrial base that remains

    “I fully anticipate that the government is going to nationalize my designs, farm them out to a whole bunch of other people. This is what we did during World War II as well. But we need to be building for that assumption.”

    Palmer Luckey, on why Arsenal-1 is built to look like a car factory

    “You only get credit for deterrence for strength that you are politically able to deploy.”

    Palmer Luckey, on why autonomous systems carry deterrent weight that crewed ones increasingly lack

    “People are overindexing on what warfare looks like in Ukraine. They’re saying, oh, this is the future of warfare. And I think it’s actually a reflection of a moment in time.”

    Palmer Luckey, on why drone dominance over vehicles and helicopters will fade as countermeasures arrive

    “A lot of our stuff totally fails to work when you get in those scenarios because my guys are computer kids.”

    Palmer Luckey, on the gap between scripted exercises and real combat conditions

    “The same interface that made it easy to learn to use also put a maximum skill ceiling on it that was very, very, very low.”

    Palmer Luckey, on the iPhone and the touchscreen-ification of everything

    “It is inevitable at this point. The only thing stopping us is that it sounds so crazy.”

    Palmer Luckey, on subterranean warfare as the next warfighting domain

    “They’re not waiting to be invented. They’re waiting to be implemented. And I’m less of an inventor and more of an implementer.”

    Palmer Luckey, on living in an age of unprecedented possibility

    Watch the full Up Next conversation with Palmer Luckey here.

    Related Reading

    • Anduril Industries official site covering Fury, Lattice, Arsenal-1, and the rest of the product line discussed here.
    • AIAA the American Institute of Aeronautics and Astronautics, host of the Up Next series and descendant of the American Rocket Society.
    • Arsenal of Democracy (Wikipedia) background on the World War II industrial mobilization Luckey wants to repeat.
    • Trophy active protection system (Wikipedia) the close-range kinetic countermeasure he expects to migrate from armored vehicles to aircraft.
    • Subterrene (Wikipedia) history of US and Soviet boring-vehicle concepts behind his subterranean warfare argument.