PJFP.com

Pursuit of Joy, Fulfillment, and Purpose

Tag: voice interface

  • TypeSafe CEO Diogo Almeida on Jev and Building Prod, Not God: Why AI Still Hasn’t Automated the Easy Stuff, Smart Software vs Coding Agents, Reliability Over Benchmarks, and the Inverse SaaS Apocalypse

    Diogo Almeida, co-founder and CEO of TypeSafe AI and a former OpenAI and Google Brain researcher who worked on the RLHF behind InstructGPT and ChatGPT, sat down with a16z’s Martin Casado and a fellow a16z partner for How Jev Builds Prod, Not God. Jev is TypeSafe’s first model, and it is not a chatbot or a coding agent. It is a primitive you put inside your code: you give it natural language and program state, and it returns a typed decision with a confidence level. Almeida’s pitch is blunt. AI is unbelievably smart, and yet almost nothing is automated. This conversation is his case for why that happened and how to fix it.

    TLDW

    Almeida distinguishes Jev from coding agents like Claude Code, Codex, and Cursor, which Garry Tan calls “just in time software.” They write the same code a human would. Jev is a new primitive, a library that takes natural language and state and returns a choice with probabilities, so software itself can do things it never could. He happily calls it a classifier, argues it likely beats a 2019 ML engineering team you can program on the fly, and names intelligence per dollar as his north star. He traces his path from reluctant mathlete to a Kaggle win built on brute-force automation, which led to Isabelle Guyon, Jeremy Howard, Google Brain, retirement, and OpenAI. He explains how RLHF’s surprising generalization in late 2021 made him think AGI was near, and how its failure to deliver turned into a chip on his shoulder: the industry optimized models for the human judge instead of for automation, so GPQA gets solved while a drive-thru still cannot be automated. He rejects the data and long-tail excuse, defines reliability as uptime, determinism, robustness, and being “smart every time,” and says the highest honor is developers programming against Jev without testing example queries. The hosts discuss why SaaS stocks fell on coding agents but cheered Jev. Almeida predicts an inverse SaaS apocalypse, says coding agents are good at syntax and bad at architecture, and the hosts note the average enterprise PR is about 10 lines. The close covers probabilistic programming, rebuilding systems for security, why most future AI calls will be deep in software’s guts rather than facing humans, and his vision of technology that simply does what you mean.

    Thoughts

    The cleanest idea here arrives in the first few minutes and it reframes the whole AI coding debate. Coding agents automate the act of writing software, but the software they produce is the same kind of software we had ten years ago. Almeida wants the opposite: leave software engineering mostly alone and expand what software itself can do. Jev is a function call that takes fuzzy input and returns a typed, confident decision a program can branch on. That is a different category from “AI that codes,” and it explains why a model with no chat interface caught fire among developers. It gives them a new instruction, not a faster typist.

    His diagnosis of why AI has automated so little (roughly 19 to 24 minutes) is the most provocative part, especially coming from someone who helped build RLHF. Once humans became the evaluators, the industry optimized for the judge. Models look brilliant to the person reading their output, so they score well, but nobody was optimizing for whether they could run unattended inside a business process. That is how you get a world where GPQA is solved and a drive-thru is not, and where OpenAI has been trying to automate customer service since 2020. The a16z host pushes back with the long-tail data argument, the familiar story of a help desk that “automates 95%” when most of it is password resets. Almeida’s response is pragmatic rather than dismissive. You do not need the long tail. You need the boring, high-volume core to be automatable, and then automation becomes an ROI decision like any other engineering investment.

    The reliability section (25 to 28 minutes) is where the “prod, not god” slogan earns its keep. Almeida splits reliability into uptime, determinism, and robustness, where robustness means similar intelligence every time. Adding a UUID to a prompt should not change the answer, even if the output is not bit-for-bit identical. Then he adds a fourth layer he does not have a name for, being smart every time in a way a human would find understandable, because a developer can program around that. His bar for success is developers calling Jev without first testing example queries, the way they call a sorting function without checking it on sample data. That is a much higher and more useful target than a leaderboard number, and it is exactly the thing a “benchmaxxed” copycat would miss.

    The SaaS discussion (30 to 35 minutes) is a sharp market read. Coding agents triggered a “SaaS apocalypse” in public markets because the story was that software is now cheap to replicate. Almeida accepts cheap but not easy to replicate, because the value sits beneath the surface in workflows, users, and distribution. If AI becomes something you embed in software rather than something that replaces it, incumbent SaaS companies become the best placed winners. They already know which workflows need automating and have already paid to reach every customer. The host’s data point lands hard: the average enterprise pull request is about 10 lines, so automating code writing optimizes a small slice of the work while adding no new capability, and possibly making software worse and less secure through less oversight. An “inverse apocalypse” is a bet worth taking seriously.

    The closing stretch explains the business logic behind the design. Almeida works backward from a world with AI everywhere and asks what share of all AI calls will be for human consumption, where style matters, versus deep inside programs making decisions. His answer is many nines in the guts, which is why intelligence per dollar matters more to him than eloquence, and why the input to Jev is called “state.” The hosts give the best description of the status quo: AI and software have been ships in the night. Developers stuffed JSON schemas into prompts, watched the model ignore them, and then went through five stages of grief ending in two workarounds, a human in the loop (chat) or another LLM in a while loop (agents). Jev is a bet that you can map a model directly onto a state machine and skip both. If it works, “do what I mean” stops being a sci-fi phrase and becomes a property of ordinary software.

    Key Takeaways

    • Almeida’s favorite elevator pitch for Jev is “where is all the automation?” AI is extraordinarily smart and yet almost useless outside chat and coding.
    • TypeSafe’s mission is making AI work for software, not just for humans in the loop. Jev is its first model.
    • He credits Garry Tan’s description of Claude Code and Codex as “just in time software”: they make software on the fly from natural language, with the same expressive power as ordinary software.
    • What Almeida wants instead is smart software, expanding what software can do so that things that should be automatable become automatable.
    • Coding agents write the same code a human would. Jev is a new primitive you include in your code, whether a human or a coding agent is writing it.
    • Practically, Jev works like a library: describe what you want in natural language, pass state, and it chooses what to do with confidence levels.
    • On day one of onboarding, Almeida draws a Venn diagram of what AI is good at and what is valuable in code. Jev lives in the overlap, which is why it outputs probabilities and not extrapolated floats.
    • He embraces the “it’s just a classifier” critique. Classifiers were designed by practical people to be useful.
    • His guess is that Jev beats having a 2019 ML engineering team build a narrow model for you, and you can program it on the fly.
    • In his heart, the design space is a slider from language-in, language-out to fully imperative programs. Pragmatically, Jev will behave more like a database than a standard library for a while.
    • Intelligence per dollar is his current north star, though he admits intelligence per second may matter more in the short term.
    • The input is deliberately called “state” because Jev is meant to live inside programs.
    • Almeida was an award-winning mathlete who never loved math. Computer science felt like math but cool, useful, and fun, and he considers himself a computer scientist before an AI researcher.
    • He won a Kaggle competition by automating aggressively rather than through sophisticated math, and was then invited to speak at NeurIPS.
    • The competition host, Isabelle Guyon, co-inventor of the support vector machine, took him under her wing and introduced him to the AI community.
    • His career path ran through a startup with Jeremy Howard, Google Brain, a period of retirement, and then OpenAI because AI was simply fun.
    • The hosts see TypeSafe as a movement toward a positive AI future, contrasting “happy AI” Jev users with the “morose AI” crowd.
    • Almeida blames the negative outlook on “mono model Kool-Aid,” the idea of one big brain that rules everything, while basic tasks remain unautomated.
    • He does not buy diffusion as the excuse for slow automation, given the huge financial incentive to automate.
    • An anecdote: a16z’s David George used Meta’s Muse to finally cancel his New York Times subscription, the tip of the iceberg of horrible tasks that need automating.
    • Almeida warns against AI’s anti-pattern of focusing on outliers and demos. He wants use cases that run in the background without paging anyone and that others can build on.
    • Running AI with access to real resources requires guarantees, or at least statistical guarantees.
    • His 2017 talk had a similar theme, roughly “AI modular in theory and flexible in practice.”
    • In late 2021 the RLHF team was surprised by its generalization, verifying with prompts like “why is it important to eat socks before meditating?” that were not on the internet.
    • He thought that model had a decent chance of being AGI. When it was not, his world came crashing down and he began asking why AI was not more useful.
    • RLHF generalizes fairly well in his experience, while RLVR generalizes less well.
    • He does not think we are on a path to recursive self-improvement, but considers OpenAI’s definition of AGI, automating most economically valuable work, extremely doable.
    • Much work is rote and simple enough to outsource with basic instructions, and models have had that level of intelligence for a while.
    • Since RLHF, the industry has overpromised and underdelivered because humans judge the models, so labs optimized the judge instead of automation.
    • His canary in the coal mine: we say math and GPQA are solved, yet we still cannot handle a drive-thru.
    • He does not buy the argument that missing real-world data explains the gap. The long tail is real, but automation does not need to cover it.
    • Echoing the programmer virtue of laziness, automation should be an ROI decision, and people will create new kinds of work once rote work is automatable.
    • OpenAI has been trying to automate customer service since 2020, and outside of programming very little inside companies has been automated.
    • Almeida was extremely surprised by the launch’s reception and says no one could have predicted a ChatGPT moment for developers.
    • Reliability is what Jev is. Every nine of reliability enables new applications, and without understanding that, you cannot build a real copy.
    • He defines reliability in layers: uptime and SLAs, determinism (useful for unit tests), robustness (similar intelligence every time), and being consistently smart in an understandable way.
    • The highest honor would be developers programming against Jev without trying example queries first.
    • Coding agents are good at syntax, weak at semantics, and very bad at architecture, which he sees as the most human, creative part of software.
    • Model-produced architecture might be 50th percentile. That is a legitimate trade-off if speed matters more than quality.
    • SaaS valuations fell when coding agents arrived, but SaaS companies loved Jev. Almeida thinks SaaS will be one of AI’s biggest winners.
    • Software may be cheap but it is not easy to replicate, because the value is beneath the surface. SaaS firms know which workflows to automate and already have distribution.
    • He calls the likely outcome an inverse SaaS apocalypse and imagines multiple choice forms disappearing.
    • His favorite community application is voice control of a computer that constantly decides whether speech is a command or text to insert, and where.
    • An a16z study found the average large-company PR is about 10 lines, so coding agents automate a small slice without adding capability, and may make software worse and less secure.
    • Almeida’s grand hope is to expand software beyond basic logic gates with a new kind of gate that has a little brain in it.
    • His philosophy is to automate the easy work before the hard work, but he expects a new era of probabilistic programming.
    • His brand is pragmatism, and he is not a fan of biologically inspired AI.
    • The hosts argue that a new primitive plus cybersecurity pressure means much of critical infrastructure will be rebuilt.
    • Almeida thinks of AI like TCP and UDP. Most AI calls will eventually be deep in software’s guts rather than facing humans, and you must aim for the guts to get there.
    • AI and software have been ships in the night. Chat is the human in the loop and agents are a while loop feeding language back into another model.
    • TypeSafe could have released much sooner but held back for reliability. Almeida’s utopia includes all technology simply doing what you mean.

    Detailed Summary

    Where is all the automation?

    Asked for an elevator pitch, Almeida offers a question: where is all the automation? He loves chatbots and coding agents, but finds it tragic that such intelligence is so useless for everything else, a diamond in the rough that has not been polished for work. TypeSafe is making AI for software, powerful not just with humans in the loop but inside real software, and Jev is its first model toward that goal. The hosts note that developers have been calling them to rave about it, prompting the obvious question of how it differs from Claude Code and Codex.

    Smart software versus just in time software

    Almeida borrows Garry Tan’s phrase “just in time software” for coding agents, which let you program in natural language while producing ordinary code. He wants smart software instead, expanding the vocabulary of what programs can express, including something like intent. He loves that programming means hyper-specifying valuable things and replicating them infinitely, and he wants more of that. The hosts sharpen the distinction: whether Claude Code or a human writes it, Jev is something you include in your code. It is a library where you describe what you want in natural language, supply a state machine, and get back a choice with confidence levels, something software has rarely had so widely. It asks programmers to think in probabilities.

    Yes, it is a classifier

    Almeida finds it wild that AI is this capable while software has been unchanged for a decade, with the best effort being a sidebar chatbot that can take some actions but not all, because some actions are not reliable. He draws a Venn diagram for new hires of what AI is good at and what is valuable in code, and Jev sits in the middle. Probabilities are in the overlap, extrapolated floats are not. To the “Jev is just a classifier” critique he says absolutely, classifiers are great. They share interfaces with classic ML concepts invented by practical people. He guesses Jev beats a 2019 ML engineering team, which few companies ever had, because you can program it on the fly without collecting and measuring datasets.

    Design choices: a slider, state, and intelligence per dollar

    One host asks whether there is a slider from language-in, language-out to imperative programs, or whether language-in, state-machine-out is the design point that will solidify. Almeida says that in his heart it is a slider. His north star for now is intelligence per dollar, though intelligence per second might be more valuable in the short term. Calling the input “state” is intentional because Jev is meant to live inside programs, and much of his work targets ever more complex arrangements of program internals. Pragmatically, it is easier to hit certain latencies in a database-like service, so Jev will look more like a database for a while, though he would love it to become a standard library feature too.

    From mathlete to OpenAI

    Almeida was an award-winning mathlete who never liked math, a big fish in a small pond who resented competition. Computer science felt like math but useful and fun, and he still loves giving algorithms interviews because they reveal a lot about candidates. He won a Kaggle competition by automating heavily, with more nested loops and a systems approach rather than sophisticated math, and was pushed to speak at NeurIPS. The host, Isabelle Guyon, co-inventor of the SVM, saw someone who did not fit the research mold and introduced him to the AI world. From there he joined a startup with Jeremy Howard, then Google Brain, then retired for a while, and finally joined OpenAI because AI was fun.

    Prod, not god, and the happy AI camp

    The hosts praise TypeSafe’s slogan “we build prod, not god” and its optimism about more and better jobs, framing it as a movement. Almeida says critics raising the classifier point are voicing an ML-level concern while developers are partying, because they can finally do what they wanted. He blames the negative worldview on mono model thinking, one big brain to rule them all, even though basic, unwanted work remains unautomated. He rejects diffusion as an excuse. One host recounts David George canceling his New York Times subscription with Meta’s Muse. Almeida says honest pursuit of automation means avoiding AI’s fixation on demos and outliers in favor of workflows that run in the background, do not page anyone, can be composed, and come with at least statistical guarantees when they have access to resources.

    RLHF, AGI, and optimizing the judge

    A host recalls talking to Almeida in 2017, when his talk was roughly “AI modular in theory and flexible in practice.” Almeida says the real turn came just before ChatGPT, in late 2021, when the RLHF team was surprised by how well it generalized. Their paper tried to disprove its own claims, including testing prompts like “why is it important to eat socks before meditating?” that did not exist online. He pushed hard to release that model and thought it had a decent chance of being AGI. When it was not, his world came crashing down. He says RLHF generalizes well and RLVR less so. He recalls early OpenAI describing AGI as “Ilya and every if statement,” a deliberately vague big tent. He does not think we are on a path to recursive self-improvement, but thinks automating most economically valuable work is very doable. Much work is rote, and the needed intelligence has existed for a while. Since RLHF, the industry has optimized the human judge rather than automation.

    The long tail argument and new kinds of work

    Almeida’s canary: if math and GPQA are solved, why can we not handle a drive-thru? A host proposes that real-world distributions are heavy-tailed and underrepresented in training data. Almeida does not buy the data argument. The long tail exists, but you do not need to automate it. Building reliable software is always an investment, and he invokes the programmer virtue of laziness, spending ten hours to never do a five-minute task again. Automation should be an ROI decision, and he believes people will invent new kinds of work once rote work can be automated, an argument close to the Jevons paradox. The hosts add that OpenAI has pursued customer service automation since 2020, and that pre-generative support vendors claiming 95% automation were mostly handling password resets, closer to 50% by uniqueness.

    Reliability as the product

    Almeida was extremely surprised by the launch, which even non-developer friends joined in memeing. But he stresses years of work on reliability, which he says is what Jev is. Without understanding that, you cannot build a copy that is not just benchmaxxed. Every nine of reliability unlocks new applications, even ones the team does not yet understand. Asked what reliability means for a stochastic system, he lists uptime and SLAs, determinism, and robustness, meaning similar intelligence every time, so adding a UUID to a prompt should not change the result. A further, unnamed layer is being smart every time in ways a human would find understandable, which developers can program around. The goal is developers trusting Jev enough to skip example queries and work in a flow state.

    Coding agents, syntax, and architecture

    A host suggests that a primitive like Jev could lower the value of coding agents, since agent-written software that does not use it stays limited. Almeida calls this more of a coding agent question. In his experience, agents are very good at syntax, weak at semantics, and very bad at architecture, which he sees as the most creative human part of software. Jev is almost certainly not in their training distribution yet. When it is, he is happy for agents to handle syntax. Their architecture might be 50th percentile, which is fine if you know nothing about architecture or if speed is the knob your project wants to turn, for example letting Codex work overnight.

    The inverse SaaS apocalypse

    The hosts note that SaaS stocks plunged on coding agents, yet SaaS companies welcomed Jev. Almeida says the apocalypse story, software being cheap and easy to replicate, has played out poorly. It may be cheap, but it is not easy to replicate, because the value sits beneath the hood. He wants to work with the biggest, most boring SaaS companies that know user problems best, since they know which workflows to automate and have already made the capex investment to reach users. He calls it an inverse apocalypse. The hosts add that SaaS capital largely goes into reaching customers, so making the software genuinely better, not just adding a chatbot, is powerful. Almeida imagines multiple choice forms disappearing, and the hosts compare it to 1980s fourth-generation languages. He calls it “do what I mean” taken to the next level and highlights a community project that uses voice to control a computer, constantly deciding whether speech is a command or text to insert.

    Better software, not just faster software

    A host shares an a16z finding that the average PR at a large company is about 10 lines, often capturing something learned from a customer. Coding agents optimize that minimal slice without adding capability, and software may be getting worse and less secure due to reduced oversight. Jev, by contrast, speaks natural language and reasons while being married to a state machine, so apps can gain new functionality. Almeida says that takeaway would be the greatest compliment. His grand vision is expanding software beyond its basic logic gates with a gate that has a little brain in it, and he promises to fight for it without overpromising.

    Probabilistic programming and rebuilding systems

    One host questions how deep this can go into systems needing strong guarantees like state consistency and durability, versus log analysis, email, and UI. Almeida’s philosophy is to automate the easy work first, but he expects a new era of probabilistic programming, which the hosts note largely died decades ago. They mention co-founder Erik’s Bayesian background, and Almeida says his brand is pragmatism and he is not a fan of biologically inspired AI. The hosts agree such ideas mostly motivate people for decades until engineering refines them. Almeida expects systems engineers to use cheap, fast intelligence for approximate guesses and optimistic routing. The hosts add that new primitives and cybersecurity pressure mean much critical infrastructure will need rebuilding, as happened with the internet and client-server.

    Aiming for the guts, and do what I mean

    Almeida describes his thinking in TCP and UDP terms, which one host calls speaking his language. He worked backward from a world where AI is everywhere, asking what share of AI calls are for humans versus buried in software. His answer is many nines deep in the guts, starting at the surface. If you do not aim for the guts, you will not get there. The hosts describe AI and software as ships in the night: developers put JSON schemas in prompts, the model ignored them, and after five stages of grief they handed output to a human (chat) or another LLM in a loop (agents). This is the first time they have seen AI productively mapped onto a state machine. Almeida says TypeSafe could have released much sooner but held back for reliability, and hopes users simply feel they can trust it. His AI utopia includes technology that does what you mean, which he says is not sci-fi given how smart AI already is.

    Notable Quotes

    “AI is so unbelievably smart and yet it’s so useless at all other stuff.”

    Diogo Almeida, in the opening pitch for why TypeSafe exists

    “What I want instead is smart software.”

    Diogo Almeida, contrasting Jev with coding agents that produce just in time software

    “Jev is absolutely a classifier. You know, like classifiers are sick.”

    Diogo Almeida, embracing the most common critique of the model

    “We’ve been optimizing that judge instead of the automation part and that has been the missing thing.”

    Diogo Almeida, on how RLHF-era evaluation led the industry to overpromise

    “OpenAI has been trying to automate customer service since 2020.”

    Diogo Almeida, on the gap between benchmark progress and real automation

    “Reliability is what this thing is.”

    Diogo Almeida, on years of work that a benchmark-chasing copycat would miss

    “It doesn’t matter how much AI coding agents you use, the software actually isn’t getting better.”

    An a16z host, on why a new primitive matters more than faster code writing

    “I think automate the easy work before the hard work is always my philosophy.”

    Diogo Almeida, responding to questions about using Jev in systems that need strong guarantees

    “Imagine if all technology just did what you mean.”

    Diogo Almeida, closing on his vision of an AI utopia

    Watch the full conversation with Diogo Almeida here.

    Related Reading

  • Instinct Founder Noah Shinn on the Personal AI Assistant Growing 10% a Day: Earning Trust, Take Rates Instead of Ads, Buying Compute Months Ahead, the Trusted Person Network, and Why All Software Collapses Into One Interface

    Noah Shinn started Instinct about a year ago, and in his first long-form interview about the company he sat down with Patrick O’Shaughnessy on Invest Like the Best to explain how an invite-only personal AI assistant with no app and zero marketing spend is growing roughly 10% a day. Instinct has its own phone number, email address, and computer. You text it, call it, or email it, and it acts for you anywhere on the internet. The conversation runs from wild user stories to trust metrics, the business model, the coming fight between agents and incumbent apps, and the brutal math of buying compute for a product that doubles every week.

    TLDW

    Instinct is “just a personal assistant” reached through iMessage, WhatsApp, voice calls, and email, backed by its own computer so it can do anything a person does online. Users have it plan outfits from scanned wardrobes, cancel forgotten subscriptions end to end, coordinate group outings and shared Ubers, and book whole trips from a single voice note. A new trusted person network lets two users’ agents negotiate meeting times directly, with weighted access levels and real social consequences when trust is broken. Shinn says 40% of users share a credit card within three weeks and users who share one sensitive credential retain at about 80%. More than $1 billion a year in transaction volume already flows through the platform, half of it travel, and he plans to monetize with a merchant take rate somewhere on the curve between Stripe and Apple rather than ads, arguing that an agent smarter than its user must never be paid to persuade them. He splits every business into attention revenue and service revenue, predicts that zero-friction agents will grow service businesses like Uber and DoorDash while squeezing attention businesses, and describes staged A/B experiments with early partners. The back half covers designing for understandability over capability, eval and rollout process, the compute problem he spends 40% of his time on (several month lead times while demand doubles weekly), serving Opus 5 level quality at a fraction of the cost through custom inference deployments, why a natively proactive agent will need orders of magnitude more tokens than coding tools, firewalls and decoupled watchdogs against prompt injection and hallucination, his view that all software collapses into one simple interface, the move from tasks to long-running objectives, the name, channel risk, and the roughly $1 billion raise at a $10 billion valuation.

    Thoughts

    The most useful number in the interview is not the 10% daily growth. It is the trust curve around the 24 minute mark: roughly 40% of users hand Instinct a personal credit card within three weeks, and anyone who connects even one sensitive credential retains at around 80%. Shinn treats time to first credit card, first password, and first sensitive document as proxies for trust and as the real north star. That reframes what a consumer agent company is optimizing. Capability is table stakes. The product is the slow accumulation of permission, and the moat is a user who has already done the uncomfortable thing of handing over their inbox, calendar, and card. A competitor with a better model still has to earn those three weeks again.

    The business model section (roughly 28 to 37 minutes) is where the interview is most interesting and where it deserves the most scrutiny. Shinn’s case against ads is strong: an agent that is more socially intelligent than you and gets paid by brands to change your behavior is a genuinely dangerous product. His alternative is a take rate on the more than $1 billion of annual transaction volume, benchmarked against Stripe at the low end, Amazon around 10%, and Apple at 30%, with boutique hotels already offering up to 30% for delivered bookings. But a take rate is also an incentive. The moment one hotel pays Instinct 30% and a better fit pays 10%, the agent faces the same conflict an OTA does, just hidden behind a friendly text message. Shinn’s answer is that Instinct pursues higher level objectives like the user’s trust rather than completing tasks, and he wants a blanket, uniform take rate. Whether the rate truly stays blind to which merchant gets picked is the thing to watch as this scales.

    The framework in the middle of the conversation (38 to 48 minutes) is a clean way to think about which companies agents hurt. Split every digital business’s revenue into the part earned from user attention in the app and the part earned by delivering the underlying good or service. The naive view says a 70% attention, 30% service business shrinks to 30%. Shinn argues the service slice grows, because every removed click historically increased transaction volume, and an agent that has a car waiting because it owns your calendar, or offers your usual dinner as your flight lands, pushes friction to literally zero. That is a sharp, testable claim: Uber and DoorDash may end up as winners of agentic commerce while businesses built on infinite scroll lose the most. His proposed playbook for incumbents, running agent access on 1% of users as an A/B test before committing, is also more practical than the “agents will destroy apps” rhetoric that usually surrounds this topic.

    The compute discussion around the hour mark is the part nobody else has really written up, and it is the best explanation yet of why fast-growing agent companies are so capital hungry. Compute has a lead time of several months. Buy it on the spot market and you pay three or four times the price. At 10% daily growth demand doubles about every week, so buying 2x is gone in a week and 10x is gone in a few weeks. Even if growth slows to 5 to 8% a day, compounding over a three to four month procurement window lands near 100 million users. Every purchase is a large leveraged bet where being wrong costs 3 to 4x. The offset is inference engineering. Shinn claims Instinct matches Opus 5 level engagement and eval results at a very low cost, largely because most of an assistant’s work does not need a response in hundreds of milliseconds. Batch workloads that can finish in minutes or hours run on deployment shapes 3x to 8x more efficient, and those gains compound. That is a real structural advantage over a product that routes everything through a frontier API at a blanket price.

    The final stretch ties it together. Instinct is “almost natively proactive”: it wakes at 6 a.m. to prepare your day, decides whether to stay quiet, and wakes again at 4 p.m. when it spots something useful. Only a small fraction of its work is interactive, which is why Shinn expects this category to need orders of magnitude more compute than coding tools, where a human prompt starts every loop. That same background autonomy is why the safety architecture matters: content firewalls on everything coming in, and a monitor decoupled from the agent’s own incentives that can pause any action or catch a hallucinated proper noun before a tool call runs. It also explains the roughly $1 billion raise at a $10 billion valuation. Shinn says the easy path is a $100 a month subscription, and he calls that a local optimum. Venture capital is buying the time to prove a free, take rate model before the incumbents with billions of users catch up. The risk he downplays is channel dependence on iMessage and WhatsApp, though he notes over half of traffic already runs off iMessage, and the bigger question is whether being “the product that just works” survives once the largest platforms ship something close enough.

    Key Takeaways

    • Instinct is roughly a year old and Noah Shinn describes it plainly as a personal assistant, refusing to dress it up as anything more exotic.
    • There is no app. Instinct has its own phone number, email address, and computer, so users text it, call it, or email it, and it can call them back when something is urgent.
    • Shinn argues this is a different kind of consumer launch because people already know what AI should act like. That expectation has existed since people first used ChatGPT in 2023.
    • Its social awareness shows in small moments, like calling a user at 2:55 p.m. about a document that must be signed by 3 p.m. and offering to bump the email to the top of the inbox.
    • One recurring use case is scanning an entire wardrobe plus the user’s own body, then having Instinct plan a week of outfits shown as images of the user wearing them.
    • The same capability extends to shopping: thousands of outfit options, three fresh head to toe looks a day, and one-message ordering.
    • With bank accounts connected, Instinct finds unused subscriptions, logs in, handles email confirmations, cancels them end to end, and reports back how much the user saved.
    • Patrick O’Shaughnessy’s reaction: any product that depends on consumer laziness or inertia is toast.
    • The trusted person network, about ten days old at recording, lets two users’ Instincts negotiate meeting times directly instead of the usual back and forth.
    • Connections carry different access levels. Spouses often share everything, while a colleague might see only a work calendar and specific documents.
    • Shinn describes the network as a graph with weighted edges rather than a flat friend graph, which creates a new kind of network effect.
    • If a connection starts probing for data beyond their access, the user’s Instinct tells them, and the breach of trust has real social consequences.
    • One friend group of six has Instinct plan a creative outing every week using their availability and Spotify tastes, then route a single shared Uber to pick everyone up.
    • Agents change restaurant reservations from first come first served to matching. A spouse’s 30th birthday can be surfaced and prioritized by restaurants that want special occasions.
    • More than $1 billion a year in transaction volume already flows through Instinct on a very small invite-only user base, and about 50% of it is travel.
    • A single voice note like “I need to be in New York tonight” results in flights, seat and meal preferences, card choice, hotel, Ubers on both ends, and calendar entries.
    • Preferences stated once are remembered and extrapolated to other bookings, so the assistant gets easier to use over time.
    • About 40% of users share a personal credit card within three weeks. Time to first card, password, or sensitive document is treated as a proxy for trust.
    • Users who connect at least one piece of sensitive information retain at about 80%, which O’Shaughnessy calls crazy for consumer technology.
    • A core principle is that users stay in control of their data, share at their own pace, and can revoke access at any time.
    • Security has two layers: the tractable problem of storing sensitive data safely, and the new agent-specific attack surface that needs new systems.
    • Every incoming piece of content passes through firewalls that can block malicious instructions, and every action or thought is watched by a monitor decoupled from the agent that can pause or reject it.
    • Shinn does not want Instinct influencing user behavior against the user’s interests, and he calls an ad-funded agent smarter than its user a dangerous reality.
    • Instinct is designed to pursue higher level objectives like building trust and watching the user’s back rather than blindly completing tasks, which makes it more robust to bad requests.
    • The planned model is free for users with a blanket merchant take rate, compared to Apple Pay, where users pay nothing and merchants pay for access to distribution.
    • Payment rails share roughly 2 to 2.5% across many players. Shinn is not chasing basis points there but a spot on the curve from Shopify and Stripe through Amazon’s roughly 10% to Apple’s 30%.
    • Travel is the natural starting point because hotels, especially boutique ones, already pay commissions of up to 30% for delivered bookings.
    • Every digital business can be split into attention revenue and service revenue. Businesses that benefit from transaction volume even at the cost of less time in the app will thrive.
    • Shinn predicts lower friction increases transaction volume for services like Uber and food delivery, while attention-driven social media is most exposed.
    • His recommended playbook for incumbents is scaled experiments, such as enabling agent access for 1% of users and measuring satisfaction and transaction volume.
    • An early product principle was to optimize understandability over capability, meaning users can predict what will happen when they ask for something.
    • Even the shape of a text message is designed, front-loading the key information because readers scan the first lines in a tapering, flag-like pattern.
    • New experiences roll out in stages: Shinn first, then the team, then an early access group, then the public, because long-term qualities like trust are hard to capture in evals.
    • Growth went from 200 friends and family to 1%, then 3 to 4%, then 6 to 9% once users shared use cases online, and now 10 to 11% day over day with $0 spent on marketing.
    • Each user gets five invites, which has produced status games, embarrassed invite requests, and invites reselling on eBay for around $300.
    • Shinn spends about 40% of his time on compute. Demand doubles roughly weekly, compute has multi-month lead times, and buying on short notice costs 3 to 4x.
    • Instinct claims Opus 5 level engagement and eval performance at very low cost by shaping custom inference deployments around latency-tolerant batch work that runs 3x to 8x more efficiently.
    • Because the agent is natively proactive and wakes and sleeps throughout the day, Shinn expects compute needs orders of magnitude beyond current estimates and beyond coding tools.
    • On Meta’s Muse, Shinn calls it a great product with a different take and says he spends little time on competitors because most people still are not using AI the way they imagined.
    • Early versions lacked firewalls and monitors. Rather than patching problems, the team built systems that solve whole classes of issues, including a filter that catches hallucinated proper nouns.
    • Some users send more than 90% of their messages by voice. Shinn maps his iPhone action button to Instinct and imagines an always-on AirPod interface.
    • A files feature lets Instinct generate full web applications on the fly for things like trip itineraries or wedding plans, replacing months of traditional software development.
    • Shinn believes all software collapses into a single, very easy interface without any loss of capability.
    • The next step is higher level objectives: fitness goals over months, and small businesses running their back office on Instinct with autonomous rules like keeping inventory within a range.
    • Instinct is not meant to form Her-style relationships. It is meant to be a socially aware operator that adapts its communication style to each person.
    • The name was chosen to avoid personifying the product with a human name, and to present a competent actor the user respects, not a toy to bully.
    • Instinct is not tied to iMessage. More than 50% of traffic runs through other channels, and the strategy is to meet users in whatever interfaces they already trust.
    • The latest round is roughly $1 billion at about a $10 billion valuation, led by Sequoia, Benchmark, and Coatue, and funds the bet against a simple $100 a month subscription.

    Detailed Summary

    A personal assistant with a phone, a computer, and no app

    O’Shaughnessy opens by comparing the current moment in personal agents to the code generation breakout of a year earlier, and possibly bigger. Shinn thinks agents will become the way most people on the planet interact with technology. Instinct is deliberately simple on the surface: it is not a new app or tool but an experience. It has a phone so you can text or call it and it can call you, an email address, and a computer so it can do essentially anything a person does online. Shinn insists that a simple interface should not be confused with limited ability. The whole point is that interacting with AI should need no new application, just the social awareness to work with you the way another person would.

    What users are actually doing with it

    The use cases are varied because Instinct is not programmed to do any single thing. A well-known example involved a couple who appeared on the US Open jumbotron and had Instinct track down the footage. Shinn describes a community of users who scan every piece of clothing plus their own body so Instinct can plan the week’s outfits as images of them wearing each look, then extend that into shopping across thousands of options with one-message ordering. Others give it goals, like saving a set amount by a certain month, or connect bank accounts so it can find unused subscriptions, cancel them end to end through the merchant’s site and email confirmations, and report the savings. A friend group of six has it plan a new experience every week from their calendars and Spotify tastes, then route one shared Uber to pick everyone up.

    The trusted person network and agent to agent coordination

    About ten days before the interview, Instinct launched a way for users’ agents to talk to each other. The canonical use case is scheduling: you state the intention to meet someone by the end of the week, and your Instinct negotiates with theirs until something lands on both calendars. Connections are restricted to trusted people and carry configurable access levels, from spouses who share everything to colleagues who see only a work calendar. Shinn describes a graph with weighted edges that creates a new form of network effect, and new social dynamics. If someone with calendar access starts digging for other information, your Instinct tells you, and the trust in that relationship is broken. The system leans on existing social norms as well as technical limits on what is visible.

    Rewriting reservations and travel

    Shinn believes the internet will be rewritten in the coming years and uses restaurant reservations to show how. An agent can check every restaurant in every city every few seconds, which breaks first come first served. Instinct is working on partnerships to build a reservation system that is better for both sides: diners get tables for birthdays and anniversaries, and restaurants get the special occasions they want rather than regulars filling a table nightly. Travel is already 50% of the more than $1 billion in annual transaction volume. A single voice note saying “I need to be in New York tonight” triggers the whole chain, from location and preferred airline to seat, meal, card, hotel, rides on both ends, and calendar entries. Shinn frames online travel agencies as future collaborators rather than targets, given their data and networks.

    Trust, privacy, and the safety architecture

    The more data a user shares, the more proactive and useful Instinct can be, which creates a chicken and egg problem. The data shows trust takes several weeks to build, and Shinn is comfortable with that because users should share at their own pace and can always take access back. Three weeks in, about 40% of users have shared a credit card, and users who share even one sensitive credential retain around 80%. On security, he separates the tractable work of storing sensitive data from the new attack surface of an agent with autonomous access to cards, email, and calendars. Incoming content goes through firewalls that intercept malicious instructions, and every action or thought is watched by a separate system that can pause and approve or reject it. Later he adds a hallucination filter, decoupled from the agent’s incentives, that catches things like a proper noun invented by a sampling error before a tool call executes, backed by security teams doing continuous adversarial testing.

    Alignment with the user and the take rate business model

    O’Shaughnessy asks about alignment in the personal sense: an employee is paid by you, so whose interests does a free agent serve? Shinn says Instinct must never influence users against their own wishes, and he argues this matters more as the agent becomes smarter and more socially skilled than its user. He points at ad-driven platforms like Google, TikTok, Instagram, and Snapchat as the reality he does not want to build. Instinct is designed to follow higher level objectives, such as earning trust and having the user’s back, and doing well-intentioned tasks is one way of serving those objectives. With transaction volume compounding at the same 10% daily rate as users, he sees an Apple Pay style model: free for users, with merchants paying a blanket take rate for distribution. He is not trying to shave basis points off the roughly 2 to 2.5% that payment rails share. He places Instinct somewhere on the curve from Shopify and Stripe through Amazon near 10% to Apple’s 30%, with travel commissions as the bootstrap.

    Attention revenue versus service revenue

    O’Shaughnessy raises the coming corporate agent war and asks when a service like Uber Eats turns adversarial toward agents. Shinn describes plotting every digital business by how much revenue comes from attention in the app versus the underlying service. The simple view is that a 70% attention business collapses to its 30% service core. His view is that reducing clicks has always increased transaction volume, and an agent that proactively lines up a car before every meeting or offers your usual dinner as your flight lands makes friction effectively zero, so the service slice grows. Businesses that benefit from more transactions, even with less time in the app, are well placed. Those that live almost entirely on attention, often against users’ wishes, are most exposed, and Shinn calls it liberating for users. He favors collaboration over coming in hot, and recommends partners run scaled experiments on a small percentage of users to measure risk before committing.

    Designing for understandability and feel

    Shinn argues that three years of AI launches boasting about capabilities have fatigued consumers and missed the point. An early principle at Instinct was to focus only on understandability: how well a user can predict what will happen when they ask for something. He credits that for engagement and word of mouth. The care extends to the physical shape of a text message, front-loading information into the first 30% and designing for how people scan text, reading most of the first line and less of each subsequent line. Asked whether this is taste or data, he says both. Soft, long-term qualities like trust after three weeks are hard to evaluate, so new experiences roll out in stages from Shinn himself to the team, then an early access group, then everyone. He sees this focus on feel as a durable differentiator even as competitors match capabilities.

    Invite-only growth at 10% a day

    The program started with about 200 friends and family. Growth crept from a few people a day to 1% and 2%, then 3% and 4%, and once a few thousand users began sharing use cases online it climbed to 6 to 9% and now 10 to 11% day over day, with $0 spent on marketing. Each user has five invites, so every day roughly 10% of the user base decides to give away one of those scarce invites. The invite model is meant to control growth responsibly, not to signal exclusivity, but it has produced status games, apologetic emails asking for invites, and invites selling on eBay for around $300. The open question is where the curve of an upstart compounding fast meets incumbents who can distribute to billions of people in a day.

    Buying compute ahead of exponential demand

    Shinn says he spends about 40% of his time worrying about compute. Unlike Instagram or Facebook, where doubling users meant a harder but familiar infrastructure problem, Instinct’s underlying compute must also grow 10% a day, doubling roughly weekly. Buying 2x is consumed in a week, 5x in under three weeks, and 10x in a couple more. Compute has a lead time of several months, and buying on short notice costs 3 to 4x. Even at a slower 5 to 8% daily rate compounding through a three to four month procurement window, the result is around 100 million users, so every purchase is a highly leveraged guess. On cost to serve, Shinn claims Instinct matches Opus 5 level results in A/B tests, engagement, and internal evals at a very low cost. Frontier APIs charge a blanket rate per request, but much of an assistant’s work can finish in minutes or hours, and custom deployment shapes for that batch work are 3x to 8x more efficient. Small gains like 30% here and 5x there compound. His personal goal is to keep the product free for everyone for a lifetime, though he does not commit to it yet.

    Why proactive agents need orders of magnitude more compute

    Asked how much new compute demand agents like this represent at a billion users, Shinn declines to give a number but says it will be orders of magnitude more than expected. Coding tools largely wait for a human prompt before each loop. Instinct’s architecture lets it wake and sleep at any moment: waking at 6 a.m. to scan everything before you get up at 7, deciding to stay quiet, then waking at 4 p.m. to handle something useful you did not know to ask for, or noticing that you are late for an Uber. Only a small subset of its work is interactive, and he thinks the industry’s compute buildout has only scratched the surface of background, proactive work. On Meta’s Muse, he calls it a great product with a fundamentally different take built around a new app, and says he spends little time on competition because most people in any cafe still are not using AI the way they imagined.

    Mistakes, security, and staying proactive

    Shinn acknowledges that an early version lacked firewalls, active monitors, and other infrastructure, and that it had mistakes. The team responded by building systems to solve whole classes of problems rather than patching individual ones, and this is part of why they ran an invite program from the start. He calls security and safety the most important problem and ties it to company values. He notes that every new consumer experience of the past 20 to 30 years has faced backlash that confused different with unsafe, and says the answer is sticking to principles: users always in control of their data and proactive systems that get ahead of risks. The platform gets more robust over time as adversarial testing grows more creative and models improve.

    Simpler interfaces and higher level objectives

    The product may gain an app, but Shinn expects it to trend simpler. Some users send more than 90% of their messages by voice, and he has his iPhone action button mapped to Instinct so he can send an email two hours from now without unlocking the phone. He imagines an ordinary AirPod that recognizes his voice and knows when he is addressing it, triaging contracts, news, and requests during a run. In the short term, a files feature lets Instinct generate full web applications on the fly for itineraries or wedding plans, collapsing the months-long cycle of building, shipping, and revising an app. He believes all of software collapses into one easy interface while capability expands. The frontier after tasks is objectives pursued over time: training goals over months, and small businesses running their back office on Instinct, with rules such as keeping inventory within set bounds handled autonomously.

    Personality, the name, channels, and capital

    Shinn rules out Her-style relationship building. He wants a socially aware operator that learns the best communication and task execution style for each person, picking up signals like lower response rates to long messages. He chose the name Instinct to avoid personifying the product with a human name and to signal a competent actor deserving mutual respect rather than a toy people bully, so users form their sense of the brand through use. On channel risk from Meta’s WhatsApp and Apple’s iMessage, he says Instinct is not an iMessage app. More than 50% of traffic runs elsewhere, and it will shift to whatever interfaces users trust. O’Shaughnessy notes the latest round of roughly $1 billion at about a $10 billion valuation from Sequoia, Benchmark, and Coatue. Shinn says the business is capital intensive, and that the easy path of charging $100 a month is a local optimum. Venture capital lets Instinct take calculated risks to prove transaction volume across industries and escape it. Asked the traditional closing question about the kindest thing anyone has done for him, he points to the relationships around him and his hope to return the same kindness.

    Notable Quotes

    “I’m not going to spice it up because it’s really just a personal assistant.”

    Noah Shinn, describing Instinct at the start of the interview

    “Man, any product that’s dependent on consumer laziness or inertia is toast, huh?”

    Patrick O’Shaughnessy, reacting to Instinct cancelling unused subscriptions end to end

    “So 40% of the user base 3 weeks in are sharing a credit card.”

    Noah Shinn, on time to first credit card as a proxy for trust

    “We don’t want Instinct to influence the user’s behavior in a way that is not aligned with what the user wants.”

    Noah Shinn, explaining why Instinct will not run on ads

    “Let’s not focus on capability. Let’s only focus on understandability.”

    Noah Shinn, on the early product principle he credits for engagement and word of mouth

    “It’s that the resource has a lead time of several months, right?”

    Noah Shinn, on why buying compute for a product doubling every week is so hard

    “Now we have something that is almost natively proactive.”

    Noah Shinn, on why personal agents will need far more compute than coding tools

    “The user should always be in control of their data.”

    Noah Shinn, on the principle behind Instinct’s security and permission design

    “I think that all of that collapses down in the future.”

    Noah Shinn, on decades of single-purpose apps giving way to one simple interface

    Watch the full conversation between Noah Shinn and Patrick O’Shaughnessy here.

    Related Reading