PJFP.com

Pursuit of Joy, Fulfillment, and Purpose

Tag: mathlete

  • TypeSafe CEO Diogo Almeida on Jev and Building Prod, Not God: Why AI Still Hasn’t Automated the Easy Stuff, Smart Software vs Coding Agents, Reliability Over Benchmarks, and the Inverse SaaS Apocalypse

    Diogo Almeida, co-founder and CEO of TypeSafe AI and a former OpenAI and Google Brain researcher who worked on the RLHF behind InstructGPT and ChatGPT, sat down with a16z’s Martin Casado and a fellow a16z partner for How Jev Builds Prod, Not God. Jev is TypeSafe’s first model, and it is not a chatbot or a coding agent. It is a primitive you put inside your code: you give it natural language and program state, and it returns a typed decision with a confidence level. Almeida’s pitch is blunt. AI is unbelievably smart, and yet almost nothing is automated. This conversation is his case for why that happened and how to fix it.

    TLDW

    Almeida distinguishes Jev from coding agents like Claude Code, Codex, and Cursor, which Garry Tan calls “just in time software.” They write the same code a human would. Jev is a new primitive, a library that takes natural language and state and returns a choice with probabilities, so software itself can do things it never could. He happily calls it a classifier, argues it likely beats a 2019 ML engineering team you can program on the fly, and names intelligence per dollar as his north star. He traces his path from reluctant mathlete to a Kaggle win built on brute-force automation, which led to Isabelle Guyon, Jeremy Howard, Google Brain, retirement, and OpenAI. He explains how RLHF’s surprising generalization in late 2021 made him think AGI was near, and how its failure to deliver turned into a chip on his shoulder: the industry optimized models for the human judge instead of for automation, so GPQA gets solved while a drive-thru still cannot be automated. He rejects the data and long-tail excuse, defines reliability as uptime, determinism, robustness, and being “smart every time,” and says the highest honor is developers programming against Jev without testing example queries. The hosts discuss why SaaS stocks fell on coding agents but cheered Jev. Almeida predicts an inverse SaaS apocalypse, says coding agents are good at syntax and bad at architecture, and the hosts note the average enterprise PR is about 10 lines. The close covers probabilistic programming, rebuilding systems for security, why most future AI calls will be deep in software’s guts rather than facing humans, and his vision of technology that simply does what you mean.

    Thoughts

    The cleanest idea here arrives in the first few minutes and it reframes the whole AI coding debate. Coding agents automate the act of writing software, but the software they produce is the same kind of software we had ten years ago. Almeida wants the opposite: leave software engineering mostly alone and expand what software itself can do. Jev is a function call that takes fuzzy input and returns a typed, confident decision a program can branch on. That is a different category from “AI that codes,” and it explains why a model with no chat interface caught fire among developers. It gives them a new instruction, not a faster typist.

    His diagnosis of why AI has automated so little (roughly 19 to 24 minutes) is the most provocative part, especially coming from someone who helped build RLHF. Once humans became the evaluators, the industry optimized for the judge. Models look brilliant to the person reading their output, so they score well, but nobody was optimizing for whether they could run unattended inside a business process. That is how you get a world where GPQA is solved and a drive-thru is not, and where OpenAI has been trying to automate customer service since 2020. The a16z host pushes back with the long-tail data argument, the familiar story of a help desk that “automates 95%” when most of it is password resets. Almeida’s response is pragmatic rather than dismissive. You do not need the long tail. You need the boring, high-volume core to be automatable, and then automation becomes an ROI decision like any other engineering investment.

    The reliability section (25 to 28 minutes) is where the “prod, not god” slogan earns its keep. Almeida splits reliability into uptime, determinism, and robustness, where robustness means similar intelligence every time. Adding a UUID to a prompt should not change the answer, even if the output is not bit-for-bit identical. Then he adds a fourth layer he does not have a name for, being smart every time in a way a human would find understandable, because a developer can program around that. His bar for success is developers calling Jev without first testing example queries, the way they call a sorting function without checking it on sample data. That is a much higher and more useful target than a leaderboard number, and it is exactly the thing a “benchmaxxed” copycat would miss.

    The SaaS discussion (30 to 35 minutes) is a sharp market read. Coding agents triggered a “SaaS apocalypse” in public markets because the story was that software is now cheap to replicate. Almeida accepts cheap but not easy to replicate, because the value sits beneath the surface in workflows, users, and distribution. If AI becomes something you embed in software rather than something that replaces it, incumbent SaaS companies become the best placed winners. They already know which workflows need automating and have already paid to reach every customer. The host’s data point lands hard: the average enterprise pull request is about 10 lines, so automating code writing optimizes a small slice of the work while adding no new capability, and possibly making software worse and less secure through less oversight. An “inverse apocalypse” is a bet worth taking seriously.

    The closing stretch explains the business logic behind the design. Almeida works backward from a world with AI everywhere and asks what share of all AI calls will be for human consumption, where style matters, versus deep inside programs making decisions. His answer is many nines in the guts, which is why intelligence per dollar matters more to him than eloquence, and why the input to Jev is called “state.” The hosts give the best description of the status quo: AI and software have been ships in the night. Developers stuffed JSON schemas into prompts, watched the model ignore them, and then went through five stages of grief ending in two workarounds, a human in the loop (chat) or another LLM in a while loop (agents). Jev is a bet that you can map a model directly onto a state machine and skip both. If it works, “do what I mean” stops being a sci-fi phrase and becomes a property of ordinary software.

    Key Takeaways

    • Almeida’s favorite elevator pitch for Jev is “where is all the automation?” AI is extraordinarily smart and yet almost useless outside chat and coding.
    • TypeSafe’s mission is making AI work for software, not just for humans in the loop. Jev is its first model.
    • He credits Garry Tan’s description of Claude Code and Codex as “just in time software”: they make software on the fly from natural language, with the same expressive power as ordinary software.
    • What Almeida wants instead is smart software, expanding what software can do so that things that should be automatable become automatable.
    • Coding agents write the same code a human would. Jev is a new primitive you include in your code, whether a human or a coding agent is writing it.
    • Practically, Jev works like a library: describe what you want in natural language, pass state, and it chooses what to do with confidence levels.
    • On day one of onboarding, Almeida draws a Venn diagram of what AI is good at and what is valuable in code. Jev lives in the overlap, which is why it outputs probabilities and not extrapolated floats.
    • He embraces the “it’s just a classifier” critique. Classifiers were designed by practical people to be useful.
    • His guess is that Jev beats having a 2019 ML engineering team build a narrow model for you, and you can program it on the fly.
    • In his heart, the design space is a slider from language-in, language-out to fully imperative programs. Pragmatically, Jev will behave more like a database than a standard library for a while.
    • Intelligence per dollar is his current north star, though he admits intelligence per second may matter more in the short term.
    • The input is deliberately called “state” because Jev is meant to live inside programs.
    • Almeida was an award-winning mathlete who never loved math. Computer science felt like math but cool, useful, and fun, and he considers himself a computer scientist before an AI researcher.
    • He won a Kaggle competition by automating aggressively rather than through sophisticated math, and was then invited to speak at NeurIPS.
    • The competition host, Isabelle Guyon, co-inventor of the support vector machine, took him under her wing and introduced him to the AI community.
    • His career path ran through a startup with Jeremy Howard, Google Brain, a period of retirement, and then OpenAI because AI was simply fun.
    • The hosts see TypeSafe as a movement toward a positive AI future, contrasting “happy AI” Jev users with the “morose AI” crowd.
    • Almeida blames the negative outlook on “mono model Kool-Aid,” the idea of one big brain that rules everything, while basic tasks remain unautomated.
    • He does not buy diffusion as the excuse for slow automation, given the huge financial incentive to automate.
    • An anecdote: a16z’s David George used Meta’s Muse to finally cancel his New York Times subscription, the tip of the iceberg of horrible tasks that need automating.
    • Almeida warns against AI’s anti-pattern of focusing on outliers and demos. He wants use cases that run in the background without paging anyone and that others can build on.
    • Running AI with access to real resources requires guarantees, or at least statistical guarantees.
    • His 2017 talk had a similar theme, roughly “AI modular in theory and flexible in practice.”
    • In late 2021 the RLHF team was surprised by its generalization, verifying with prompts like “why is it important to eat socks before meditating?” that were not on the internet.
    • He thought that model had a decent chance of being AGI. When it was not, his world came crashing down and he began asking why AI was not more useful.
    • RLHF generalizes fairly well in his experience, while RLVR generalizes less well.
    • He does not think we are on a path to recursive self-improvement, but considers OpenAI’s definition of AGI, automating most economically valuable work, extremely doable.
    • Much work is rote and simple enough to outsource with basic instructions, and models have had that level of intelligence for a while.
    • Since RLHF, the industry has overpromised and underdelivered because humans judge the models, so labs optimized the judge instead of automation.
    • His canary in the coal mine: we say math and GPQA are solved, yet we still cannot handle a drive-thru.
    • He does not buy the argument that missing real-world data explains the gap. The long tail is real, but automation does not need to cover it.
    • Echoing the programmer virtue of laziness, automation should be an ROI decision, and people will create new kinds of work once rote work is automatable.
    • OpenAI has been trying to automate customer service since 2020, and outside of programming very little inside companies has been automated.
    • Almeida was extremely surprised by the launch’s reception and says no one could have predicted a ChatGPT moment for developers.
    • Reliability is what Jev is. Every nine of reliability enables new applications, and without understanding that, you cannot build a real copy.
    • He defines reliability in layers: uptime and SLAs, determinism (useful for unit tests), robustness (similar intelligence every time), and being consistently smart in an understandable way.
    • The highest honor would be developers programming against Jev without trying example queries first.
    • Coding agents are good at syntax, weak at semantics, and very bad at architecture, which he sees as the most human, creative part of software.
    • Model-produced architecture might be 50th percentile. That is a legitimate trade-off if speed matters more than quality.
    • SaaS valuations fell when coding agents arrived, but SaaS companies loved Jev. Almeida thinks SaaS will be one of AI’s biggest winners.
    • Software may be cheap but it is not easy to replicate, because the value is beneath the surface. SaaS firms know which workflows to automate and already have distribution.
    • He calls the likely outcome an inverse SaaS apocalypse and imagines multiple choice forms disappearing.
    • His favorite community application is voice control of a computer that constantly decides whether speech is a command or text to insert, and where.
    • An a16z study found the average large-company PR is about 10 lines, so coding agents automate a small slice without adding capability, and may make software worse and less secure.
    • Almeida’s grand hope is to expand software beyond basic logic gates with a new kind of gate that has a little brain in it.
    • His philosophy is to automate the easy work before the hard work, but he expects a new era of probabilistic programming.
    • His brand is pragmatism, and he is not a fan of biologically inspired AI.
    • The hosts argue that a new primitive plus cybersecurity pressure means much of critical infrastructure will be rebuilt.
    • Almeida thinks of AI like TCP and UDP. Most AI calls will eventually be deep in software’s guts rather than facing humans, and you must aim for the guts to get there.
    • AI and software have been ships in the night. Chat is the human in the loop and agents are a while loop feeding language back into another model.
    • TypeSafe could have released much sooner but held back for reliability. Almeida’s utopia includes all technology simply doing what you mean.

    Detailed Summary

    Where is all the automation?

    Asked for an elevator pitch, Almeida offers a question: where is all the automation? He loves chatbots and coding agents, but finds it tragic that such intelligence is so useless for everything else, a diamond in the rough that has not been polished for work. TypeSafe is making AI for software, powerful not just with humans in the loop but inside real software, and Jev is its first model toward that goal. The hosts note that developers have been calling them to rave about it, prompting the obvious question of how it differs from Claude Code and Codex.

    Smart software versus just in time software

    Almeida borrows Garry Tan’s phrase “just in time software” for coding agents, which let you program in natural language while producing ordinary code. He wants smart software instead, expanding the vocabulary of what programs can express, including something like intent. He loves that programming means hyper-specifying valuable things and replicating them infinitely, and he wants more of that. The hosts sharpen the distinction: whether Claude Code or a human writes it, Jev is something you include in your code. It is a library where you describe what you want in natural language, supply a state machine, and get back a choice with confidence levels, something software has rarely had so widely. It asks programmers to think in probabilities.

    Yes, it is a classifier

    Almeida finds it wild that AI is this capable while software has been unchanged for a decade, with the best effort being a sidebar chatbot that can take some actions but not all, because some actions are not reliable. He draws a Venn diagram for new hires of what AI is good at and what is valuable in code, and Jev sits in the middle. Probabilities are in the overlap, extrapolated floats are not. To the “Jev is just a classifier” critique he says absolutely, classifiers are great. They share interfaces with classic ML concepts invented by practical people. He guesses Jev beats a 2019 ML engineering team, which few companies ever had, because you can program it on the fly without collecting and measuring datasets.

    Design choices: a slider, state, and intelligence per dollar

    One host asks whether there is a slider from language-in, language-out to imperative programs, or whether language-in, state-machine-out is the design point that will solidify. Almeida says that in his heart it is a slider. His north star for now is intelligence per dollar, though intelligence per second might be more valuable in the short term. Calling the input “state” is intentional because Jev is meant to live inside programs, and much of his work targets ever more complex arrangements of program internals. Pragmatically, it is easier to hit certain latencies in a database-like service, so Jev will look more like a database for a while, though he would love it to become a standard library feature too.

    From mathlete to OpenAI

    Almeida was an award-winning mathlete who never liked math, a big fish in a small pond who resented competition. Computer science felt like math but useful and fun, and he still loves giving algorithms interviews because they reveal a lot about candidates. He won a Kaggle competition by automating heavily, with more nested loops and a systems approach rather than sophisticated math, and was pushed to speak at NeurIPS. The host, Isabelle Guyon, co-inventor of the SVM, saw someone who did not fit the research mold and introduced him to the AI world. From there he joined a startup with Jeremy Howard, then Google Brain, then retired for a while, and finally joined OpenAI because AI was fun.

    Prod, not god, and the happy AI camp

    The hosts praise TypeSafe’s slogan “we build prod, not god” and its optimism about more and better jobs, framing it as a movement. Almeida says critics raising the classifier point are voicing an ML-level concern while developers are partying, because they can finally do what they wanted. He blames the negative worldview on mono model thinking, one big brain to rule them all, even though basic, unwanted work remains unautomated. He rejects diffusion as an excuse. One host recounts David George canceling his New York Times subscription with Meta’s Muse. Almeida says honest pursuit of automation means avoiding AI’s fixation on demos and outliers in favor of workflows that run in the background, do not page anyone, can be composed, and come with at least statistical guarantees when they have access to resources.

    RLHF, AGI, and optimizing the judge

    A host recalls talking to Almeida in 2017, when his talk was roughly “AI modular in theory and flexible in practice.” Almeida says the real turn came just before ChatGPT, in late 2021, when the RLHF team was surprised by how well it generalized. Their paper tried to disprove its own claims, including testing prompts like “why is it important to eat socks before meditating?” that did not exist online. He pushed hard to release that model and thought it had a decent chance of being AGI. When it was not, his world came crashing down. He says RLHF generalizes well and RLVR less so. He recalls early OpenAI describing AGI as “Ilya and every if statement,” a deliberately vague big tent. He does not think we are on a path to recursive self-improvement, but thinks automating most economically valuable work is very doable. Much work is rote, and the needed intelligence has existed for a while. Since RLHF, the industry has optimized the human judge rather than automation.

    The long tail argument and new kinds of work

    Almeida’s canary: if math and GPQA are solved, why can we not handle a drive-thru? A host proposes that real-world distributions are heavy-tailed and underrepresented in training data. Almeida does not buy the data argument. The long tail exists, but you do not need to automate it. Building reliable software is always an investment, and he invokes the programmer virtue of laziness, spending ten hours to never do a five-minute task again. Automation should be an ROI decision, and he believes people will invent new kinds of work once rote work can be automated, an argument close to the Jevons paradox. The hosts add that OpenAI has pursued customer service automation since 2020, and that pre-generative support vendors claiming 95% automation were mostly handling password resets, closer to 50% by uniqueness.

    Reliability as the product

    Almeida was extremely surprised by the launch, which even non-developer friends joined in memeing. But he stresses years of work on reliability, which he says is what Jev is. Without understanding that, you cannot build a copy that is not just benchmaxxed. Every nine of reliability unlocks new applications, even ones the team does not yet understand. Asked what reliability means for a stochastic system, he lists uptime and SLAs, determinism, and robustness, meaning similar intelligence every time, so adding a UUID to a prompt should not change the result. A further, unnamed layer is being smart every time in ways a human would find understandable, which developers can program around. The goal is developers trusting Jev enough to skip example queries and work in a flow state.

    Coding agents, syntax, and architecture

    A host suggests that a primitive like Jev could lower the value of coding agents, since agent-written software that does not use it stays limited. Almeida calls this more of a coding agent question. In his experience, agents are very good at syntax, weak at semantics, and very bad at architecture, which he sees as the most creative human part of software. Jev is almost certainly not in their training distribution yet. When it is, he is happy for agents to handle syntax. Their architecture might be 50th percentile, which is fine if you know nothing about architecture or if speed is the knob your project wants to turn, for example letting Codex work overnight.

    The inverse SaaS apocalypse

    The hosts note that SaaS stocks plunged on coding agents, yet SaaS companies welcomed Jev. Almeida says the apocalypse story, software being cheap and easy to replicate, has played out poorly. It may be cheap, but it is not easy to replicate, because the value sits beneath the hood. He wants to work with the biggest, most boring SaaS companies that know user problems best, since they know which workflows to automate and have already made the capex investment to reach users. He calls it an inverse apocalypse. The hosts add that SaaS capital largely goes into reaching customers, so making the software genuinely better, not just adding a chatbot, is powerful. Almeida imagines multiple choice forms disappearing, and the hosts compare it to 1980s fourth-generation languages. He calls it “do what I mean” taken to the next level and highlights a community project that uses voice to control a computer, constantly deciding whether speech is a command or text to insert.

    Better software, not just faster software

    A host shares an a16z finding that the average PR at a large company is about 10 lines, often capturing something learned from a customer. Coding agents optimize that minimal slice without adding capability, and software may be getting worse and less secure due to reduced oversight. Jev, by contrast, speaks natural language and reasons while being married to a state machine, so apps can gain new functionality. Almeida says that takeaway would be the greatest compliment. His grand vision is expanding software beyond its basic logic gates with a gate that has a little brain in it, and he promises to fight for it without overpromising.

    Probabilistic programming and rebuilding systems

    One host questions how deep this can go into systems needing strong guarantees like state consistency and durability, versus log analysis, email, and UI. Almeida’s philosophy is to automate the easy work first, but he expects a new era of probabilistic programming, which the hosts note largely died decades ago. They mention co-founder Erik’s Bayesian background, and Almeida says his brand is pragmatism and he is not a fan of biologically inspired AI. The hosts agree such ideas mostly motivate people for decades until engineering refines them. Almeida expects systems engineers to use cheap, fast intelligence for approximate guesses and optimistic routing. The hosts add that new primitives and cybersecurity pressure mean much critical infrastructure will need rebuilding, as happened with the internet and client-server.

    Aiming for the guts, and do what I mean

    Almeida describes his thinking in TCP and UDP terms, which one host calls speaking his language. He worked backward from a world where AI is everywhere, asking what share of AI calls are for humans versus buried in software. His answer is many nines deep in the guts, starting at the surface. If you do not aim for the guts, you will not get there. The hosts describe AI and software as ships in the night: developers put JSON schemas in prompts, the model ignored them, and after five stages of grief they handed output to a human (chat) or another LLM in a loop (agents). This is the first time they have seen AI productively mapped onto a state machine. Almeida says TypeSafe could have released much sooner but held back for reliability, and hopes users simply feel they can trust it. His AI utopia includes technology that does what you mean, which he says is not sci-fi given how smart AI already is.

    Notable Quotes

    “AI is so unbelievably smart and yet it’s so useless at all other stuff.”

    Diogo Almeida, in the opening pitch for why TypeSafe exists

    “What I want instead is smart software.”

    Diogo Almeida, contrasting Jev with coding agents that produce just in time software

    “Jev is absolutely a classifier. You know, like classifiers are sick.”

    Diogo Almeida, embracing the most common critique of the model

    “We’ve been optimizing that judge instead of the automation part and that has been the missing thing.”

    Diogo Almeida, on how RLHF-era evaluation led the industry to overpromise

    “OpenAI has been trying to automate customer service since 2020.”

    Diogo Almeida, on the gap between benchmark progress and real automation

    “Reliability is what this thing is.”

    Diogo Almeida, on years of work that a benchmark-chasing copycat would miss

    “It doesn’t matter how much AI coding agents you use, the software actually isn’t getting better.”

    An a16z host, on why a new primitive matters more than faster code writing

    “I think automate the easy work before the hard work is always my philosophy.”

    Diogo Almeida, responding to questions about using Jev in systems that need strong guarantees

    “Imagine if all technology just did what you mean.”

    Diogo Almeida, closing on his vision of an AI utopia

    Watch the full conversation with Diogo Almeida here.

    Related Reading