PJFP.com

Pursuit of Joy, Fulfillment, and Purpose

,

Poteto on pstack: Stop Being the Meat Proxy for Your Coding Agents

Poteto, the creator of pstack, explains how to stop being the “meat proxy” between your coding agents and their tools, and how that led to 2,500 merged pull requests in one month. In this live conversation with Matt Pocock, Lauren Tan (Poteto online), a former React team engineer at Meta now working at Cursor, which Tan describes as now part of SpaceX AI, walks through the trust ladder, the “Michelin kitchen” alternative to the software factory, and what it takes to let agents merge their own code overnight.

TLDW

Tan’s argument is that the limit on agentic coding is no longer the model. It is the human standing between the agent and everything it needs: the browser, the profiler, the bug reports in Slack. Remove yourself from each of those spots, starting with verification, and trust builds until agents can run in parallel and merge without you. The work that makes this possible is mostly unglamorous: lint rules, a codebase with one way to do each thing, scripts for the deterministic parts, and routines that pull outside context in. Tan is clear that it takes a long time to get there and that it only works where the output can be verified.

Thoughts

Ask anyone who ships 2,500 pull requests in a month how they reviewed them, and the honest answer is that they did not, at least not before the code landed. Tan says so plainly. Agents merge their own work, and the review happens the next morning by reading the commit history. What makes that less reckless than it sounds is the unit of correction. When Tan finds a bad pattern, the fix does not go to the pull request or to the agent that wrote it. It goes to the environment: a new lint rule, a tighter type, an amended skill, so that no agent can make that mistake again. Pocock’s phrase for it is “sampling instead of blocking,” and it is the real change in mindset here. A reviewer who blocks every change is quality control. A reviewer who samples and then changes the process is running a system. Most teams still do the first and wonder why agents do not make them faster.

The “meat proxy” line is funny, and it is also the most portable idea in the conversation. Tan coins it while describing performance work on Cursor’s agents window: copying numbers by hand between an agent and Chrome DevTools. But the bigger version comes later. Even with a fast inner loop, Tan was still the courier for everything outside it, carrying bug reports from Slack, Linear and X to the agents and keeping their picture of the goal from going stale. The question Tan keeps asking is “where am I the bottleneck in this process?” and it is a good one to borrow, because the answer is usually boring. You are pasting an error message. You are describing a screenshot. You are relaying what a customer said. Each of those is a place where a tool, a subscription or a skill could stand instead of you.

The part I expect people to skip is the part that does the work. Tan built an internal framework called Dune, described as a kind of in-house Next.js for Electron apps, where every feature lives in its own directory and strict lint rules leave only one way to do each thing. It grew out of watching an early version of the product pile up in eight files of 10,000 lines or more. This is where the claim that domain expertise matters more than ever becomes concrete. The conventions in that framework are a senior engineer’s taste, written down as constraints an agent runs into at the right moment. Someone without that taste can install the same skills and will not get the same result. The skills are downloadable. The judgment about what the codebase should forbid is not.

A smaller idea deserves more attention than it gets: split every skill into the part that needs judgment and the part that does not, and turn the second part into a script. Tan noticed agents rebuilding the same verification tooling from scratch in every session, each one differently, and then throwing it away. Wrapping the mechanical steps in a small CLI made every agent faster and more consistent, and left the model to do only the thinking. The same logic covers migrations, which Tan does with code mods and not with an agent improvising each file. It also explains the prediction near the end that skills will get shorter. As models improve, the instructions shrink to a workflow, and the fixed steps move into code where they belong.

The limits are stated honestly, and they matter more than the headline number. Asked about changes that cannot be undone, in medicine, law or finance, Tan says “I don’t really have the answer” and falls back on a single test: can the agent verify the work in a way you trust? Where it can, a one-way door starts to look like a two-way door. Where it cannot, none of this applies. The 2,500 figure needs the same care. Tan says much of it is “gardening,” small fixes to keep the codebase clean, and that full autopilot with ten verifier agents per pull request burns a lot of tokens. Counting pull requests tells you how finely the work was sliced, not how much value shipped. The transferable part is the sequence: verification first, then constraints, then outside context, and only then autonomy.

Key Takeaways

  • The “meat proxy” is the human relaying information between an agent and a tool it cannot reach. Tan’s whole method is finding those spots and removing yourself from them.
  • Verification is the one skill Tan says everyone needs, with or without pstack: give the agent “hands and eyes” so it can run the app, click through it like a user, and take traces and snapshots.
  • A loop is only a loop if the agent can check its own work. Once it can, you can hill climb: give it a score to improve and let it keep trying.
  • Even frontier models take shortcuts and do the easy thing. Tan’s skills aim to make the easy thing the right thing.
  • Domain expertise is worth more, not less. As models improve, the bottleneck becomes your ability to express intent clearly, which favors doctors, lawyers and other experts who are a little technical.
  • Precise words carry compressed intent. Tan borrowed Pocock’s instruction to remove “tautological tests” because one word did the work of a paragraph.
  • Put deterministic steps in scripts and CLIs, and leave only judgment to the agent. Otherwise each agent reinvents the tooling, slowly and differently.
  • Every time an agent makes a mistake, ask how to turn it into a lint rule or make it impossible in the codebase. Constraints in the environment do not have to be remembered.
  • Tan separates an inner loop (agents building toward a snapshot of your intent) from an outer loop (bug reports, feature requests and other context arriving from outside). Connecting them is what lets work start without you.
  • Coordinator agents, which Tan compares to an executive chef or chief of staff, do no work themselves. They hold the task list and delegate to sub-agents. Tan runs more than ten.
  • Sometimes a queue beats instant action. One of Tan’s routines logs bad React patterns to a document, and reading it every few days reveals that many entries are the same problem.
  • Review by sampling, after merge. Read the commit history in the morning, revert what is wrong, and fix the environment when several agents repeat a mistake.
  • Mine your own chat transcripts. The moments where you corrected an agent are your real process, and each one can become a skill or a lint rule.

Chapters

01:31 Climbing the Trust Ladder and the Meat Proxy Problem

After leaving the React team at Meta, Tan took a month off and started a side project, then noticed hours going into micromanaging a single agent. The skills written to fix that became the basis of pstack. On joining Cursor in March, Tan was asked to fix a laggy agents window and spent the first weeks in flame graphs and heap snapshots by hand. That is where the meat proxy line comes from: being the link between the agent and Chrome DevTools.

06:49 Why Domain Expertise and Language Matter More Now

Pocock asks whether expertise is losing value. Tan argues the opposite: the constraint is now transferring your intent and vision to the agent, and experts have the clearest vision. The two compare notes on wording. Pocock looks for terms an agent latches onto and repeats in its reasoning, such as TDD, and Tan points to “tautological” as a word that packs a lot of meaning into a small space.

10:36 The Michelin Kitchen Instead of the Software Factory

Tan finds “software factory” accurate but short on craft, and prefers a Michelin kitchen. A home cook does everything alone and gets stressed when the family crowds in. A chef stops cooking every dish and runs the kitchen: ordering, storage, prep and timing. Engineers are in the same position, no longer writing the code but still responsible for what goes out under their name.

15:58 Verification, Loops and the CLI Inside the Skill

Verification was the first skill Tan built at Cursor and the first that moved trust upward, because earlier skills still left a human checking the output. Every app there now has its own verification skill, maintained automatically. The CLI inside it began as a way to save context window and stayed as a way to keep mechanical steps out of the model’s hands. Tan calls it glue around Playwright and the Chrome DevTools protocol, nothing novel.

25:08 The Environment Is the Job: Dune, Lint Rules and God Files

Tan says the engineer’s new job is the environment. Without trust you micromanage, and micromanaging leaves no time to sharpen your knives. Drawing on TypeScript’s type narrowing, Tan describes constraining the space of possible code until there is one right answer. The internal Dune framework does this with a directory per feature, a registry that discovers them, and restrictive lint rules.

32:47 Where 2,500 Pull Requests Come From: Inner and Outer Loops

Tan does not open 2,500 chats. With the kitchen built, each project runs like a restaurant in a chain, and Tan moves between them. The missing piece was outside context, which Tan had been carrying in by hand. Subscribing agents to a Slack channel lets them triage a bug report, reproduce it and confirm it still exists on main. Tan finds terms like “company brain” too abstract: it is just teaching the agent to fetch what you would otherwise fetch for it.

40:25 Coordinator Agents, Gardening and Routines

Tan’s setup has two layers. Personal agents with connectors to Slack, email, X and Linear watch for issues and forward them. Cursor projects receive them: a coordinator agent in the cloud that picks a team shape and delegates. Grouping related bug reports under one coordinator matters because several slightly different reports often point to one cause. Tan borrows “context, not control” from time at Netflix, and notes that many of the pull requests are routine upkeep.

48:46 Review by Sampling and the Dark Factory

You cannot taste every dish across several restaurants, so you sample and correct the process. Tan admits the factory is dark in one sense: agents merge their own pull requests around the clock. A full autopilot mode in pstack spawns verifier agents that fuzz the running app, find regressions and fix them before landing. The first night was frightening. Now, Tan says, sleep is better.

55:36 One-Way Doors, Verifiability and Proofs

Pocock asks about work where a merge cannot be walked back. Tan’s answer rests on how verifiable the domain is, and concedes the industry has not solved the hard cases. The hope is for agent-oriented programming languages that combine code with formal proofs, with a language called Bend as the example, so that code which compiles and proves correct could simply merge.

1:00:11 Skills Are Just Process: Recall and Your Own Knives

Asked how pstack and Pocock’s skills fit together, Tan says they are complementary and that everyone should build their own set, the way a chef carries personal knives from job to job. Past chats are a treasure trove because they are the process as it really happened. Tan’s recall skill came from wanting to carry context from one chat into the next without writing an essay each time. Skills, Tan expects, will keep getting smaller.

Notable Quotes

“I was sort of the meat proxy in a way, right? I was the meat proxy between my agent and Chrome DevTools.”

Poteto, on doing performance work by hand at Cursor

“A lot of the skills that I’ve built have been around how do I make the easy thing the right thing?”

Poteto, on why even frontier models need guard rails

“If the agent can’t actually see the result of its work there’s no way it can actually iterate.”

Poteto, on why verification comes before everything else

“I almost feel like the new job of the engineer is really to spend time on the environment.”

Poteto, on what replaces writing code

“How do I turn this into a lint rule? How do I make it so that the code base makes this impossible?”

Poteto, on what to ask each time an agent makes a mistake

“You don’t want to be in a position where you’re not tasting your food ever again.”

Poteto, on reviewing by sampling when agents merge their own pull requests

“Trust to me is really about trust in your own tools.”

Poteto, on why everyone should build their own skills

The detail on how the loops are wired together is worth hearing first hand, along with Pocock’s questions about where it breaks. Watch the full conversation here.

Related Reading

  • Poteto on GitHub where the skills and the earlier open source project mentioned in the conversation are published.
  • Cursor the editor whose agents window and projects feature are discussed throughout.
  • Total TypeScript Matt Pocock’s site, for the host’s own work on TypeScript and agent skills.
  • Formal verification (Wikipedia) background on proving code correct, the idea behind the one-way door answer.
  • Brigade de cuisine (Wikipedia) how a professional kitchen divides work, the model behind the Michelin kitchen metaphor.