PJFP.com

Pursuit of Joy, Fulfillment, and Purpose

,

Software Factory Tutorial: Build an AI Agent Pipeline (9 Best Videos)

A software factory is a repeatable system in which AI coding agents take a request, build the change, prove it works, get it reviewed and ship it, with people supervising instead of typing. This post distills nine videos on the subject, published between September 10 and October 5, 2026, into one step-by-step tutorial, followed by notes on each video. The one to watch if you only watch one is Steve from Builder.io’s deep dive, embedded above, because it shows a factory that has been running on real products for months.

TLDW

Every builder in this set describes the same assembly line: isolate the task, build to written rules, make the agent prove the result, review in a loop, then decide by risk whether a person has to approve the merge. The factory is the workflow and its checks, not the model and not a product you buy. The videos disagree about where it should run, how much should merge without a human, and whether a fully “dark” factory with nobody reading the code is wise at all. None of the nine is actually dark, and the most experienced builders keep a person on production actions.

The 9 Videos at a Glance

#VideoChannelLengthBest for
1Build an agentic software factory: deep diveSteve (Builder.io)17:35The full loop from feedback to shipped fix, with open-source skills
2Building a Software Factory that actually works (Full Course)Greg Isenberg31:29The clearest build of the inner loop: isolate, build, prove, ship
3What is An AI Software Factory? The Future of Coding with AI AgentsAI With Sanket18:57Non-developers who need to judge whether a factory demo is real
4The AI Software Factory: Agents Run the Entire SDLC5 Minutes Tech5:12A five minute map of the enterprise version
5Inside OpenAI’s Agentic Software FactoryOwain Lewis10:48Risk classification and auto-merge, drawn from how OpenAI works
6I Built A Claude DevOps Agent To Run My Software FactoryOwain Lewis16:55Building the on-call agent that watches production
7Forget the Terminal. This Is the Future of AI CodingLeon van Zyl17:04Seeing a factory as a 3D office with a CEO agent. Fun, lighter on method
8AI Engineer Paris 2026 Opening Keynotes: Mistral, Langfuse & SizzyAI Engineer1:13:35Sandboxes per task, agent safety and the economics. Long, three talks
9How To Build a Dark Software Factory in 3 MonthsCraft vs Cruft9:06The skeptic’s case. Watch before you bet a team on going dark

Thoughts

The most useful agreement in the set is about what a software factory is not. Ross Mike, in Greg Isenberg’s video, says it is a handful of markdown files and a workflow. Steve’s factory is a set of skills and one YAML file that the agent is allowed to edit. Sanket says an older model in a well-built environment can beat a newer one in a poor one. Yet three of the nine videos end with someone showing a product: a 3D office, a custom orchestrator, an enterprise platform. Both things are true. The method is free and portable, and the interface on top is where people are now competing. Learn the method first, because it survives a change of tool.

Strip away the vocabulary and every video is about verification. Steve repeats the word three times in a row. Ross Mike makes the agent record the bug before and after. The 5 Minutes Tech pipeline has a review loop on every artifact. The sharpest warning comes from Sanket: a planner, a coder and a reviewer that share one model and one misunderstanding will all agree, and the customer still gets double booked. That is a real weakness in the simpler setups, where the same agent writes the code and the proof. The builders who address it do so the same way, with a check that comes from somewhere else: a separate review service, a second model family as watchdog, or tests written before the code.

Where the factory runs is the clearest disagreement, and it hides a security problem nobody connects. Steve insists on a real laptop with his keys, browser sessions and passkeys, because cloud environments could not reproduce his setup. Kitze, at AI Engineer Paris, went the other way after an agent in a git worktree wiped his production database, and now gives every task a throwaway virtual machine. Then Mistral’s talk demonstrates an agent reading a GitHub issue that contains a prompt injection and leaking a secret key. Put those together. A factory that collects public issues automatically and runs on a machine holding real credentials is exposed to exactly that attack. If your inputs come from strangers, the sandbox is not optional.

What every builder skipped is the ledger. Steve says dozens of merged pull requests a day fit inside a $200 plan. Kitze pays for four such plans, runs two load balancers, built two Linux machines, and says more than once that he has not shipped much. Ray Myers quotes a claim that teams should burn $1,000 in tokens per engineer per day. Nobody shows a before-and-after of customer outcomes. Clemens of Langfuse supplies the frame they are missing: electricity took decades to raise factory output because the factories had to be rebuilt around it, and investment in a new technology lowers productivity before it raises it. Many of these setups are in that dip right now. Count the cleanup, as Sanket says, not only the moment the code appears.

On the dark factory itself, the evidence in these videos points one way. Owain Lewis, who has built the on-call agent twice, keeps a human approval on every rollback because agents jump to conclusions. OpenAI, as he describes it, merges low-risk changes automatically and sends anything touching infrastructure or a database to a person. Steve cuts production releases once a day so a bad change is easy to undo. Ray Myers’ joke lands because it is accurate: if the models could do all of software engineering, nobody would need a tutorial on building the factory. The workable version is lights dimmed by risk class, with the lights fully on wherever a mistake is expensive.

The Tutorial: Software Factory Step by Step

The steps are in the order a person does the work, not the order of the videos. Model names, prices and plan limits are as stated in each video on its date.

Step 1: Pick one boring task and define done

Do not start with a whole app. Start with one kind of change whose result you can check and undo, such as fixing a clear error message. Then write down what must stay true when real people use it. Sanket’s example is a court booking page. “Build a booking page” is a bucket of unanswered questions. “Two confirmed bookings must never overlap, even if two customers try at the same moment” is a rule you can test. Write three or four examples like that before any agent writes code. In the enterprise pipeline this is the requirements phase, where ambiguity is flagged before implementation and not during it.

Shown in: Sanket 03:01, 5 Minutes Tech 00:00, Steve 14:21.

Step 2: Write the workflow down as skills

The factory lives in a few plain files. Ross Mike uses an AGENTS.md file, which is sent to the agent with every message, plus five or six skills. His rule for AGENTS.md is to leave out anything the agent can read from the code and put in only the workflow it would not follow by itself. The larger version is a knowledge base: architecture standards, security policy, coding conventions and past decisions, loaded per task, with agents allowed to propose additions that a person approves. Kitze warns that skills scattered across machines as loose files become a mess, and serves his from one place over MCP so every agent gets the same set.

Shown in: Ross Mike 03:49, 5 Minutes Tech 01:30, Kitze 35:20.

Step 3: Give every task its own workspace

Agents working on the same branch overwrite each other. The first fix is a git worktree per feature, branched from the main line and deleted after the merge:

git fetch origin
git worktree add -b feature/landing-page ../myapp-landing-page origin/main
# the agent works in ../myapp-landing-page and opens a pull request
git worktree remove ../myapp-landing-page

Ross Mike runs fifteen features at once this way. A worktree separates files and nothing else. It shares your database, your credentials and your machine, which is how Kitze lost a production database to an agent that was “testing”. For anything that touches data, give each task a throwaway container or small virtual machine with its own database and seed data and no production secrets, and put a limit on how many run at once. Kitze queues them at eight. The enterprise pipeline does the same with one container per agent per task and short-lived credentials.

Shown in: Ross Mike 05:21, Kitze 39:07, 5 Minutes Tech 03:00.

Step 4: Give the builder rules it cannot skip

Models get the job done, and will do it sloppily if sloppy works. Ross Mike has a code structure skill that makes the agent write in a service layer style a new developer, or a new agent, can follow. Kitze goes further because agents quietly disabled his lint rules to reach “no errors”. He stacks strict linters (he names Oxlint, Ultracite and a library called anti-slop), a small command line checker of his own for rules a linter cannot express, and a cheap model that reads each pull request against rules written in plain English. The principle is the same in both: the rules in your head have to be moved into something that fails the build.

Shown in: Ross Mike 11:25, Kitze 41:49.

Step 5: Make the agent prove its work

An agent saying it worked is not evidence. Ross Mike’s agents record the bug happening, fix it, then record the working result, and put both in the pull request. When the change has no visible surface, they attach measurements. His example is a page that went from about 850 milliseconds to about 60. If the after recording shows the feature is not finished, the skill sends the agent back to build without being told. Steve’s factory only takes a task at all if an agent can reproduce the issue, fix it and verify the fix. Leon van Zyl uses a separate QA agent that attaches screenshots to its report. For rules that matter, follow Sanket and anchor at least one check outside the agent that wrote the code.

Shown in: Ross Mike 14:26, Steve 03:01, Leon van Zyl 08:21, Sanket 12:14.

Step 6: Review in a loop, with a different reviewer

Put a second set of eyes on every pull request and make the building agent answer it. Ross Mike uses the code review service Greptile, which leaves comments and a confidence score out of five. His agent goes back to build, proves again and resubmits until the score is five, and only then does he look. He names CodeRabbit and Macroscope as alternatives. Steve’s agents must address or rebut every review comment and every CI failure, and a separate scheduled agent checks that they did. OpenAI, per the article Owain Lewis walks through, splits review among specialists for data, infrastructure, cloud and security. Give every loop a retry limit and escalate to a person when it is hit. A loop with no exit only spends money.

Shown in: Ross Mike 22:06, Steve 04:33, Owain Lewis 01:30, 5 Minutes Tech 02:15.

Step 7: Classify risk and decide who merges

This is the step that turns a fast agent into a factory. Have an agent label each pull request by risk. Low risk with green checks can merge without a person: documentation, a CSS fix, a dependency patch with full test coverage. Anything touching infrastructure, a database or a payment flow waits for a human. Owain Lewis adds two rules from his own setup: never auto-merge when CI is failing, and never auto-merge a branch that needs a rebase, because the code may no longer be what was reviewed. Risk is relative. A solo developer on a side project and a cloud provider should draw the line in very different places. Leon van Zyl’s app makes the same choice a single checkbox.

Shown in: Owain Lewis 04:32, Steve 05:18, 5 Minutes Tech 02:15, Leon van Zyl 09:06.

Step 8: Feed the factory from real signals

Once the inner loop is trusted, connect the inputs. Steve’s collect skill gathers error reports, logs, performance and uptime numbers, in-app feedback and GitHub issues every few hours, and runs each through a policy. Clear, verifiable issues go to agents. Everything else is batched for people. If a report is too vague, the agent asks the reporter for the page address or the time it happened, and checks back later for the answer. His strongest claim is that a factory is only as good as its inputs, so time spent on telemetry, source maps and feedback channels pays back directly. Given the prompt injection demo in the Mistral talk, treat input from the public as untrusted and run that work in a sandbox.

Shown in: Steve 02:15, Steve 12:51, Mistral 61:07.

Step 9: Release in stages and put an agent on call

Agents will produce many small merges, so slow the last step down. Steve keeps a beta environment in step with the main branch, has his own team use it, and cuts a production release once a day so a bad change is easy to roll back. Owain Lewis builds the other half: an agent that receives production alerts from a queue, reads logs, metrics and recent deployments, and reports a diagnosis with a recommended action. It can also be told to watch the next deployment for five minutes and flag regressions. He keeps a human approval on the rollback itself, and has the agent write what it learned from each incident to a lessons file so the next diagnosis is better. Mistral shows the same pattern, from a Slack alert to an agent investigation waiting for the engineer.

Shown in: Steve 13:36, Owain Lewis 05:20, Owain Lewis 11:23, Mistral 53:31.

Step 10: Add watchdogs, then schedule it

Automate last. Steve ran every skill by hand for months, starting with a dry run where the agent only listed what it thought it should fix, and scheduled each one only when he noticed he was pressing the same buttons every morning. Two additions matter at scale. A watchdog, run a few times a day, finds work that stalled, such as a pull request that was opened and never shipped, and nudges it along. For harder work he points a model from a different lab at the thread to check it independently. A weekly lookback hunts for the same complaint returning after a supposed fix and hands the pattern to a larger model to solve at the root. A dedicated always-on machine with plenty of memory comes only after all of that works.

Shown in: Steve 08:19, Steve 10:35, Steve 15:08.

Cheat Sheet

StationWhat it doesWhat to copy
Define doneTurns a request into rules and examples that can be testedThree or four examples written before any code
SkillsHolds the workflow in plain filesAGENTS.md with only what the agent cannot infer
IsolateOne workspace per taskWorktree for files, container or VM for anything touching data
BuildWrites to enforced rulesCode structure skill plus linters the agent cannot switch off
ProveEvidence before and afterRecording, screenshots or numbers in the pull request
ReviewA different reviewer, in a loopReview service or second model, with a retry limit
GateRisk label decides who mergesLow risk and green merges. Data, infra, payments wait for a person
CollectPulls in errors, metrics and feedbackA policy: verifiable goes to agents, the rest to people
ReleaseStaged rolloutBeta follows main, production cut once a day
On callAgent investigates alerts and watches rolloutsHuman approval on rollback, lessons written to a file
WatchdogFinds stalled work and repeat failuresDaily nudge, weekly lookback, run by hand first

Five questions to ask of any software factory, from Sanket’s video: what work does it automate, what counts as done, what checks the evidence, what happens when it fails, and can it handle the next change without breaking the last one.

The Videos, One by One

1. Build an agentic software factory: deep dive

Steve (Builder.io), 17:35, published September 23, 2026. Watch on YouTube.

This is the video embedded at the top of the post.

Steve runs most of the maintenance of his open-source Agent Native framework and the apps built on it through scheduled agent loops. The video walks the whole cycle (collect, babysit, review, ship, repeat) and then spends its second half on what goes wrong and the automations that catch it. The skills are open source.

Key takeaways

  • The default policy is narrow: an agent may take an issue only if it can reproduce it, fix it and verify the fix. Everything else is set aside for people.
  • He runs the factory on a second laptop that is always on, because verification needs his real environment. When it is blocked it messages him on Telegram.
  • As of this video he prefers Codex with the Luna model on maximum settings for volume work, and says dozens of merged pull requests a day fit inside a $200 ChatGPT Pro plan.
  • A watchdog runs about four times a day to find work that stopped partway. A weekly lookback compares the last 30 days with the 30 before to catch problems that keep returning.
  • His tips: keep everything in one monorepo, invest in the quality of your inputs, separate beta from production, and start with one skill run by hand.

Chapters

  • 00:00 The basic loop and where people can take over
  • 02:15 Input sources and the collect skill
  • 04:33 Babysitting a pull request and the approval policy
  • 06:03 Why it runs on a laptop and not in the cloud
  • 08:19 Watchdog and the weekly lookback
  • 12:51 Monorepo, input quality, beta and production, starting small

“I spend most of my time feeling like a scientist of a biological organism, always inspecting, microscoping, and assessing if the factory’s working as efficiently and as effectively as it possibly can.”

Steve of Builder.io, on what the job becomes

2. Building a Software Factory that actually works (Full Course)

Greg Isenberg, 31:29, published September 14, 2026. Watch on YouTube.

Greg Isenberg hosts Ross Mike, who draws his own factory on a whiteboard: isolate, build, prove, ship. It is the most teachable video in the set and assumes no engineering background. Greg’s comparison to a physical factory near the end (a station per order, an assembly line, quality control, shipping) is a good summary to borrow.

Key takeaways

  • A software factory is model and harness agnostic. It is a workflow, skills and some domain knowledge.
  • Most reports of an agent deleting work come from several agents sharing one branch. A worktree per feature fixes that.
  • Working code is not well-written code. He has one model review another’s output and gives the builder a structure to follow.
  • Before-and-after proof makes review possible for someone who does not read code. Greg compares it to tapping through stories.
  • He does not read most of the code any more. He enters only when the review score reaches five out of five, and then clicks merge.

Chapters

  • 02:17 What a software factory is and why it matters
  • 03:49 The AGENTS.md file and what belongs in it
  • 05:21 Isolate: one worktree per feature
  • 11:25 Build: the code structure skill
  • 14:26 Prove: evidence-driven testing and before and after
  • 22:06 Ship: the review loop and the confidence score
  • 29:05 Why a factory is not a product

“A software factory is literally just a bunch of markdown files.”

Ross Mike, on why you do not need to buy one

3. What is An AI Software Factory? The Future of Coding with AI Agents

AI With Sanket, 18:57, published September 27, 2026. Watch on YouTube.

Sanket explains the idea to people who do not write code, using one small booking app from request to production. It builds nothing, and it is the best video here on how to tell a working factory from a convincing demo. Skip it if you already run agents with tests and reviews. Watch it if someone is about to sell you one.

Key takeaways

  • His definition: a repeatable system that turns a software request into a checked change, with explicit rules for when to continue and when to stop.
  • The model is one worker. The harness around it (tools, instructions, records, safety rules) is the workshop, and the same model behaves differently in a better one.
  • More agents are not more independent experts. They can share a model, missing information and the same mistake.
  • “Production ready” means nothing until you ask under which criteria, with which checks, and who accepts the remaining risk.
  • The real test of a factory is the second change to a working app, not the first version.

Chapters

  • 02:16 From autocomplete to coding agent to factory
  • 03:01 A practical definition, and turning a request into rules
  • 05:20 Model, tools and harness
  • 06:54 Why an army of agents is not a team of experts
  • 11:28 What production ready means
  • 13:01 The better test: ask for a change
  • 16:48 Five questions to ask

“Agreement is not the same thing as evidence.”

Sanket, on agents that review each other’s work

4. The AI Software Factory: Agents Run the Entire SDLC

5 Minutes Tech, 5:12, published October 3, 2026. Watch on YouTube.

A narrated diagram of the enterprise software factory in five minutes: seven phases of the delivery lifecycle, three mechanisms that run through all of them, and how the system is deployed. There is no demo and no build. It is the fastest way to get the whole map, and the only video that covers audit, identity and cost as requirements.

Key takeaways

  • Coding was never the main constraint. Requirements, design, review, testing and operations are, so speeding up one step moves the bottleneck.
  • The seven phases run from requirements to operations, and an incident in production opens a new work item, which closes the loop.
  • Three mechanisms sit in every phase: a knowledge base, review loops with a retry limit, and human gates set by risk.
  • Each agent task runs in its own container with a clone of the repository and short-lived credentials, destroyed when the task ends.
  • At enterprise scale it adds three requirements: a record of who and which model produced each artifact, a least-privilege identity per agent, and tight tracking of token spend.

Chapters

  • 00:00 Why faster coding only moves the bottleneck
  • 00:45 The seven phases of the lifecycle
  • 01:30 Knowledge base, review loops and risk-based gates
  • 03:00 How it runs: events, an orchestrator, sandboxes and MCP
  • 03:45 Audit, identity and cost control
  • 04:30 How the human role changes

“They spend more time defining policies, curating the knowledge base, reviewing high-risk artifacts, and approving the decisions agents should not make alone.”

5 Minutes Tech, on what engineers, analysts and testers do in a factory

5. Inside OpenAI’s Agentic Software Factory

Owain Lewis, 10:48, published September 21, 2026. Watch on YouTube.

Owain Lewis walks through a Pragmatic Engineer article on how OpenAI structures development around agents, then rebuilds two pieces himself: a risk classifier that labels and merges pull requests, and an on-call agent. The diagram walkthrough is the value. The demos are short and the on-call part is done more fully in his next video.

Key takeaways

  • OpenAI’s diagram says “builder”, not “engineer”, because the person describing the task may be a product manager.
  • Review is split among specialist agents for data, infrastructure, cloud and security.
  • A classifier decides whether a change is low or high risk, and low-risk changes can merge and deploy without a person.
  • A severity bot investigates production incidents, and a performance factory sends regressions back into the pipeline as tickets.
  • He uses agents during rollout: deterministic code deploys the change, and the agent watches metrics and logs and can revert.

Chapters

  • 00:45 The architecture of the factory
  • 01:30 Specialist reviewers and risk classification
  • 03:01 The severity bot and the performance factory
  • 04:32 Demo: classify open pull requests and merge the safe ones
  • 06:48 Demo: an on-call alert agent
  • 08:18 Agents in the deployment process

“12 months ago, people would have said that that was crazy or reckless to use AI agents for the deployment process, but I think it makes a lot of sense right now.”

Owain Lewis, on letting agents watch a rollout

6. I Built A Claude DevOps Agent To Run My Software Factory

Owain Lewis, 16:55, published October 5, 2026. Watch on YouTube.

The follow-up builds the operations end of the factory: a site reliability agent on the Claude Agent SDK that takes alerts from a queue, investigates, and reports to a web console. Two demos show it diagnosing a bad deployment and a latency problem. Owain is the most candid person in the set about where agents fail.

Key takeaways

  • The agent follows an observe, orient, decide, act loop: see the alert, gather evidence with its tools, judge, then act or notify.
  • He recommends a human in the loop here, because agents often over-weight one signal and reach the wrong conclusion.
  • A lessons file gives the agent memory of past incidents. Outages tend to repeat, so this improves diagnosis over time.
  • The same agent can be told to watch the next deployment for five minutes and report errors or slower responses.
  • The quality of the agent comes down to the tools you give it. Writing them in code lets you limit exactly what it may do.

Chapters

  • 00:46 What downtime costs and why an agent responds first
  • 03:02 The observe, orient, decide, act loop
  • 05:20 Architecture: alert, queue, agent, console
  • 08:23 The triage loop and the lessons file
  • 09:53 Demo: a bad deployment and an approved rollback
  • 13:38 The agent’s tools and permissions
  • 15:09 Watching a rollout

“They can often jump to conclusions and so you have to be very careful about relying too much on whatever they tell you.”

Owain Lewis, on the limits of an on-call agent

7. Forget the Terminal. This Is the Future of AI Coding

Leon van Zyl, 17:04, published October 1, 2026. Watch on YouTube.

Leon van Zyl built a browser-based 3D office where each worker at a desk is a real Claude Code or Codex session and you manage them through a CEO agent, from the office or from a phone view. Under the cartoon it is a conventional factory on GitHub issues and pull requests. It is open source and installs with one command. Watch it for the interface idea more than the method.

Key takeaways

  • He sees two existing flavors of factory: one wired to GitHub issues, and one run through an org-chart tool such as Paperclip. His combines them.
  • The CEO agent creates issues, assigns them by role and proposes new hires, which you approve or let it make on its own.
  • The whiteboard is a live view of real GitHub issues. Finished work goes to a QA agent that can send it back to the developer.
  • Auto-merge happens only after QA and every check pass, and can be switched off.
  • As of this video he built it with Opus 5.5. The first prompt gave about half the app in one hour and 16 minutes for about $35.

Chapters

  • 00:00 Two flavors of software factory
  • 01:31 Tour of the office and the settings
  • 03:02 The CEO agent and the phone
  • 06:51 The whiteboard and GitHub issues
  • 08:21 QA reports, evidence and auto-merge
  • 10:37 How it was built and what it cost
  • 12:52 Installing it and hiring the first agents

“So that prompt gave me about 50% of what you’re seeing here.”

Leon van Zyl, on the single prompt that started the project

8. AI Engineer Paris 2026 Opening Keynotes: Mistral, Langfuse & Sizzy

AI Engineer, 1:13:35, published September 24, 2026. Watch on YouTube.

Three keynotes in 73 minutes, and only the middle one is about factories by name. Clemens of Langfuse gives the economic history of why new technology takes decades to show up in productivity. Kitze, the developer behind Sizzy, recounts his own road to a factory. Mistral’s VP of engineering covers enterprise agents and agent safety. The recording starts partway through the first talk.

Key takeaways

  • Clemens: steam, electricity and computers all arrived long before the gains did. AI infrastructure spending is about 1.8% of US GDP this year by his figures, already above the 1990s fiber boom.
  • Kitze names two failure modes: the “slop grenade thrower” who passes on unread agent output, and the “meat proxy” who only relays prompts between a manager and a model.
  • His method settled on one chat, one isolated environment and one pull request per task, after a worktree wiped his production database.
  • He admits the tinkering took him a long way from shipping, and says a plain terminal setup with Herdr and Pi was where he shipped most.
  • Mistral shows a prompt injection hidden in a GitHub issue making an agent leak a secret key, then blocks it with a policy enforced at the network level.

Chapters

  • 00:00 Clemens: general purpose technologies and the productivity puzzle
  • 14:47 Two clocks: invention and deployment
  • 21:42 Kitze: slop grenades and meat proxies
  • 26:15 Nobody agrees what a software factory is
  • 39:07 A throwaway machine per task, then guardrails
  • 53:31 Mistral: from a Slack alert to an agent investigation
  • 58:51 Agent safety and the prompt injection demo

“Almost nobody has an actual software factory and people are just scraping it together and we don’t agree on what a software factory should be.”

Kitze, at AI Engineer Paris 2026

9. How To Build a Dark Software Factory in 3 Months

Craft vs Cruft, 9:06, published September 10, 2026. Watch on YouTube.

Ray Myers answers the question he keeps getting, how to build a dark software factory, with a deadpan plan: take the model labs at their word, wait three months, and ask the model to build it. The joke carries a serious argument about trust. It is nine minutes, it contains no build steps, and it is the necessary counterweight to the other eight.

Key takeaways

  • “Dark” means lights off: nobody looks at the code any more.
  • He ties the idea to the labs’ own claims, citing Dario Amodei saying models may do all of software engineering end to end within 6 to 12 months.
  • The teams that report a working dark factory describe a great deal of engineering around the agents, which contradicts the claim that agents are ready to take over.
  • He cites Tessl describing the dark factory as real for a handful of elite teams and contested as a general model, and a claim that such teams should spend $1,000 in tokens per engineer per day.
  • His objection is professional: society trusts engineers to deliver systems that are secure, private and maintainable, and he cannot recommend handing that to a component he does not trust.

Chapters

  • 00:00 What a dark software factory is
  • 01:31 The labs’ definition of AGI and the capital behind it
  • 03:03 The three month plan
  • 04:35 What the elite teams actually built
  • 06:51 An argument from trust
  • 07:37 Professional ethics

“Anything you would have had to do to set up a dark software factory would have been software engineering.”

Ray Myers, on why nobody should need to build one if the claims were true

Start where Steve says he started: one skill, run by hand, on one kind of issue you can verify. Add the proof step from Ross Mike before you add a second agent, and the risk gate from Owain Lewis before you let anything merge itself. If you watch only one video, make it the Builder.io deep dive. If you watch two, add Ray Myers, so you hear the case against before you turn the lights down.

Related Reading