I gave the agent one prompt and stopped touching the keyboard.
Hour after hour, tickets closed, commits landed, and tests went green without any of my attention. It was my second AGI moment, after the Opus 4.5 release.
That first project ended with about 26,000 lines of code, and the implementation run never asked me anything. The longest run lasted about 10 hours.
No ticket needed me. The only inputs were decisions I had reserved for myself in advance, and the agent flagged them early and kept working. Every pause was mine.
I’ve used the same flow on three very different projects:
- A native mobile app, ported from a web prototype.
- A backend, rewritten from scratch.
- An internal CRM application at my company.
Different languages, frameworks, and databases. Together, more than 100 tickets with no human in the implementation loop.
Your work moves before and after the run
- Before: planning. It spread across a couple of days but took me a few hours at most.
- During: nothing, apart from a few planned decisions.
- After: reading ticket comments, a quick code read, testing the app, and running the release myself.
My setup is boring
No dozens of custom skills. No multi-file convention docs, no folder of code examples, no clever harness. A goal prompt, a few markdown files, and checks the agent can run itself.
I didn’t know how to start an autonomous loop like this. So I asked an LLM, and its answer became the first version of this setup. I just asked my question.
Dax put it better:
a weird inversion with LLMs is the models improve faster than the tinkerers
— dax (@thdxr) October 5, 2026
when i see people with custom workflows and setups they’re all addressing problems that don’t exist anymore
the person naively using vanilla codex is more likely to be experiencing state of the art
The checklist
1. One source of truth.
- Straight port: the old code is the spec and stays read-only.
- Rewrite: the old code holds the problems you want gone, so written decisions define the new behaviour.
Pick one rule and write it down.
2. Every open question answered, with a recommendation. The agent reads the code and writes the questions, each with options and a recommended answer. I answer in batches, mostly “as recommended”. The few I overrule are where I matter.
3. A plan reviewed by another provider’s model. A Claude model wrote my backend plan. A GPT model reviewed it cold:
“I would not approve the plan for autonomous implementation yet.”
It was right; one issue would have let an agent delete the old code too early. Answering questions gives you decisions. A cold review finds where they contradict each other.
4. A ticket graph. Each ticket is a markdown file with Status: and Depends on: lines, and a map orders them. The agent takes the first ticket whose dependencies are done.
5. Architecture before features. The prototype showed me what not to build again. My prototype backend had every feature implemented twice, two API generations, several error formats, and dozens of database triggers.
For the rewrite I redesigned the database from scratch and removed the triggers. Every write goes through one command runner that handles validation, concurrency guards, and safe retries. The API contract is frozen, and the mobile client is generated from it.
Then a brand-new agent got only the architecture guide and built a small feature in a throwaway copy. I fixed whatever confused it. Agents copy patterns, so the first one has to be right.
6. A list of what only you can do. Credentials, approvals, production. The agent should know them on day one and work around them.
Let the agent test itself, and get out of its way
This decides whether a run lasts ten minutes or ten hours. After hours alone, an agent will think it’s done when it isn’t. Only checks it can run itself protect you from that. And every check it can’t run is a place where it stops and waits for you.
What I give the agent:
- An executable spec. For the port, a script ran the old logic and saved thousands of input/output cases. The new code replays them.
- Tests first, against a real local database and storage emulator, not mocks.
- Mechanical rules. An architecture checker fails the build on a wrong import or an unsafe cast.
- One command that means “done”: types, lint, architecture, formatting, API contract, tests. Green means the ticket can close.
- Fresh reviewers. Two read-only reviewers, one against the spec and one against the standards. They found real defects the tests missed.
What I take away, because it blocks the agent:
- Approval prompts for local work. Only production is locked.
- Branches and pull requests. One commit per ticket, straight to main. My first prompt asked for PRs, and the agent stopped after one ticket to wait for me.
- Questions that can wait. The agent records them, marks the ticket, and moves on.
- Shared resources. Parallel agents each get their own dev server port.
- Taste. The port had no styling and a fixed table mapping web patterns to native ones.
My remaining gap: no agent ever used the UI. After a later restyle, one screen went blank while every test stayed green. Anthropic’s and OpenAI’s published setups let the agent drive the UI. That’s my next layer.
The goal prompt
Each run starts with one goal-mode prompt, which keeps the agent working until a condition is met. A GPT model wrote my backend prompt from three short requests. I pasted it at every restart. Its parts:
- Goal and stop rule. Finish the scope without asking between tickets. Stop only when everything is done or what’s left needs a human.
- Restart routine. Read the progress file first; after compaction, check it against the real files.
- The ticket loop. Failing test, pass, repeat. Verify, review, commit.
- Limits. “Authorization to commit directly to main is not authorization to deploy.”
- Done. Every criterion has evidence. If blocked, finish everything else and say exactly what’s needed.
progress.md and TODO.md
Long runs get compacted many times, so the state lives on disk.
progress.md is a restart checkpoint:
- where to resume,
- each ticket’s state and evidence,
- open questions,
- decisions the agent made.
The agent updates it after every meaningful step. The run survives compaction, a restart is one paste, and I read its questions afterwards instead of being interrupted.
TODO.md points the other way: things only I can do.
The test for both: can a fresh agent with no memory continue from them? One catch: when subagents work in parallel, the file falls behind and git becomes the truth.
Roles, not personas
I don’t write “You are a senior code reviewer.” A reviewer is a model, a task, and criteria. I assign models by role (more on picking them in The right model, not the best model):
- Integrator (Opus 5.5): owns the ticket map, the progress file, and every commit.
- Implementers (GPT-6 Luna, GPT-6.1 Sol, Opus 5.5): subagents doing tickets.
- Reviewers (GPT-6 Astra, Opus 5.5): read-only.
- Advisor (Fable 5.1): a background model commenting on the main session. My weakest role: expensive, and often late.
The flow doesn’t depend on the tool. The port ran in Claude Code with no subagents. The backend ran in omp (oh-my-pi), where the setup got much easier to use, more efficient, and easier to control.
omp is built for this way of working:
- Models are configured by role (default, fast, slow, planning, subagent tasks, advisor). Each role can come from a different provider, so Claude and GPT models share one session on my existing subscriptions.
- Subagents are first-class.
- A multiple-choice ask tool with a recommended option turns my gates into one-click answers.
- Compaction is a setting. I set it to a quarter of the model’s window.
- Cost per message. omp estimates it, so I see which role eats the budget.
More on that in a separate post, coming soon.
Production stays human
The agent could commit to main. It could not deploy.
The harness refused production writes even after I approved them. So the agent wrote release scripts, and I ran them. Why the agent shouldn’t hold your production keys: The code your AI writes is untrusted.
OpenAI’s harness team made human code review optional. For code, agents and checks can be the gate. Production can’t be undone. I made the case against ritual review in Code review is becoming a fiction.
Results and limits
Across three projects: more than 100 tickets, a 10-hour run, about 26,000 lines in one project, the old backend deleted, and every app in use. Review, tests, gates, and planned audits caught every bug I know of, except that blank screen. Nothing broke in production.
The limits:
- My own app has one user: me. The company application has real users.
- Each project had a working app to check against, and for the rewrite, written decisions about what to change. Without a definition of correct, you’re back in the loop.
- Parallel agents in one folder collided. I prefer one ticket at a time; if you go parallel, use a worktree per agent.
- My review was light by design: a quick code read, using the app, and every ticket comment. Why that’s enough: I don’t read most of the code my agents write.
For a packaged version, try /implement-spec from Matt Pocock’s skills. Start where “correct” is clearly defined.
Then spend your hours before the run and after it.