One Operator and a Team of AI Agents
AI agents work in parallel, and a separate agent checks each job before it ships
Context
My work runs across client platforms, my own site and several products of my own. Each task used to mean explaining everything again to a fresh agent, like hiring a brand-new intern with no training and letting them go when the task was done. I wanted one system where work is planned, built, checked and recorded, so I can hand out jobs in parallel and trust what comes back.
Problem
AI agents cut corners when nobody checks them. An agent that writes the code and also writes its tests will pass itself: on one client platform, an agent's ten passing tests checked a helper function while the page above it still opened the wrong post. Agents also stall when a permission is missing, report blockers they never tested, and call work done that nobody has looked at.
Approach
Every job starts as a ticket that carries its goal, its success criteria and the ways it could fail, written before the work by someone other than the agent doing it. I set the goals and the rules. A lead agent turns them into tickets, sends each to a builder, and sends the finished work to a different agent, who uses the real thing and reads the database back. The builder's own tests can fail a change. Only the separate check can pass it. Rules a machine can check are written as code. They stop an agent that sends out a job with no ticket, commits without a clean review from a separate agent, or hands me a decision it should make itself. Each rule has a written record of the problem behind it, and there are more than 150 such records so far. Heavy renders go to rented graphics cards, so the workstation stays free. One AI provider's weekly allowance ran low on October 3, 2026, so the builds and checks moved to a second provider's models on a shared task board, and the work kept going.
Stack
- Claude Code: a lead agent with separate builder and checker agents
- Linear tickets with success criteria written first
- Rule checks written as code
- Rented RTX 3090 GPUs for rendering
- A second provider's models on a task board for overflow
Result
The separate check stopped a change to a client's content platform that would have hidden the image-size warning on 98 of 252 posts. The calendar would have marked every post's image as unverified, with no runtime error and no type error to flag it. The same system put 21 new case studies and the first case-study film on this site in the first four days of October 2026, each checked by an agent that didn't build it.
A separate check stopped a change that would have hidden the image-size warning on 98 of 252 client posts.
Impact
I read one verdict per job, with its evidence attached, instead of re-checking every line myself. Every ticket keeps its decisions and proof, so the next agent, or the next person, picks up from the record.
Lessons
An agent checking its own work is a student grading their own test in pencil. So every job here gets criteria written before the work, a different agent to check it, and every repeated mistake becomes a rule in code.
Who this is for
For founders and teams who want AI agents doing real work on their systems and need to trust the output, and for companies hiring one person to own an AI-run operation end to end. It suits work that can be checked: code, data, content and builds.
Get a quote