The Agent Fleet Still Needs an Owner
The leading agent on the Remote Labor Index now completes roughly one project in five at a client's standard. I'm interested in what happens between the agent saying 'done' and the customer agreeing.
The Enterprise's computer can do an extraordinary amount of work. Picard still has to be captain.
That seems worth remembering as we turn increasingly capable agents loose on real assignments. The software can plan, research, build, and announce that it has finished. Somebody has to decide whether the result is any good, and somebody has to answer for it when it isn't.
As of September 21, Scale's Remote Labor Index puts GPT 6 Astra at a 20.83% automation rate and Fable 5.1 at 17.92%. The benchmark uses freelance projects and expert reviewers to judge whether the finished deliverable meets a professional reference standard. Its collection contains 240 projects, with 230 reserved for quantitative evaluation and ten released publicly.
Roughly one project in five for the leader. That's substantial progress, with most projects still falling short. It measures particular agent systems under particular conditions, across selected kinds of work. Reading the result as “a fifth of all jobs are now automated” would be nonsense. So would assuming that a human working closely with an agent gets the same result as an autonomous benchmark run.
What interests me is the standard being applied. Would a reasonable client accept the work? That's a question I recognize. It's the one that belongs at the end of every demo.
The customer gets the result
Consider a website that builds successfully but won't let a customer complete a purchase. Or a report with handsome charts and a wrong central comparison. The work can look finished to the person reviewing the process while remaining useless to the person who needed it.
An agent's completion message tells you where to start checking. Open the file. Try the workflow. Use the thing as its intended user would. That last part sounds almost embarrassingly obvious, which may be why it's so easy to skip when the rest of the process has been impressive.
The temptation to skip it will grow. If an agent produces ten times as much work, a human may have ten times as much to review. Unless the checks improve too, faster production can leave you with more unfinished business and less time to notice it.
For a small venture, I would begin with a service whose results can be checked at a sensible cost. A narrow job with clear requirements gives you somewhere to learn. You can find out which mistakes recur, which checks catch them, and how often a human has to intervene before expanding the offer.
That can be a useful business long before an agent can handle every kind of remote work. The customer doesn't need general intelligence. They need the particular thing they paid for to work.
The founder's time belongs in the accounts
I still owe this series the revenue and cost accounting promised in April. This month's benchmark gives me a useful occasion to spell out what I think those accounts have to include. A list of launches wouldn't tell readers whether the experiment was paying its way.
Start with all the attempts that went into the accepted result. If it took ten drafts, count ten. Add the subscriptions and usage charges, then include the time spent explaining the assignment, checking it, repairing it, and helping the customer afterward. The founder's hours cannot quietly disappear from the calculation just because nobody issued an invoice for them.
That matters particularly when the attraction is supposed to be freedom from work. If you spend every evening correcting your agents, you may have built something productive. You may even enjoy it. But you haven't yet demonstrated that the operation gives you more time to live.
I would want to know whether it can repeat a good result without requiring the same rescue next time. An ingenious intervention can save a project. If the intervention becomes part of every delivery, it is part of the service you are selling and should be priced accordingly.
None of this requires a complicated dashboard. A short record of each completed job, its costs, and the human minutes involved would be much more informative than an impressive count of tasks the agents say they performed.
Decide who can say yes
There are also decisions to make before the work starts. What may the agent change? How much may it spend? Can it promise something to a customer? When should it come back to a person?
On the Enterprise, the computer doing more doesn't make the captain less responsible for an order. A business has the same practical problem. If software sends the wrong thing to a customer, “the agent did it” is a poor answer from the person operating the service. The customer needs a correction and somebody they can reach.
These responsibilities need a place in the organization. In a one-person business they land on the founder. In a larger company, someone needs the authority and time to deal with them. Putting a person's name beside an agent isn't enough if that person is supervising more work than they can reasonably inspect.
This is also work beginners can learn. Give someone a small assignment to supervise and a more experienced colleague to help when the result is doubtful. Let them practice deciding what is acceptable and explaining why. We don't need to reserve every meaningful decision for the same few senior people while wondering where their replacements will come from.
Actionable Steps
Pick one service or workflow you want to run with agents. Write down what the user must be able to do when it is finished. Keep a record of the failures as well as the successes, and include the human time needed to put things right.
Then ask the person using it. Their answer may be less flattering than the agent's report. It is also the answer that tells you whether to keep going, fix something, or choose a different problem.
I remain optimistic about what small teams and individuals can build with these tools. I want more of us to have that chance. But when I imagine a post-labor life, it includes being able to leave the desk. That's a result I'd like the experiment to earn.
LLAP. 🖖
Related reading elsewhere
Follow the experiment
One essay a month on where the post-labor economy is actually heading — capability, evidence, policy, and a live experiment in post-labor income. No hype, receipts included.
Follow Kenn on LinkedIn