People assume handing over a multi-step task means specifying less. My experience is the opposite: the cost of a bad brief scales with how much work it triggers.
The difference between a conversation and a delegated task is not sophistication. It is who holds the sequence.
Why the brief matters more, not less
People assume that handing over a multi-step task means specifying less, because the thing works out the steps itself. My experience is the opposite.
In a conversation, a vague instruction costs one message and I correct it immediately. In a delegated task, a vague instruction runs for a while and produces something coherent, confident and subtly not what I wanted. The cost of a bad brief scales with how much work it triggers.
Delegating to software has the same requirement as delegating to a person: if you cannot define done, do not hand it over.
How I brief a delegated task
1. State the outcome, not the procedure
What should be true when this is finished. Procedures constrain the approach unnecessarily and outcomes are what I can actually check. The exception is any step where the method genuinely matters, which I then state explicitly.
2. Name what must not happen
Do not modify the source files. Do not invent figures. Do not send anything. This is more important in a delegated task than in a conversation, because there is no intermediate moment where I would notice.
3. Supply the material rather than describing it
The actual files, the actual data. A task working from a description of the data will produce something shaped like an answer rather than the answer.
4. Define done in checkable terms
Not looks right. A specific artefact, in a specific place, containing specific things. If I cannot state that, I am not ready to delegate the task and I should do it myself.
5. Ask for the working, not just the result
What was changed, what was assumed, what it could not resolve. The assumptions are where the errors are, and they are invisible in a finished-looking output.
6. Verify against the definition, every time
Not a glance. Check the thing I said would be true. The temptation to skim increases as the outputs get more consistently good, and that is exactly when it becomes risky.
What I hand over, and what I do not
| What it is good for | Where it stops | What it costs me |
|---|---|---|
| Reformatting and restructuring material I supplied | Verifiable against the source. | A real check, not a skim. |
| Cross-checking a set of documents for inconsistencies | It finds things I miss, and I confirm each one. | Confirming every flag. |
| Assembling something from scattered pieces | Strong, provided the pieces are supplied rather than found. | Reading the assumptions it states. |
| Anything that sends, pays, publishes or deletes | Not delegated, at any confidence level. | One approval step, permanently. |
What it is honestly worth
The gain is on tasks with many mechanical steps and a checkable end state. Work that would take an afternoon of tedious assembly and produces an artefact I can inspect. That is a genuine category and it is larger than I expected.
The gain is much smaller on work requiring judgement I cannot specify. If I could not write down what good looks like, delegating it produces something plausible that I then have to think about from scratch, which is slower than doing it myself.
The industry context is worth holding onto. Gartner predicted in June 2025 that over 40 percent of agentic AI projects would be cancelled by the end of 2027, citing escalating cost, unclear business value and inadequate risk controls. That is a forecast rather than a measurement, from a research firm rather than a vendor. It matches what I see: the failures are about scope and governance rather than capability.
How it breaks
Verification decays. The most likely failure and the hardest to notice. After a run of good outputs, checking becomes skimming. Sample properly at intervals rather than trusting the streak.
Delegating something unverifiable. If there is no way to tell afterwards whether it was right, you have not saved work; you have accepted risk you cannot price.
Scope creeps within a task. A broadly worded brief produces broadly interpreted work. Narrow briefs, checked, beat wide ones every time.
It touches something live. Keep anything that sends, pays, publishes or deletes behind a human. This is the same line as everywhere else on this site, and it is the one I would never move.
How to tell whether it is working
Count how often the output needed correction, and how long the correction took. If corrections take longer than the task would have, the brief is the problem rather than the tool. And if you have stopped checking, the number you are tracking is no longer meaningful.