AI, usefully · 4 min read ·
An AI agent needs a budget, not just a goal
Give an AI agent this instruction:
Research the market thoroughly and return the best answer.
When should it stop?
After five searches? Fifty? When two sources agree? When it has checked every plausible explanation? “Be thorough” gives the agent a direction, but no definition of enough.
This matters because an agent can repeatedly search, call tools, ask specialist agents for help and revise its answer. Each step consumes time and computing resources. More activity may improve the result—but it can also repeat the same work, collect weaker evidence and make the final answer harder to audit.
A production agent therefore needs two things:
- A goal describing what success looks like.
- A budget describing how much work it may do to get there.
The budget is more than money
Cost is the obvious constraint, but it is only one part of an execution budget.
A useful budget may limit:
| Budget | What it controls |
|---|---|
| Model calls | How often the agent can ask a model what to do next |
| Tool calls | Searches, database queries and external API requests |
| Time | How long the task may run |
| Tokens | How much information models read and generate |
| Specialist agents | How widely the supervisor can delegate |
| Retries | How many times a failed step may run again |
| Data | Which records, files or date ranges may be processed |
These limits belong in application logic. Telling a model to “avoid unnecessary searches” may influence its behaviour, but it does not enforce a maximum.
The application should count the work and stop or change course when a limit is reached.
Match the effort to the question
Not every task deserves the same architecture.
“What is the delivery status of order 104?” probably needs one lookup.
“Compare five vendors across security, pricing and implementation risk” may benefit from parallel research and several specialist assignments.
The mistake is allowing the second workflow to become the default for the first.
Anthropic reported that its multi-agent research system was valuable for broad, parallel research but used substantially more tokens than ordinary chat interactions. It also found that tasks requiring tightly shared context or many dependencies were poorer candidates for multiple agents.
The lesson is not that multi-agent systems are wasteful. It is that their additional cost should buy something specific: greater coverage, independent investigation or faster parallel work.
If one model call and one tool can answer the question reliably, adding a supervisor and four specialists mostly creates more places for the workflow to fail.
Give the supervisor an effort policy
A supervisor should not have to invent the amount of effort from scratch.
An effort policy might say:
| Request type | Initial plan |
|---|---|
| Direct record lookup | One read-only tool call |
| Defined comparison | Retrieve the required sources, calculate, validate |
| Broad research question | Divide into distinct research areas, then synthesize |
| High-impact action | Research and prepare only; stop for approval |
| Ambiguous request | Ask a focused question before using expensive tools |
This is still a starting policy, not an automatic guarantee of quality. But it gives the system a sensible default.
The supervisor can then escalate when the evidence justifies it. A conflicting source may warrant another search. A missing identifier should trigger clarification, not ten speculative lookups.
Define what “enough” means
A budget tells the agent when it must stop. A completion rule tells it when it should stop early.
For a research task, completion might require:
- every requested dimension is covered;
- important claims have supporting sources;
- material disagreements are visible;
- the agent has no unresolved gap that would change the recommendation.
Notice what is absent: “use all twenty searches.”
The budget is a ceiling, not a target.
Once the completion rule is satisfied, more calls may add volume without adding value. The supervisor should be able to finish with budget remaining.
It should also be allowed to stop without an answer. If the evidence cannot support the requested conclusion, exhausting the budget should produce a clear account of what is missing—not a confident guess created because the workflow ran out of time.
Retries need their own limit
A failed tool call creates a tempting loop:
Call tool → receive error → try again → receive error → try again
Some failures are temporary. Others will never improve through repetition: access is denied, an identifier is invalid or the requested record does not exist.
Retry policies should distinguish those cases.
A temporary timeout might allow a small number of attempts with increasing delays. A permission failure should stop immediately. Invalid input should return for correction.
The retry count should stay in the task state so a restarted workflow does not quietly begin the same attempts again.
Measure useful work, not visible activity
A highly active agent can look impressive while accomplishing very little.
For each run, record:
- whether the task succeeded;
- which tools were used;
- how many calls and tokens it required;
- total time;
- repeated or failed operations;
- whether a human had to repair the result.
Then compare similar tasks.
If a new workflow uses three times as many calls without improving correctness or coverage, the extra orchestration has not earned its place. If parallel specialists cut a complex investigation from an hour to ten minutes while preserving source quality, the expense may be justified.
Cost, latency and reliability should be evaluated together.
Try this
Before your next multi-step AI task, add this to the brief:
Propose the smallest plan likely to answer the question. State the tools required, the maximum number of searches or tool calls, what would justify additional work and the conditions for stopping. If the available evidence cannot support an answer within the budget, return the unresolved gaps instead of guessing.
Then inspect what the assistant proposes.
The goal is not to make AI do less. It is to make every additional step earn its cost.