An AI agent can work for hours while your estimator, project manager, designer, purchaser, or controller does something else. That does not make every agent hour equivalent to an employee hour. The useful unit is accepted work: an output that is supported by the right sources, inside the agent's authority, reviewed at the right level, and safely entered into the business process. Everything else creates supervision work now or recovery work later.

OpenAI's recent research describes a shift from short chatbot interactions to delegated, long-horizon agent tasks and reports that its heaviest internal users run many agent tasks in parallel. Anthropic's trustworthy-agent guidance makes the operational counterweight clear: agents plan and act with less direct oversight, which increases the importance of human control, secure interactions, transparency, and privacy. A building business should read those ideas together. More agent capacity creates leverage only when the company also designs the capacity to supervise it.

Separate runtime from accepted work

Do not report that an agent ran for six hours as if the business received six hours of value. Record what deliverable it attempted, what source packet it used, what actions it took, what stopped it, and what a person accepted. A purchasing agent that reviews 200 order lines but sends 45 ambiguous exceptions to a manager may still help. The decision depends on review minutes, exception quality, missed risks, and the cost of repairing any wrong action—not the length of the run.

Price the four kinds of supervision

  • Setup supervision: preparing the approved files, rules, permissions, and definition of done before a run starts.
  • Review supervision: checking evidence, calculations, scope, and proposed actions before work is accepted.
  • Exception supervision: resolving missing information, conflicting records, unusual conditions, and decisions outside the agent's authority.
  • Recovery supervision: reversing a bad update, correcting downstream records, notifying affected people, and improving the test set after failure.

These costs land on different roles. An office coordinator may prepare a bid packet, an estimator may review scope, an owner may approve margin, and accounting may repair an incorrect cost-code update. Add their minutes separately. A workflow that saves junior assembly time but consumes scarce owner judgment can reduce total capacity even when its demo looks fast.

Set a supervision ratio before adding parallel agents

For each workflow, divide accepted deliverables by total human supervision hours. Then track the percentage accepted without material correction, the number of exceptions per run, the age of the review queue, and the recovery time after mistakes. This is not a universal benchmark. It is a local capacity model that shows whether adding another agent run will clear work or merely enlarge the queue waiting for your best people.

Suppose an agent prepares five draft change-order packets overnight. If the project manager needs twelve minutes to verify each packet and the owner needs five minutes to approve price, that batch requires 85 minutes of next-day supervision before anything moves. Schedule that review capacity. If the team can reliably review only three packets, launching ten is not scale; it is hidden work in progress.

Match review depth to consequence

Not every output needs the same gate. A weekly internal risk draft can tolerate corrections before the meeting. A client promise, purchase commitment, payroll change, schedule update, or price decision needs stronger evidence and explicit approval. Define review tiers by consequence, reversibility, source quality, and authority. The agent should know which tier applies and stop when it cannot prove the requirement was met.

Give supervisors a queue they can actually work

A useful review item contains the requested deliverable, relevant source links, the agent's proposed result, unresolved conflicts, actions already taken, the exact decision needed, and a deadline tied to the underlying workflow. Sort the queue by consequence and blocking impact rather than arrival time. Preserve accept, revise, reject, escalate, and reopen states so the business can learn where supervision time is going.

Test the capacity model in shadow mode

Run the agent beside the existing process without allowing consequential writes. Measure setup minutes, review minutes, exception minutes, accepted-without-change rate, missed issues, false alarms, and recovery work. Use real messy packets, not only complete examples. Expand runtime or parallelism only when the accepted-work rate rises without overwhelming the named reviewers or weakening the control boundary.

Publish the operating method, not the runtime claim

Google says AI Overviews and AI Mode rely on the same foundations as Search: useful, original, accessible content and accurate structured data, with no special AI-only schema required. The part an answer summary cannot replace is the business artifact: your supervision categories, role-based review minutes, consequence tiers, queue design, acceptance data, and decision about where additional agent capacity will actually help.

Continue with controlled capacity

Sources Read

Next step, if this note maps to a problem on your desk: Private Training — a private working session for your team ($1,500+).

Related Field Notes