The cheapest AI run is not necessarily the least expensive work. A low-cost model can produce a plausible bid comparison, schedule summary, or product recommendation and still create hours of checking, correction, and client-facing risk. The useful unit of cost is not the token. It is the accepted outcome.
OpenAI's July guidance makes the same distinction: leaders should measure useful work per dollar, including tasks completed, decisions improved, and workflows ready to scale. It recommends counting model and tool usage, attempts, completion rate, latency, and human review when comparing models. That is a better operating lens for a building business than a dashboard full of messages and credits.
Activity is not an outcome
A team can use AI every day without improving a single business process. More chats, documents, or automated runs may simply move unfinished work downstream. If a project manager still has to reconstruct the source packet, verify every total, and rewrite the client note, the system has generated activity rather than capacity.
Start by naming the accepted outcome. For a quote comparison, it might be a complete scope table with every number traceable to the current vendor documents, all exclusions exposed, and a qualified reviewer signing off. For a meeting follow-up, it might be an approved action register assigned to named owners. Acceptance is a business state, not a model score.
Build a cost-per-accepted-outcome record
- Run cost: model, retrieval, document processing, and tool charges for the job.
- Attempts: initial run, retries, and any manual restarts required to finish.
- Review time: minutes spent checking sources, calculations, policy, and presentation.
- Correction time: minutes spent repairing missing, unsupported, or incorrectly formatted work.
- Acceptance result: accepted, accepted with edits, rejected, or escalated.
- Business effect: cycle time reduced, risk avoided, capacity created, or revenue protected.
Keep the record attached to the workflow run. Over ten or twenty real jobs, the pattern becomes more useful than a demo. You can see whether a more capable model reduces retries, whether a better source packet cuts review time, or whether the task should remain deterministic instead of agentic.
Define the quality bar before comparing models
OpenAI's agent-building guide recommends establishing evaluations with the most capable model first, then testing whether smaller models still meet the accuracy target. The order matters. If the team optimizes price before defining acceptable work, it can quietly lower quality while celebrating a cheaper run.
Use representative building-industry cases: an outdated drawing, conflicting quote totals, an allowance without a basis, a product with a missing lead time, and a client request outside the approved scope. Grade source use, completeness, calculation accuracy, escalation, and whether the final artifact is genuinely review-ready.
Treat corrections as workflow evidence
Production systems improve when they preserve what happened after launch. OpenAI's Presence description emphasizes reviewing sessions, escalations, and quality signals, then testing proposed changes before controlled rollout. A smaller company can use the same loop without enterprise infrastructure: log the correction, classify its cause, update the source packet or instruction, rerun the relevant eval, and approve the change.
This separates model problems from operating problems. Repeatedly missed exclusions may mean the quote template is inconsistent. Wrong product status may mean the catalog source is stale. Excessive review may mean the acceptance criteria are vague. Buying a larger model will not repair every broken input or undefined handoff.
Publish the operating proof
Google's guidance for AI Overviews and AI Mode still calls for helpful, original, technically accessible content and structured data that matches the visible page. There is no special AI-only schema. A cost-per-accepted-outcome framework gives searchers and answer systems concrete, source-grounded guidance while preserving the business value that a summary cannot deliver: the templates, thresholds, review roles, and evaluation cases needed to run the workflow well.
Continue with the operating system
- AI Can Cross Job Lines. Accountability Cannot
- Coverage Evals Stop Agents Missing What Matters
- Explore practical AI paths for your team
Sources Read
Next step, if this note maps to a problem on your desk: Private Training — a private working session for your team ($1,500+).