A production AI agent does not stay finished. Your estimating rules change. A vendor changes its lead-time language. A project manager invents a better handoff. A customer asks a question your workflow never anticipated. Every one of those changes can make yesterday's reliable agent wrong in a new way.
OpenAI's July 2026 Presence announcement describes production agents as systems that need policies, evaluations, approved actions, escalation rules, and controlled improvement after launch. Anthropic makes the same operating problem visible from another direction: an agent's behavior depends on the model, harness, tools, and environment, and it must learn when to act versus when to return a decision to a person. The lesson for a building business is simple: treat an agent update like a change order, not an informal prompt tweak.
Production reveals requirements the kickoff missed
A pilot can cover the happy path. Live work exposes the exceptions: the proposal with two alternates, the selection without a final approval, the invoice that references an old purchase order, or the lead who asks for a promise your team cannot make. These are not merely bad outputs. They are evidence that the workflow's sources, rules, or escalation boundary are incomplete.
OpenAI says production sessions, escalations, and quality signals should reveal where an agent needs attention, after which teams can test a proposed change against the production version and approve a controlled rollout. That is a useful operating pattern even if a ten-person remodeler never buys an enterprise agent platform.
Open a change record before editing the agent
- Trigger: the failed session, escalation, policy change, source revision, or user correction that exposed the gap.
- Impact: which clients, projects, roles, systems, and past outputs could be affected.
- Proposed change: the prompt, source rule, tool permission, validation, or escalation path that should change.
- Acceptance cases: the exact examples the new version must pass, including the original failure.
- Approver: the person who owns the affected business rule and can authorize the rollout.
- Rollback: the prior version and the condition that sends the workflow back to it.
This record keeps the team from solving a visible failure while quietly creating three new ones. If an agent missed an expedited freight charge, the fix is not automatically “always flag freight.” The business owner may need to define which vendors, thresholds, project phases, and contract types require the flag. The agent change should encode that rule and preserve the evidence behind it.
Turn every real failure into a regression test
Keep a small evaluation set made from real work: accepted examples, corrected examples, ambiguous cases, and actions that must escalate. Before a new version reaches the team, run both the new failure and the old cases. Score source selection, required-field coverage, policy compliance, tool use, uncertainty, escalation, and the final artifact the next person receives.
A selections agent might need to prove that it uses the current approval log, refuses to invent availability, separates allowance from upgrade cost, and routes any client-facing commitment to the project manager. A bid-comparison agent might need to catch scope exclusions without choosing a contractor. The evaluation should test the business boundary, not whether the prose sounds polished.
Separate advice changes from permission changes
Changing how an agent drafts a summary is not the same as letting it send the summary. Adding a new price book is not the same as letting it update an estimate. Expanding a knowledge source is not the same as granting access to email, accounting, or project records. Anthropic's guidance distinguishes model behavior, instructions, tools, and environment because each layer changes the risk in a different way.
Use a stricter review whenever a change expands data access, tool access, financial authority, client communication, scheduling, purchasing, or system-of-record writes. Those releases should have named approval, a limited first audience, observable logs, and an immediate stop path.
Roll out in rings
Start the revised agent on saved test cases. Then use it in shadow mode, where it produces work without taking action. Next, release it to one trained operator or one low-risk workflow. Compare acceptance, corrections, escalations, and downstream rework against the prior version. Expand only when the evidence supports it.
The useful dashboard is not a count of chats. Track version, workflow state, source set, reviewer, approval status, escalation reason, first-pass acceptance, correction category, and rollback events. That record tells you whether the agent is improving or merely changing.
Publish the method, not the magic
Google says there is no separate AI-only schema requirement for AI Overviews or AI Mode. Helpful, original, technically accessible content and accurate structured data remain the foundation. The business-relevant proof an answer summary cannot replace is the operating method: what triggered a change, which cases tested it, who approved it, how it rolled out, and what the team learned.
Continue with the operating system
- Self-Improving AI Agents Need A Promotion Gate
- Define The Deliverable Before You Hire The AI
- Explore practical AI paths for your team
Sources Read
- Introducing OpenAI PresenceOpenAI
- Trustworthy agents in practiceAnthropic
- Google's Guide to Optimizing for Generative AI Features on Google SearchGoogle Search Central
Next step, if this note maps to a problem on your desk: Private Training — a private working session for your team ($1,500+).