OpenAI introduced Presence on July 22 as a managed product for deploying voice and chat agents in high-value customer and internal workflows. The product is aimed at eligible enterprise customers, but its most useful lesson applies to a much smaller remodeler, builder, design firm, showroom, supplier, distributor, or trade contractor: getting an agent to work once is not the finish line.
OpenAI describes a production system built around a specific job, limited knowledge and system access, company policies, approved actions, simulations, evaluations, guardrails, escalation rules, and a controlled process for improving the agent after launch. That is a different mental model from buying software, configuring it, and assuming the installation is complete.
Launch creates evidence you could not collect in a demo
A test can cover known requests: find the latest selection sheet, draft a response to a warranty question, summarize an open purchase-order list, or route a new lead. Production introduces the language, missing information, policy conflicts, stale files, unusual job conditions, and tool failures that the test set did not anticipate.
That does not mean testing failed. It means the first release established a baseline. The operating question becomes whether the business captures those new cases, labels the correct outcome, and improves the workflow without breaking cases that already worked.
Give the agent a maintenance owner
Someone must own the agent after the consultant, software vendor, or enthusiastic employee finishes the first version. That owner does not need to be a programmer. They do need authority to review failures, confirm the governing source, gather corrections from the team, pause risky actions, and approve changes.
For a selection-status agent, the owner might be the operations manager. For a product-support agent, it might be the showroom manager. For a bid-intake workflow, it might be the preconstruction lead. Assign the role to the person who can judge the work, not simply the person who knows the AI tool best.
Separate business changes from model changes
Agent behavior can drift even when the model stays the same. A vendor changes its lead-time language. The company revises its warranty policy. A form adds a field. A folder moves. A team begins using a project status differently. Each change can make yesterday's correct instructions incomplete today.
Keep a change record that names the trigger, source document, instruction or tool affected, proposed revision, test results, approver, rollout date, and rollback path. Version the workflow as a package: instructions, connected sources, permissions, tools, model, eval set, and escalation rules. Otherwise the team will know that the agent changed without being able to explain why.
Production review should feed regression tests
OpenAI says production sessions, escalations, and quality signals can reveal where an agent needs attention, and that proposed updates can be tested against the production version before a controlled rollout. A building business can use the same loop at a smaller scale.
When the agent misses a scope exclusion, cites an outdated price sheet, routes a warranty request incorrectly, or drafts a message that needs a major correction, save a privacy-safe version of that case. Record the expected result and why it matters. Add it to the repeatable test set before changing the workflow. Then run both the new case and the older passing cases.
This protects against the common repair that fixes one visible failure while quietly creating two new ones. The scorecard should measure business behavior: correct source used, required fields covered, prohibited action avoided, appropriate escalation, acceptable reviewer edits, and a complete audit record.
Use a simple post-launch operating rhythm
- Name one workflow, its owner, its authoritative sources, and its allowed actions.
- Log every run, source used, tool action, warning, escalation, approval, and material edit.
- Review high-risk sessions and a sample of ordinary sessions on a fixed cadence.
- Turn meaningful failures and corrections into labeled regression cases.
- Test proposed changes against the current production version before rollout.
- Require approval and a rollback path for changes to permissions, actions, policies, or source systems.
- Retire the workflow when nobody owns it or its sources can no longer be trusted.
The practical budget for an agent is therefore not only the launch cost. It includes review time, source maintenance, test maintenance, incident handling, and controlled improvement. A narrow agent with a healthy operating loop is more valuable than a broad demo that slowly becomes untrustworthy.
The same evidence makes the business easier to evaluate from the outside. Google says AI Overviews and AI Mode do not require special AI-only schema; helpful original content, crawlable text, accurate structured data, and visible support remain the foundation. Publishing a real workflow's job, boundaries, tests, changes, and measured results gives buyers and search systems something more substantial than an unsupported claim that a company is “AI powered.”
- Related: A Scheduled AI Task Still Needs An Exception Queue.
- Related: Measure AI By Accepted Work, Not Activity.
Sources Read
- Introducing OpenAI PresenceOpenAI
- A practical guide to building agentsOpenAI
- Optimizing your website for generative AI features on Google SearchGoogle Search Central
Next step, if this note maps to a problem on your desk: Discovery Call — a 1-on-1 leverage assessment for your business ($1,500 · 90 min).