Most teams start an AI project by choosing a tool. A building business gets a better result by starting one step earlier: name the exact deliverable the business needs. A quote comparison, client update, finish schedule, lead follow-up, and procurement recommendation may all use AI, but they require different sources, different room for judgment, and different approval rules.
Anthropic's June 2026 Economic Index classified the primary outputs of Claude conversations into more than 30 artifact types. It found that the nature of the output shapes how people work with AI: translating a document is largely constrained by the source text, while building a website leaves much more to the model's judgment. The study examines Claude usage rather than construction operations, but its operating lesson transfers cleanly. The less defined the artifact, the more hidden discretion the AI receives.
A task name is not a deliverable
“Help with estimating” does not tell a worker what finished means. Should the output be a coverage checklist, a normalized bid table, an exception report, a client allowance summary, or a recommendation ready for approval? Each artifact answers a different business question. If the team leaves that choice to the AI, it may produce the most fluent output rather than the one the next person can safely use.
The same problem appears in marketing, selections, purchasing, and project management. “Follow up with this lead” hides decisions about tone, qualification, scheduling, price claims, and whether the message may be sent. “Review this proposal” hides whether the team wants a summary, a scope-gap list, a risk flag, or an approval recommendation. Naming the artifact exposes those decisions before the run begins.
Write an artifact contract
- Purpose: the business decision or handoff this artifact must support.
- Shape: the required table, checklist, draft, comparison, exception report, or approval packet.
- Sources: the exact systems, files, revisions, and fields the AI may use.
- Required fields: the facts, citations, exclusions, uncertainties, and next action that cannot be omitted.
- Acceptance test: the mechanical checks and human judgment that determine whether the work is usable.
- Authority: who may review, approve, send, purchase, schedule, price, or update the system of record.
For a subcontractor comparison, the contract might require one row per scope item, source-file links, revision dates, alternates, exclusions, normalized totals, unresolved conflicts, and a named approval request. That is far safer than asking an agent to “pick the best bid.” It also creates a reusable interface between the AI worker and the estimator who owns the decision.
Match autonomy to the artifact
Not every deliverable needs the same control. A verbatim document translation can be checked against its source. A client-facing scope recommendation asks the AI to interpret evidence, weigh tradeoffs, and influence a commitment. As judgment and consequence rise, narrow the tool permissions, add evidence requirements, strengthen the evaluation cases, and move human approval earlier.
Anthropic's trustworthy-agent guidance makes the surrounding system explicit: agent behavior depends on the model, harness, tools, and environment, and users should consider what data, permissions, and operating environment the agent receives. OpenAI's production guidance similarly emphasizes shared business context, clear permissions and boundaries, and evaluations that teach the system what good work looks like. The practical unit to govern is therefore not “AI” in the abstract. It is one artifact-producing workflow.
Evaluate the deliverable your team will accept
Build evaluation cases from real rejected work. Test a missing drawing revision, conflicting quote totals, an unsupported product claim, a scope gap, an expired lead time, and a request to act without approval. Score whether the artifact contains every required field, cites the current source, exposes uncertainty, preserves business rules, and routes the next decision to the right person.
Then measure accepted work: first-pass acceptance, reviewer corrections, missing-field rate, unsupported claims, time to approval, and downstream rework. Token cost and model speed matter, but only after the artifact survives the real handoff. A cheaper draft that creates an hour of reconstruction is not a cheaper workflow.
Make the method visible
Google says AI Overviews and AI Mode do not require special AI-only schema. Helpful, original, technically accessible content and accurate structured data remain the foundation. For Datum, the useful public proof is the operating detail an answer summary cannot replace: the artifact contract, source rules, acceptance tests, approval boundary, and examples of what the team rejected and improved.
Continue with the operating system
- AI Training Needs A Job Packet, Not A Prompt List
- Measure AI By Accepted Work, Not Cheap Tokens
- Explore practical AI paths for your team
Sources Read
- Anthropic Economic Index report: CadencesAnthropic
- Trustworthy agents in practiceAnthropic
- Introducing OpenAI FrontierOpenAI
- Google's Guide to Optimizing for Generative AI Features on Google SearchGoogle Search Central
Next step, if this note maps to a problem on your desk: Private Training — a private working session for your team ($1,500+).