Contact Discuss an opportunity
Proof Engine cover card reading 'Define done before delegation' over overlapping dark rounded frames lit in red

The Acceptance Criteria Is The New Prompt

A familiar scene is starting to repeat inside product, GTM, and operations teams.

Someone wants to improve an AI workflow. The conversation turns quickly to prompts, model choice, agent frameworks, examples, system messages, and which tool should be connected next. The first output gets better. The summary becomes cleaner. The research brief has a sharper structure. The follow-up email sounds more polished. The CRM note looks more complete.

Then the real question appears: is the work actually done?

The team often cannot answer. The prompt improved the artifact, but the artifact still may not be acceptable for the workflow. A sales brief can be well written and still fail to change account prioritization. A support summary can be concise and still hide the unresolved customer risk. A customer feedback analysis can group comments into themes and still fail to support a product decision. A CRM update can look plausible and still distort the shared account memory.

Prompt quality matters. Model quality matters. Tooling matters. But work quality needs a different kind of standard.

Before teams write a better prompt, they need something older, less fashionable, and more operational: acceptance criteria, Definition of Done, and a shared view of what acceptable work looks like.

Where The Idea Came From

The concept did not come from AI. It came from the need to make work inspectable.

Product and software teams had to solve this problem long before LLMs entered the workflow. A team could not rely on someone saying a feature was “basically done” if the code had not been tested, documented, reviewed, or integrated. A story could not be accepted only because the implementation existed. The team needed a shared standard for what had to be true before the work could move forward.

Scrum formalized one version of this through the Definition of Done. In the Scrum Guide, the Definition of Done is the quality standard that makes it clear when an increment is complete enough to inspect and use. It creates transparency around the state of the work. Developers are accountable for creating increments that meet it. Without that shared definition, the word “done” becomes local opinion rather than an operating standard.

Acceptance criteria usually operate closer to the specific item of work. They describe what must be true for a feature, story, workflow, or output to be accepted. Acceptance testing grew from the same practical need: the team and stakeholder need a way to verify whether the work satisfies the expected behavior. Definition of Ready can appear upstream as well: what must be known before the team should begin.

The point is not to import Scrum ceremonies into AI work. Most teams do not need more ritual. The useful part is the operating discipline underneath it: before work is delegated, the team defines the state that would make the output trustworthy enough to use.

AI makes that discipline more important because AI produces fluent artifacts quickly. Fluency can create the illusion of completion. A summary reads well, so it feels complete. A brief has headings, so it feels structured. A CRM update mentions the right account, so it feels safe. The surface quality can improve faster than the underlying work quality.

That is the trap.

The New Failure Mode

The old failure mode for weak AI usage was a bad draft.

The new failure mode is a confident system acting from an output nobody defined, reviewed, or accepted properly.

This matters because AI workflows are moving closer to systems of record and customer-facing decisions. They read CRM, call transcripts, support tickets, product docs, financial records, and internal notes. They draft customer messages, classify risk, suggest prioritization, update fields, create tasks, and route exceptions. The workflow may still have a human in the loop, but the human often reviews the final text rather than the evidence behind the work.

That means a clean-looking output can quietly change the company’s operating memory.

If the CRM says an opportunity is qualified, someone will trust it. If a call summary says the buyer asked for pricing, someone may send pricing. If a product feedback cluster says enterprise customers want reporting, a product team may design around that assumption. If an AI assistant updates a field that feeds a forecast, the effect can spread beyond the person who approved the first output.

The practical implication: every serious AI workflow needs a definition of done before it gets meaningful autonomy.

This connects directly to the argument in Outcome-Priced AI Needs Outcome Receipts. If a vendor, internal team, or operator wants to price, evaluate, or trust AI based on outcomes, the system needs evidence that the outcome happened. Acceptance criteria sit one step earlier. They define what would count as acceptable work in the first place.

Without that standard, the receipt has nothing stable to prove.

What Acceptance Criteria Look Like For AI Workflows

For AI work, acceptance criteria should define more than the desired output format. They should describe the work contract around the output.

A useful template includes nine parts:

  1. Input
  2. Output
  3. Evidence
  4. Reviewer
  5. Destination
  6. Boundary
  7. Confidence
  8. Exception handling
  9. Decision impact

Input defines which sources the system may use. A revenue workflow might allow CRM records, call transcripts, meeting notes, product pages, approved case studies, account websites, and recent emails. It might exclude stale notes, unverified Slack comments, private documents, competitor rumors, or any source without a timestamp. If the input boundary is unclear, the output becomes hard to inspect.

Output defines what must be produced, updated, routed, or decided. A draft email is different from a sent email. A suggested CRM update is different from an applied field change. A customer feedback theme is different from a roadmap recommendation. Many teams describe the output as “summarize”, “research”, “classify”, or “write”. Those verbs are too loose for workflow design. The output should describe the usable artifact and its intended state.

Evidence defines what must accompany the output so a human can inspect it. This may include citations, source links, transcript timestamps, customer quotes, screenshots, structured records, prior field values, proposed new field values, confidence labels, or a list of missing sources. OpenAI’s own guidance on evidence of value points teams toward workflow records, usage data, timestamps, quality reviews, structured feedback, and direct observation. That same logic applies inside the workflow itself: the system should preserve enough evidence to make the work reviewable.

Reviewer defines who accepts the output. A sales manager may accept call review labels. RevOps may accept CRM field updates. A product owner may accept customer feedback clusters. Legal may accept policy drafts. A founder may accept an investor proof memo. Acceptance without ownership becomes theater. The reviewer does more than glance at the artifact. The reviewer is accountable for deciding whether the output can enter the next step of work.

Destination defines where accepted work lives. Does it become a CRM activity, account note, customer follow-up, product backlog item, support knowledge base article, weekly review artifact, decision memo, or audit log? AI work that lives only in chat is often hard to reuse. The destination matters because it determines whether the work becomes part of the operating system or disappears into a private conversation.

Boundary defines what the AI can suggest, draft, classify, update, or trigger. It also defines what requires explicit human approval. This is where teams prevent accidental escalation. A system may be allowed to draft a follow-up but not send it. It may suggest an opportunity stage change but not apply it. It may classify a support ticket but not close it. It may identify a compliance concern but only route it to a human reviewer.

Confidence defines how the system should express uncertainty. Confidence should not mean a vague percentage from the model. It should tell the reviewer what kind of evidence supports the output: directly stated, observed, inferred, estimated, contradicted, stale, missing, or unknown. A buyer saying “we need security review before procurement” is a different signal from an AI system inferring that security may be involved because the account is enterprise.

Exception handling defines what happens when the workflow cannot complete the work safely. Missing transcript? Conflicting CRM fields? Stale source? Policy exception? Low confidence? The system should route the case, ask for human input, mark the output incomplete, or refuse to take the next action. A mature workflow does not hide exceptions behind a smooth answer.

Decision impact defines what the accepted output is supposed to change. This is the part teams often skip. If the output will not change prioritization, routing, follow-up, forecast, product discovery, support escalation, renewal risk, or a customer decision, the workflow may be producing content rather than useful work.

The useful check is simple: when this output is accepted, what decision becomes easier, safer, faster, or more accurate?

Example 1: Account Research Agent

The weak version of an account research agent is easy to build. The team asks the agent to “research this account” and receives a polished company summary. It describes the business, market, leadership team, recent news, funding, and possible pain points. The output looks useful. It may even be accurate.

The workflow value is unclear.

If the account summary does not change prioritization, qualification, routing, outreach, or meeting preparation, it is a content artifact. It may help a rep feel prepared, but it has not entered the revenue workflow in a meaningful way.

A stronger version starts with acceptance criteria.

For example:

  • The agent must identify one current business trigger from an approved source.
  • It must state why that trigger may make the product relevant.
  • It must separate verified facts from inferred buyer hypotheses.
  • It must identify likely stakeholder roles and mark confidence for each.
  • It must cite every external source and include the date accessed.
  • It must recommend one outreach angle tied to the trigger.
  • It must identify one reason the account should be disqualified or deprioritized if present.
  • It must route low-confidence or missing-source cases to human review.

Now the definition of done is sharper. The brief is accepted only when it can support a decision: prioritize, disqualify, route, or prepare a reviewed first message. If the output cannot support one of those decisions, it is incomplete.

This is also where workflow readiness becomes visible. If the team cannot define what makes an account worth prioritizing, the AI system cannot solve that gap with a better prompt. OpenAI’s workflow readiness guidance starts from a real workflow problem and asks whether the work is frequent, painful, stable enough to improve, connected to systems, and ready for governance. Account research often looks ready because it is repetitive. In practice, it may be unready because the team’s ICP, routing logic, or disqualification criteria are weak.

The acceptance criteria expose that before the team overbuilds the agent.

Example 2: Customer Feedback Analysis

Customer feedback analysis is another common AI workflow. The weak version is a thematic summary.

The model groups feedback into categories: pricing, onboarding, integrations, reporting, performance, support, and usability. The result looks organized. The team can paste it into a product review.

The problem is that surface categories often mix different kinds of evidence.

“Reporting” could mean an executive buyer wants dashboards for internal visibility. It could mean a support user cannot export a CSV. It could mean an enterprise prospect needs compliance reporting before procurement. It could mean an existing feature exists but users cannot find it. It could mean the product has a permission model problem rather than a reporting problem.

A stronger workflow defines the acceptance criteria for each theme.

Each theme should include source count, customer segment, customer type, workflow stage, representative quote, evidence strength, affected job, revenue or retention relevance, and open uncertainty. It should distinguish feature request, workflow blocker, emotional complaint, implementation constraint, and underlying job. It should mark whether the evidence came from support tickets, sales calls, churn interviews, product analytics, customer interviews, or internal interpretation.

The output is done only when it can support product judgment.

That judgment might be a discovery question, a backlog decision, an experiment, a customer follow-up, a roadmap memo, or a decision to ignore a noisy theme. The AI system is not there to make the product decision. It is there to make the evidence more inspectable so the product owner can make a better decision.

This connects to Proof Engine’s broader methodology: evidence matters when it changes what the team does next. A feedback summary without acceptance criteria can make the team feel informed. A feedback artifact with source, segment, uncertainty, and decision impact can change what the team builds or validates.

Example 3: Sales Follow-Up Agent

Sales follow-up is where the problem becomes very concrete.

The weak version asks an AI assistant to write a follow-up email from a call transcript. The email is polite, concise, and on brand. It thanks the buyer, summarizes the conversation, and proposes a next meeting.

It can still be wrong.

The buyer may have asked for proof that the product works with their current CRM. They may have named a CFO concern. They may have said the next step depends on a sales manager. They may have expressed hesitation around implementation effort. They may have asked for a comparison against a current vendor. If the follow-up misses that proof request, the email can be well written and still weaken the deal.

Acceptance criteria for a sales follow-up agent should include:

  • buyer-stated problem;
  • agreed next step;
  • named stakeholder or missing stakeholder;
  • proof requested by the buyer;
  • unresolved objection or risk;
  • internal owner for the next action;
  • source timestamp from the call;
  • human approval before sending;
  • CRM destination for the accepted note and follow-up.

The follow-up is done when it preserves commitment and improves buyer movement. It is incomplete if it merely sounds good.

This is one reason the AI workflow question should start with the work rather than the model. The model can make the email sound better. The acceptance criteria make the workflow more commercially useful.

How This Changes The Build Process

When acceptance criteria exist first, the build process becomes more honest.

The team can see which data sources are missing. It can see whether a human reviewer has capacity. It can see which outputs are low risk and which touch systems of record. It can see whether the workflow has a real owner. It can see whether the organization has a clear view of what the AI system should improve.

It also prevents premature autonomy.

Some workflows are not ready for agents because the underlying process is unclear. Some workflows need better source systems before AI can act safely. Some workflows are valuable but should remain draft-only. Some workflows should start as a manual or semi-automated loop because the team has not yet learned what good output looks like.

This is the difference between building a demo agent and building a useful internal product.

A demo agent proves that something can be generated. A useful internal product proves that a real workflow can be improved, reviewed, adopted, and trusted. For teams deciding where AI belongs, the first useful artifact may be a workflow map and acceptance criteria, not an agent.

That is the center of Proof Engine’s AI Workflow / Internal Product Build work. The useful work is mapping the workflow, defining acceptable output, identifying the evidence trail, building the smallest useful system, and evaluating whether it changes real work.

The prompt is still part of the system. It is just downstream of the definition of done.

Before adding an agent, the sharper question is: what would make this work acceptable, inspectable, and useful enough to act on?

If the team cannot answer that, the next prompt will probably improve the artifact without improving the workflow.

Practical Close

For any AI workflow under consideration, write the acceptance criteria before writing the prompt.

Define the input. Define the output. Define the evidence. Define the reviewer. Define the destination. Define the boundary. Define confidence. Define exceptions. Define the decision the output should change.

Then ask whether the workflow is ready.

If your team is adding AI agents to a workflow, Proof Engine can help define the workflow, acceptance criteria, evidence trail, and review model before you automate the wrong version of the work.

Sources