Contact Discuss an opportunity
Proof Engine cover card reading 'Not every win proves the same thing' over black horizontal bands with glowing amber waveform edges

Case Studies Need Evidence Levels

Two companies can say that an idea was validated while describing completely different realities.

One team interviewed twelve potential users and heard a recurring problem. Another placed a prototype in a live workflow and observed repeated use. A third converted a pilot into a paid deployment. A fourth retained and expanded several production customers.

All four may use the word validated.

The problem is not that one kind of evidence is legitimate and the others are not. Each can be appropriate for a different decision. The problem is that the label hides the distance between them.

When a buyer, founder, or investor reads a case study, they need to know what actually happened, under which conditions, and what the evidence is strong enough to support. A case based on interviews may justify building a prototype. It does not justify a claim of repeatable revenue. A strong pilot may justify a paid rollout for one workflow. It does not automatically prove that the same result will scale across markets.

Case studies need evidence levels because decision quality depends on understanding both what has been learned and what remains uncertain.

A Level Is Not A Grade

An evidence level should not become another vanity badge.

Early-stage evidence is not weak simply because it is early. If the decision is whether to spend two weeks prototyping a workflow, discovery evidence may be enough. Requiring production data before any experiment would make learning impossible. The standard should be proportional to the commitment being considered.

Nesta’s Standards of Evidence make a similar broader point in the context of social innovation: evidence standards should increase as an intervention matures and larger commitments depend on the result. Its framework moves from a coherent theory of change toward measured impact, stronger causal confidence, independent replication, and systems that support consistent replication.

The model proposed here is different and narrower. It is designed for product, GTM, AI workflow, and commercial case studies. It does not claim causal impact in the scientific sense. It gives readers a practical way to distinguish the kinds of evidence a business project has produced.

The purpose is not to rank teams. It is to prevent the evidence from being asked to support a decision larger than it can carry.

Level 1: Discovery Evidence

Discovery evidence clarifies the problem and the assumptions around it.

It may come from customer interviews, workflow mapping, field observation, desk research, expert conversations, support records, call reviews, or analysis of an existing process. At this level, the team can describe who appears to have the problem, how the work happens today, where friction occurs, and why a proposed change might matter.

Good discovery evidence is concrete. It preserves the buyer’s language, identifies recurring moments, distinguishes roles, and exposes contradictions. It may show that a problem is frequent, that a workflow has an owner, or that current alternatives are unsatisfactory.

What it can support: a sharper hypothesis, narrower segment, better problem definition, prototype decision, interview plan, or test design.

What it cannot support by itself: demonstrated adoption, willingness to pay, operational reliability, causal impact, or repeatable commercial demand.

The distinction is important because positive interviews are easy to overread. People can accurately describe a pain and still decline to change their behavior, expose their data, involve a manager, run a pilot, or pay.

Level 2: Demand And Intent Evidence

Demand and intent evidence appears when a potential buyer does something that costs more than expressing an opinion.

Examples include replying to a targeted offer, applying for access, introducing another stakeholder, sharing a real workflow, providing sample data, requesting security review, agreeing to a pilot, signing a Letter of Intent, or discussing a credible budget path.

The strength of the signal depends on the cost and specificity of the action. An email signup is lighter than a scheduled workflow session. A friendly introduction is lighter than access to operating data. An unsigned statement of interest is lighter than an LOI that identifies the proposed scope, decision date, and paid-conversion conditions.

Strategyzer’s testing guidance emphasizes the difference between what customers say and what they do. That distinction is useful here. Stated interest can guide the next test. Commitment behavior begins to reveal whether the opportunity can move through a real buying process.

What this level can support: prioritizing a segment, investing in a prototype or pilot, refining an offer, and mapping the buying path.

What it cannot support by itself: repeated usage, workflow fit, realized value, retention, or scalable economics.

Level 3: Behavioral Evidence

Behavioral evidence comes from people interacting with an artifact under conditions that approximate the intended job.

The artifact might be a prototype, concierge service, MVP, AI workflow, application flow, demo environment, or manual version of the future product. The team observes what users complete, repeat, abandon, correct, accept, or route for review.

Useful measures include completion, repeated runs, time on task, accepted outputs, setup friction, reviewer corrections, exception patterns, data-sharing behavior, and whether the artifact changes the next step in the workflow.

Behavioral evidence is stronger than preference because the user must perform work. It can reveal that a feature described as essential is ignored, that a supposedly simple setup creates friction, or that a human-review step is necessary for trust.

What it can support: improving the workflow, defining acceptance criteria, selecting the right level of automation, preparing a controlled pilot, or abandoning a poor interaction model.

What it cannot support by itself: reliable operation at production scale, organizational adoption, paid continuation, retention, or robust unit economics.

Level 4: Operational Pilot Evidence

Operational pilot evidence appears when the system is used inside a real process with a defined owner, baseline, review cadence, and decision rule.

The distinction from a product test is organizational. The workflow has consequences. Real users interact with real or appropriately controlled data. Exceptions must be handled. The system has to coexist with current tools and management routines. Someone is accountable for deciding whether the output is acceptable.

At this level, evidence may include measured workflow time, accepted-output rate, error and exception rate, follow-up speed, CRM completeness, usage by intended roles, manager review effort, risk events, and a business outcome connected to the pilot hypothesis.

The pilot should also record controls: source provenance, consent, security, human approval, escalation, reversibility, and stop conditions. A result is not operationally credible if the team cannot explain how it was produced or how an unsafe output would be caught.

What it can support: paid conversion for the tested scope, controlled expansion, production hardening, or a decision to stop.

What it does not automatically support: generalization across teams, markets, data environments, or implementation partners. One successful pilot remains evidence under particular conditions.

Level 5: Production And Commercial Evidence

Production and commercial evidence combines durable use with economic commitment.

Signals include paid continuation, renewal, retention, expansion, repeated implementation, stable operational performance, and credible economics. At this level, the product has survived beyond the novelty of the pilot and become part of ongoing work.

The strongest evidence is not merely that customers paid once. It is that the value can be reproduced without exceptional founder effort, uncontrolled service labor, or hidden risk. Implementation time, support burden, gross margin, adoption, and the conditions required for success all matter.

What this level can support: stronger scale decisions, investment in repeatable acquisition, broader hiring, standardization, or expansion into adjacent segments.

Even here, claims need boundaries. Production proof in one geography or workflow may not transfer automatically to another. Commercial evidence can decay as the product, market, buyer, or operating environment changes.

Evidence Has More Than One Dimension

A single badge can still mislead.

A product may have strong behavioral evidence and weak commercial evidence. A workflow may operate reliably but lack a repeatable acquisition path. A buyer may sign an LOI while the technical integration remains unproven. A company may have paid customers but little evidence that the product caused the desired outcome.

That is why a useful case should name at least the relevant dimensions:

  • problem evidence;
  • demand and commitment evidence;
  • behavioral or product evidence;
  • operational evidence;
  • commercial evidence.

The case does not need a complicated scorecard on the page. It needs precise language. “Users completed eight prototype runs with two human-review points” is more informative than “the AI workflow was validated.” “A production workflow is in buildout” is different from “deployed in production.” “The team observed willingness to share partial revenue data” is different from “customers purchased financing.”

Precision increases trust because the reader can match the evidence to their own decision.

Applying The Levels To Proof Engine Cases

The current Proof Engine case-study portfolio contains projects at different stages and with different public-evidence boundaries. That makes it a useful place to apply this framework.

The objective is not to relabel the cases as winners and losers. It is to clarify what each one demonstrates.

AI Sales Assistant

The AI Sales Assistant case tested an inbound qualification workflow and compared a shorter flow with a longer, more consultative one. The shorter flow performed better in the validation context, and the work surfaced trust requirements around privacy, brand tone, and qualification transparency.

The public case supports discovery and behavioral evidence around the qualification flow. It also supports a product decision: narrow the proposition from broad AI sales automation toward faster response, qualification, and useful handoff to a human seller.

It should not be presented as proof of repeatable production revenue or autonomous sales replacement.

Creator Capital

The Creator Capital case explored a financing opportunity for creators. The work helped narrow the broad creator-economy story toward creators with recurring or semi-recurring income, while testing trust, willingness to share financial information, offer comprehension, and capital-side constraints.

The public evidence sits primarily across discovery and demand/intent. It supports a sharper segment and financing wedge. It does not by itself demonstrate funded transactions, portfolio performance, or repeatable commercial economics.

This limitation does not make the work less useful. It shows which commitment the evidence was designed to support: a more focused product and validation path.

AI Workflow Automation

The AI Workflow Automation case began by scoring several workflow candidates and selecting one with visible frequency, manual effort, data availability, buyer urgency, and feasibility. The prototype used two human-review checkpoints.

Public evidence includes prototype runs, completion with review, setup friction, estimated manual-time reduction, and stronger response to specific outcome language than to a broad category claim about AI agents.

This is behavioral evidence with estimated value signals. It helps define the workflow and the role of human review. It should not be converted into a claim that the system has already demonstrated production-scale savings.

The Legal Operations AI case is a public-safe, anonymized account of a controlled workflow around notifications, deadlines, routine drafts, and human review. The prototype has been implemented, while the production workflow is described separately as being in buildout.

That wording is an evidence boundary. Prototype implementation supports behavioral learning about workflow design, required controls, and review. Production buildout describes current work, not completed operational proof.

The case becomes more credible when those states remain distinct.

Fix-And-Flip Marketplace

The fix-and-flip marketplace case is intentionally public-safe and centered on validation architecture. It describes risky assumptions, both sides of a capital marketplace, manual pilot design, unit economics, legal feasibility, and decision gates.

The strongest public evidence is not traction. It is a coherent discovery and test system for a complex opportunity where premature software could hide unresolved service, legal, underwriting, and trust problems.

Calling that distinction out protects the reader from inferring a market result that the public case does not claim.

Add Proof Markers To Every Claim

An evidence level is useful, but each important claim should still carry proof markers.

At minimum, ask:

  • Source: Where did the signal come from?
  • Date: When was it observed?
  • Sample: How many users, accounts, runs, or records were involved?
  • Environment: Interview, prototype, controlled pilot, or production?
  • Reviewer: Who checked or accepted the result?
  • Status: Measured, observed, reported, estimated, planned, or unknown?
  • Limitation: What can this evidence not establish?
  • Decision impact: What changed because of it?

OpenAI Academy’s evidence-of-value guidance uses a similar status discipline, distinguishing measured, observed, reported, estimated, planned, and unknown evidence. That vocabulary is especially helpful for AI case studies, where a polished prototype can make an estimate feel like an achieved outcome.

The markers do not all need to appear in a dense table. They should be available in the case or supporting material, and the prose should never erase them.

Better Case Studies Produce Better Decisions

The standard question for a case study is, “Does this make the company look credible?”

A better question is, “What decision can a reader responsibly make from this evidence?”

If a case shows discovery evidence, it can demonstrate how the team understands an uncertain market. If it shows behavioral evidence, it can demonstrate how assumptions were tested through use. If it shows operational pilot evidence, it can demonstrate how the system performed under real constraints. If it shows production and commercial evidence, it can support a stronger claim about repeatability.

At Proof Engine, the goal is not to make every case sound equally mature. It is to show how the right evidence changed the next decision. That is the philosophy behind Meet The Proof Engine Case Studies and the wider Proof Engine methodology.

The next time you read or write a case, ask two questions:

What does this evidence allow us to believe?

What does it still not prove?

The credibility of the case lives in the distance between those answers.

FAQ

What are evidence levels in a case study?

Five levels that separate what has actually been shown: discovery evidence, demand and intent evidence, behavioural evidence, operational pilot evidence, and production and commercial evidence. Each supports a different size of decision, which is what the word “validated” usually hides.

Does a lower evidence level mean weaker work?

No. The standard should be proportional to the commitment being considered. Discovery evidence is enough to justify a two-week prototype; requiring production data before any experiment would make learning impossible. The failure is asking evidence to support a decision larger than it can carry.

What proof markers should a claim carry?

Source, date, sample, environment, reviewer, status, limitation and decision impact. Status matters most: a measured result, an observed sample, a reported benefit and an estimate are not the same thing, and a polished prototype makes an estimate feel like an achieved outcome.

What can a successful pilot prove, and what can it not?

It can support paid conversion for the tested scope, controlled expansion, production hardening or a decision to stop. It does not automatically generalise across teams, markets, data environments or implementation partners. One pilot remains evidence under particular conditions.

Sources And Continuation Paths