The short answer
Your startup idea needs AI only if it helps deliver a valuable customer outcome better than a practical alternative, after accounting for review, errors, integration, and operating costs. Validate the problem and buying decision first; test AI’s contribution separately. NIST’s AI Risk Management Framework explicitly includes consideration of viable non-AI alternatives. (nvlpubs.nist.gov)
Use the framework below to decide what to investigate, what to test, and when to stop. The document-processing example is entirely hypothetical: its workflow, prices, and calculations illustrate a decision process, not actual customer findings or product performance.
1. Define the customer outcome without mentioning AI
Start with a statement that would remain useful if your preferred technology disappeared:
“For [specific customer], help complete [specific job] with less [cost, delay, effort, or risk], compared with [current approach].”
For a hypothetical document-processing startup:
“Help operations coordinators at small wholesale distributors turn incoming purchase orders into review-ready order records, with less manual retyping.”
That statement leaves several questions open. Are coordinators actually spending time retyping? Is retyping the bottleneck, or are missing information and approvals causing the delay? Does the buyer want faster processing, fewer corrections, or simply a better connection between existing systems?
Treat each answer as a hypothesis.
Steve Blank’s original customer-development guidance distinguishes testing a customer-problem hypothesis from collecting feature requests. The task is to determine whether the proposed solution addresses a need—not to assemble everything prospective customers say they would like. (steveblank.com)
Write down:
- User: Who performs the work?
- Buyer: Who can authorize spending?
- Trigger: What starts the workflow?
- Outcome: What does successful completion look like?
- Alternative: How is it handled today?
- Disproof: What finding would make you abandon this idea?
For the example, a useful disproof might be: “Most target businesses already import these orders automatically, and the remaining manual workload is too small to justify switching.”
2. Interview people about actual work, not your pitch
Customer discovery is a way to investigate market potential before committing to commercialization. NSF’s I-Corps program uses this process to help research teams assess their inventions beyond the laboratory. (nsf.gov)
For your own discovery, recruit people who perform or supervise the target workflow. Include businesses with different current approaches, including those satisfied with what they have. Interview users and buyers separately when their responsibilities differ.
Here is an original interview script for the hypothetical document-processing idea:
- “Walk me through the most recent purchase order you received.”
- “Where did it arrive, and what happened next?”
- “Which information did you copy, check, or ask someone to clarify?”
- “Can you show me a redacted example and the resulting order record?”
- “What went wrong or took longer than expected?”
- “What have you already tried to improve this process?”
- “How do you know when an order is ready to proceed?”
- “Who would approve a change to this workflow, and what would they need to see?”
Ask permission before observing work or recording conversations. Do not request confidential documents when a redacted example would answer the question.
Keep three columns in your notes:
| Observation | Interpretation | Still unknown |
|---|---|---|
| A coordinator retyped information during an observed task | Manual entry exists in this workflow | How frequently it occurs |
| A manager described a previous software purchase | There may be a purchasing route | Whether this project qualifies |
| A prospect praised the concept | The presentation was appealing | Whether they will adopt or pay |
Avoid opening with “Would an AI tool save you time?” Instead, investigate the work before presenting a solution. This keeps the conversation focused on testing the problem rather than gathering approval for your idea. (steveblank.com)

3. Map existing workarounds and test simpler alternatives
Create a workflow map from receipt to completion. Mark each handoff, correction, approval, and system update. Then identify the step your product would change.
Compare possible approaches rather than assuming a model belongs in the workflow:
| Approach to investigate | Question it should answer |
|---|---|
| A structured order form | Can better inputs remove the retyping problem? |
| An existing system’s import feature | Is the required capability already available? |
| Templates and explicit rules | Are the documents consistent enough for fixed logic? |
| A human-operated service | Will customers pay for the outcome before automation? |
| AI-assisted extraction | Does interpreting varied documents add enough value? |
These are candidate tests, not claims that any approach will work.
NIST recommends considering the resources needed to manage AI risk alongside viable non-AI systems, approaches, or methods. Apply that comparison to the whole workflow, not just the model’s output. (nvlpubs.nist.gov)
For the hypothetical startup, test whether suppliers would submit orders through a structured form. If that removes most manual entry, the business opportunity may be an integration product rather than an AI product.
Alternatively, if buyers cannot influence incoming formats, investigate extraction—but keep calculations, required-field checks, and approval rules explicit wherever possible.
Write down the strongest alternative and why a customer would choose it. “We use AI” is not a sufficient reason to switch.

4. Test willingness to pay with a concrete next step
Separate discovery from the sales test. First understand the workflow; then offer a clearly defined outcome.
For the hypothetical startup, propose a limited, human-reviewed pilot that returns draft order records. State:
- Which documents are included.
- Which fields will be returned.
- What remains the customer’s responsibility.
- Whether humans, automation, or both will perform the work.
- The proposed price and delivery arrangement.
- How documents will be handled and deleted.
Do not imply that a manually delivered service proves automated performance.
Use a practical evidence ladder:
- Interest: The prospect wants to learn more.
- Participation: They arrange a workflow walkthrough.
- Commitment: They obtain approval for a bounded pilot.
- Purchase: An authorized buyer accepts a defined paid offer.
- Continuation: They choose to keep using the outcome after trying it.
This ladder is a decision aid, not a universal ranking or statistical validation standard. Record exactly what happened. A paid pilot establishes a purchase of that pilot; it does not establish broad demand or a scalable business.
Ask the buyer what spending the offer would replace and what approval steps remain. If payment is blocked, investigate why. The obstacle might be weak value, procurement timing, data restrictions, or an unsuitable offer. Do not interpret every refusal as either total rejection or hidden enthusiasm.
5. Test whether AI improves the workflow reliably
Generative AI can produce confident but incorrect content. NIST’s Generative AI Profile identifies this risk as confabulation; plausible wording is not evidence of correctness. (nvlpubs.nist.gov)
Before testing, define the exact task. “Process documents” is too broad. “Extract specified fields into a draft record and flag missing information” is testable.
For the hypothetical pilot, prepare authorized examples with independently checked answers. Include different layouts, missing fields, ambiguous dates, duplicate orders, and poor scans if those conditions belong to the intended workflow.
NIST recommends comparing outputs with known ground truth and cautions against extrapolating performance from narrow, anecdotal assessments. (nvlpubs.nist.gov)
Use a practical evaluation sheet:
- Was each required field correct, missing, or unsupported?
- Could the reviewer locate the supporting information?
- How much work did checking and correction require?
- Were exceptions routed appropriately?
- Did a failed extraction remain a draft rather than trigger an action?
Distinguish a formatting mistake from an incorrect quantity or delivery address. Track critical errors separately.
Keep test documents separate from examples used to tune the system. Re-test material changes. For this pilot, require explicit approval before a draft order moves downstream.
Also check access, retention, deletion, and third-party handling before sharing documents. NIST identifies leakage and unauthorized disclosure or use of sensitive data as generative-AI risks. (nvlpubs.nist.gov)
6. Calculate value after review and switching costs
Consider this illustrative calculation, not a forecast.
Assume a prospective customer handles:
- 100 documents per month.
- Six minutes of manual work per document.
- Two minutes of review and correction per document with the proposed workflow.
- An estimated labor value of $30 per hour.
- A proposed monthly price of $150.
Estimated time released:
100 × (6 − 2) ÷ 60 = 6⅔ hours per month
Estimated value of that capacity:
6⅔ × $30 = $200 per month
After the proposed subscription, the customer has $50 of estimated monthly value before setup, training, and other switching costs.
Now change one assumption: review takes four minutes.
100 × (6 − 4) ÷ 60 × $30 = $100
At a $150 price, the offer no longer pays for itself on this labor-capacity calculation alone.
Do not label released time as cash savings unless spending actually decreases. Ask what the customer would do with that capacity.
Then calculate your own delivery economics separately: model usage, infrastructure, human support, exception handling, onboarding, and maintenance. Customer value and provider margin answer different questions.
Replace every assumption with observed evidence before using the calculation to justify development. If reduced mistakes or faster turnaround matter, investigate those benefits separately rather than assigning an invented dollar value.
7. Set evidence gates—and allow a non-AI “go”
Before the next experiment, write down the evidence needed to proceed and the findings that would stop it.
NIST’s framework supports an initial go/no-go decision informed by context, impacts, and risks. It also calls for defining business value and organizational risk tolerances. (nvlpubs.nist.gov)
For this hypothetical startup, use these gates:
| Gate | Evidence to seek | Reason to pause |
|---|---|---|
| Problem | Documented instances of consequential manual work | Minor inconvenience with no priority |
| Buyer | Identified approver and purchasing route | Enthusiasm without authority |
| Demand | Acceptance of a defined offer | Repeated praise without action |
| Alternative | Comparison with the strongest simpler option | Existing tools solve the need adequately |
| Feasibility | Workflow testing including review and exceptions | Checking removes the benefit |
| Data | Authorized access and acceptable handling arrangements | Documents cannot be used appropriately |
| Economics | Customer value and plausible delivery margin | Value or costs depend on unsupported assumptions |
Choose numerical thresholds for your own experiment before seeing results. Treat them as decision rules, not universal proof of product-market fit. A small discovery sample can justify another test without establishing market-wide demand.
Finish with one of four decisions:
- Proceed with AI: The task benefits from it, and evidence supports a bounded next step.
- Proceed without AI: The problem is valuable, but a simpler approach is stronger.
- Revise: Change the segment, workflow, offer, or scope.
- Stop: Current evidence does not justify more investment.
Your next step is not necessarily a prototype. It is the smallest honest experiment that resolves the most important uncertainty.