The short answer
Before expanding an AI deployment, measure whether it improves a defined business outcome after including human review, rework, and operating costs. Compare it with the existing workflow, set quality limits before testing, and track actual use. Expand only when the evidence supports both business value and acceptable risk—not simply faster output.
1. Define the decision your pilot must answer
Start with a decision, not a tool demonstration:
Should we expand AI-assisted drafting for routine support requests, keep testing it, or stop?
Then write a one-page pilot charter containing:
- Eligible work: the request types, languages, and channels included.
- Excluded work: cases that must remain outside the pilot.
- Primary outcome: one business measure, such as cost per resolved request.
- Quality requirements: factual accuracy, appropriate escalation, and resolution standards.
- Accountability: who reviews outputs and who can pause the pilot.
- Decision rules: what evidence permits expansion, requires revision, or triggers shutdown.
For a first support-drafting pilot, consider limiting AI to suggesting replies from approved documentation. Require an agent to check and send every response. Exclude sensitive disputes or actions requiring separate authorization.
Evaluate that complete arrangement—not just the model’s writing ability. NIST’s Generative AI Profile calls for evaluation under conditions resembling deployment and warns against extrapolating capabilities from narrow or anecdotal assessments. (nvlpubs.nist.gov)
2. Collect a baseline with consistent definitions
Before introducing AI, record how the existing workflow performs. Choose a collection period that includes normal variations in demand, staffing, and request complexity. Do not assume a particularly quiet week represents ordinary operations.
For support drafting, build a baseline covering:
| Measure | Definition to establish before testing |
|---|---|
| Handling time | Active staff time, including research, writing, review, corrections, and follow-up |
| Resolution rate | Resolved eligible requests divided by all eligible requests, using a fixed follow-up window |
| Factual-error rate | Audited sent replies containing at least one factual error divided by audited sent replies |
| Escalation rate | Eligible requests transferred for additional authority or expertise divided by eligible requests |
| Cost per resolved request | Allocated workflow operating costs divided by requests meeting the resolution definition |
Define “resolved” carefully. For example, your pilot could require completion without a related reopening during a specified follow-up window. Apply the same rule to both groups, and allow that window to finish before calculating results.
Keep active handling time separate from elapsed customer waiting time. Record case type, language, channel, and agent experience so you can inspect whether the groups handled comparable work.
NIST’s experimental-design guidance identifies operators and time of day as examples of factors that can affect measured results. Applied here, those are reasons to document workflow differences rather than attribute every change to AI. (itl.nist.gov)
3. Use a comparison group, not just before-and-after results
Where feasible, randomly assign eligible work or agents to:
- Control: the existing workflow.
- Pilot: the AI-assisted workflow with its required review process.
Random assignment is a standard experimental-design approach described in NIST’s statistical handbook. Blocking—forming comparable groups before randomizing—can help account for important differences such as agent experience or request category. (itl.nist.gov)
Choose the assignment unit deliberately. Assigning by agent may be practical if people would otherwise reuse AI-derived knowledge on control requests. If you randomize agents or teams, ask the analyst to account for that grouping; do not analyze every ticket as though it were an independent assignment.
Keep routing, staffing rules, documentation access, and quality auditing comparable. Log model, prompt, and knowledge-source versions. Separate training and troubleshooting from the main evaluation period.
Analyze all work assigned to the pilot group, including requests where an agent declines to use AI. Then examine AI-used cases separately as a diagnostic view. This avoids presenting enthusiastic users’ results as the expected effect of offering the tool to everyone.
Before launch, ask an analyst to estimate the sample needed to detect your minimum worthwhile improvement and assess quality limits. If volume is insufficient, label the result inconclusive rather than relaxing the rules afterward.
4. Set quality thresholds before looking at savings
Create a review rubric that distinguishes:
- Minor presentation issues: awkward wording or formatting.
- Material factual errors: incorrect instructions, policy statements, or account details.
- Critical incidents: sensitive-data disclosure or an unauthorized consequential action.
Evaluate both the initial AI draft and the reply actually sent. The first shows what reviewers must repair; the second shows what reaches customers. NIST specifically recommends checking sources and citations in generated outputs during testing and ongoing monitoring. (nvlpubs.nist.gov)
Audit a representative sample from both groups, using the same rubric. Where practical, hide group labels from reviewers. Have reviewers independently score some overlapping cases and reconcile disagreements before relying on the ratings.
Predefine your acceptable quality difference. You might require material-error performance to be no worse than control within an agreed margin, while making a critical incident an immediate pause trigger. Those are organization-specific decisions, not universal benchmarks.
Do not equate polished writing with correctness. A randomized experiment involving 758 consultants found benefits on tasks within the tested AI’s capabilities but worse correctness on a task outside them. The controlled consulting exercises do not establish what your support workflow will achieve; they illustrate why speed and quality need separate evaluation. (doi.org)
5. Include human-review costs and meaningful adoption metrics
Build two cost views:
- Operating economics: labor, AI fees, allocated licenses, additional auditing, maintenance, and incident handling.
- Total pilot investment: operating costs plus integration, training, evaluation setup, and management time.
For the operating scorecard, use:
Cost per resolved request = allocated operating costs for eligible work ÷ eligible requests meeting the resolution definition
Include the cost of unresolved requests in the numerator. Include escalation and follow-up work wherever it occurs. Avoid counting review time twice if it is already included in handling time.
Treat time savings as potential capacity, not automatic cash savings. Before promising a financial benefit, state how released time would be used: reducing overtime, absorbing demand, or reallocating staff.
Track adoption through observable workflow events:
- Drafts offered divided by eligible pilot requests.
- Drafts used divided by drafts offered.
- Drafts requiring substantial correction.
- Requests completed without AI.
- Reasons for rejecting suggestions.
- Use by agent experience, case type, and language.
A high acceptance rate is not a quality score. In Generative AI at Work, a study of 5,172 support agents found an average 15% increase in issues resolved per hour, with uneven effects across workers. Its staggered rollout analysis concerned one firm and one tool; it is not a transferable ROI estimate. (arxiv.org)
6. Work through a hypothetical support-drafting scorecard
Every figure below is an illustrative assumption, not a measured result or recommended benchmark.
Suppose control and pilot each receive 1,000 comparable eligible requests. All AI-assisted replies require agent review. Both groups complete the same resolution follow-up window.
| Metric | Control | AI pilot | Illustrative rule |
|---|---|---|---|
| Mean active handling time per assigned request | 12 minutes | 9.6 minutes | Diagnostic improvement |
| Resolved requests | 900/1,000 | 900/1,000 | No unacceptable deterioration |
| Sent replies with factual errors | 8/200 audited | 10/200 audited | No increase beyond agreed margin |
| Escalations | 100/1,000 | 110/1,000 | Maximum agreed increase: 1 percentage point |
| Critical incidents observed | 0 | 0 | Any critical incident pauses the pilot |
| Operating cost per resolved request | $6.78 | $5.67 | At least 10% lower, with quality gates met |
Assume a loaded labor rate of $30 per hour. The pilot’s 9.6 minutes includes 1.8 minutes of checking and editing. Both groups’ handling totals include escalation and follow-up labor.
Control calculation
- Labor: 1,000 × 12 minutes ÷ 60 × $30 = $6,000
- Additional quality auditing: $100
- Cost per resolved request: $6,100 ÷ 900 = $6.78
Pilot calculation
- Labor: 1,000 × 9.6 minutes ÷ 60 × $30 = $4,800
- AI fees and allocated licenses: $200
- Additional quality auditing: $100
- Cost per resolved request: $5,100 ÷ 900 = $5.67
The assumed operating-cost reduction is approximately 16.4%. But expansion is not yet justified.
Audited factual errors increased from 4% to 5%, and escalations rose from 10% to 11%. Ask an analyst to assess uncertainty against the predefined margins. The error samples alone do not establish a reliable quality difference, while observing zero critical incidents does not prove zero risk.
If setup cost another assumed $2,000, report that separately. The illustrative $1,000 operating-cost difference per comparable batch would offset setup after two such batches only if the benefit persisted and could be realized economically.
7. Specify pause, stop, and expansion conditions
Write the response plan before the pilot starts.
Pause immediately for a critical incident, failed review safeguard, unauthorized data exposure, or missing logs that prevent meaningful assessment.
Revise and retest when review effort removes the expected benefit, adoption remains low, results are too uncertain, or a particular request category performs poorly.
Stop the approach when repeated revisions do not deliver worthwhile economics or the workflow cannot operate within your risk limits.
Expand gradually only when the primary outcome meets the business threshold, quality requirements are supported by evidence, and ongoing ownership is clear.
Name the person authorized to disable the feature and document how the old workflow resumes. Retain a review checkpoint for each expansion into a new language, channel, or request type.
This approach aligns with NIST’s AI RMF provisions for post-deployment monitoring, override, incident response, recovery, change management, and decommissioning. (nvlpubs.nist.gov)
8. Check for mistakes before approving rollout
Use this final review to challenge an attractive headline:
- Was the comparison fair? Check case mix, staffing, routing, and documentation changes.
- Was the whole workflow measured? Include checking, corrections, escalation, and follow-up.
- Were all assigned requests included? Do not quietly remove rejected drafts or failures.
- Was quality assessed independently? Separate factual correctness from tone and fluency.
- Did the follow-up window finish? Do not count recently closed requests as durable resolutions.
- Were results segmented? Inspect the languages and request categories you intend to expand.
- Are the benefits actionable? Explain how saved capacity becomes business value.
- Can the workflow be paused? Confirm the owner, fallback, and monitoring schedule.
Prepare a short decision memo showing the outcome difference, uncertainty, quality findings, adoption, costs, and unresolved questions. State exactly where the evidence applies.
The goal is not to prove that AI works in general. It is to establish whether this AI-assisted workflow, for this work, with these safeguards, earns the next stage of investment.