An automation headline can be numerically true and still be operationally misleading. “Four hours became 40 minutes” may omit the number of tasks, the people who checked the output, the failures that were quietly corrected, the setup work, and the meetings that absorbed the released time. “Inquiries fell from 12 to three” may be good news, or it may mean customers stopped asking, cases went unanswered, or the measurement window changed.
The right question is not “Is the headline impressive?” It is:
What evidence would let another person reconstruct the comparison, find its blind spots, and decide whether the result is likely to transfer?
This guide provides that audit. It requires a baseline window, comparable task volume, labor minutes, rework, error severity, unanswered cases, quality sampling, operator training, software cost, implementation time, maintenance, and an account of what happened to the saved capacity. It also separates observed association from causal proof.
The official guidance and vendor documentation cited here were checked on 2026-07-24. Product behavior and documentation can change. Recheck the official pages before using this article for procurement, policy, or deployment.
Start by downgrading the headline to a testable claim
Source material supplied for this article reports several striking before-and-after figures:
- processing 80 notices fell from four hours to 40 minutes;
- inquiries fell from 12 to three;
- one activity fell from 30 minutes to 90 seconds; and
- time released by automation can return as additional meetings.
Those figures are useful leads, not verified performance results. The source material does not, by those numbers alone, establish the baseline dates, task mix, denominators, sampling method, error definitions, review labor, implementation effort, software cost, or counterfactual. The figures should therefore be disclosed as self-reported observations that require substantiation, not as proof that AI caused an improvement.
Turn each headline into a claim record before evaluating it.
| Headline fragment | Minimum testable translation | Missing evidence that could reverse the conclusion |
|---|---|---|
| “80 notices: four hours to 40 minutes” | For 80 notices of a defined type, measured under comparable conditions, total elapsed or labor time changed from 240 to 40 minutes | Whether the metric is elapsed time or hands-on labor; notice complexity; review and correction time; failed notices; operator experience |
| “Inquiries: 12 to three” | In the same-length before and after windows, a named inquiry category changed from 12 to 3, with intake volume and unanswered cases reported | Whether fewer inquiries mean fewer defects, less demand, missing logging, suppressed reporting, or changed classification |
| “30 minutes to 90 seconds” | A specified task’s median, mean, or percentile duration changed from 1,800 to 90 seconds for a disclosed number of cases | Distribution, queue time, retries, human review, exceptions, quality, cherry-picked best runs |
| “Saved time returned as meetings” | Capacity released by the workflow was traced to named uses and measured rather than assumed | Whether meetings created value, shifted work, increased coordination load, or erased the intended benefit |
A credible case study states which interpretation applies. If it cannot, the correct grade is “unsubstantiated,” not “false.” Absence of evidence and evidence of absence are different.
Define the unit before measuring the improvement
An audit begins with five definitions:
- Unit of work: one notice, inquiry, lead, invoice, support case, document, or completed customer outcome.
- Start event: the timestamp that begins the clock, such as valid intake received.
- End event: the event that ends the clock, such as approved notice delivered, not draft generated.
- Eligible population: every item that should enter the workflow, including exceptions.
- Successful outcome: a completed item meeting defined quality and safety requirements.
Without those definitions, teams often compare manual completion with automated draft generation. That is not a like-for-like comparison. The after measurement must include validation, correction, escalation, recovery, and downstream handoff when those steps were part of the original task.
The U.S. Bureau of Labor Statistics defines labor productivity at a broad economic level as output relative to hours worked. A single company case study is not a BLS productivity statistic, but the discipline is still useful: measure output and labor inputs separately rather than calling reduced elapsed time “productivity” by itself (BLS calculation method).
Keep four clocks separate
| Clock | What it measures | Example start and end | Why it matters |
|---|---|---|---|
| Elapsed time | Calendar time experienced by the requester | Valid intake to accepted completion | Includes waiting and queueing; may improve without reducing labor |
| Hands-on labor | Minutes people actively work on the item | Reading, correcting, approving, recovering | Required for capacity and cost claims |
| Machine runtime | Time the automation is executing or waiting | Trigger to final machine step | Useful for reliability and infrastructure analysis |
| Time to acceptable output | Time until the result passes the defined quality gate | Intake to approved, usable result | Prevents draft speed from being presented as completion speed |
Report at least a median and a tail statistic such as the 90th or 95th percentile when timing varies. A mean alone can hide a small number of very slow or failed cases. Preserve the count of timed observations and the rules for excluded observations.
Require a real baseline window
A screenshot of one slow manual run is not a baseline. Use a window long enough to include ordinary variation: different operators, busy and quiet periods, common input types, and known exceptions. The correct duration depends on volume and seasonality, so there is no universal number of days.
The baseline record should state:
- exact start and end dates;
- business days, shifts, or operating hours covered;
- incoming, eligible, attempted, completed, and unanswered counts;
- operators included and their training or tenure;
- task categories and complexity mix;
- source systems and workflow version;
- outages, campaigns, policy changes, staffing changes, and seasonality;
- timing collection method;
- quality sampling method;
- all exclusions, with counts and reasons.
Then freeze the comparison rule before reading the after results. If the after period has twice the volume, easier cases, or a newly trained team, normalize or stratify the data rather than comparing raw totals.
NIST’s AI Risk Management Framework calls for documented test sets, metrics, tools, deployment-like conditions, production monitoring, and defined human oversight. It also notes the value of independent review (NIST AI RMF Core). That does not certify any automation claim; it supports a traceable measurement design.
Measure every item entering and leaving the funnel
The case-study funnel should reconcile. For each window:
incoming items
- ineligible items
= eligible items
eligible items
= completed items + unanswered items + open items at cutoff
attempted automated items
= straight-through completions + human-reviewed completions
+ escalations + failed items + open items at cutoff
Counts may overlap only where the schema explicitly says so. For example, a completed item can also be a reworked item, but it cannot simultaneously be an unanswered item. Define late-arriving completions and the cutoff rule.
Unanswered cases belong in the denominator. Removing them makes a fast workflow look better precisely when it fails to serve people. If inquiries fell from 12 to three, audit:
- the number of customers or transactions that could have generated an inquiry;
- the inquiry categories and classification rule;
- logging coverage before and after;
- unresolved and abandoned contacts;
- response time and resolution quality;
- product, staffing, or policy changes that could also explain the decline.
Fewer inquiries can be an outcome metric only when the case study shows that demand and observability did not collapse.
Count labor minutes, not just button time
For each item, include all human work caused by the process:
total labor minutes
= intake and preparation
+ first-pass processing
+ human review and approval
+ correction and rework
+ exception handling
+ incident investigation
+ customer recovery
+ reporting and administration
+ allocated maintenance
Keep implementation and initial training separate from recurring operation, then show both. Mixing them hides payback; excluding them hides total investment.
Useful formulas:
labor minutes per incoming item
= total recurring labor minutes / incoming items
labor minutes per accepted item
= total recurring labor minutes / accepted completed items
rework rate
= items requiring correction / eligible items
unanswered rate
= unanswered items / eligible items
straight-through rate
= accepted items requiring no human touch / eligible items
gross labor minutes released
= baseline labor minutes at after-period volume - observed after labor minutes
net operating value
= gross labor hours released x loaded labor cost per hour
- incremental recurring software and infrastructure cost
simple payback windows
= one-time implementation and training cost / net operating value per comparison window
“Released” is more accurate than “saved” until the organization shows what happened next. A person may spend fewer minutes on one process while total hours, overtime, or backlog remain unchanged.
Rework and error severity are first-class outcomes
Automation can reduce average handling time while increasing expensive mistakes. Record:
- items requiring any correction;
- correction minutes;
- repeat corrections;
- duplicate or missing outputs;
- false accepts and false rejects where applicable;
- customer-visible failures;
- financial, privacy, safety, or authorization incidents;
- time to detect and time to recover;
- downstream work created outside the automated team.
Do not collapse every error into one percentage. Use a severity model defined before grading.
| Severity | Operational definition | Reporting rule |
|---|---|---|
| Critical | Unauthorized action, material safety or privacy impact, irreversible harm, or another predefined stop condition | Report every case separately; never average it away |
| Major | Wrong outcome requiring substantial correction, customer recovery, missed deadline, or material downstream work | Report count, rate, cause, and recovery labor |
| Minor | Local defect corrected without changing the material outcome | Report count and correction labor |
| Cosmetic | Formatting or style defect with no operational consequence | Keep separate so cosmetic volume does not dilute serious failures |
Severity definitions must fit the domain and be reviewed by people who understand the work. A typo in an internal draft and a wrong payment instruction are not equivalent errors.
The machine-gate guide explains how deterministic checks can enforce schemas and prohibited actions. The human-review coverage guide helps allocate review to risk rather than claiming that every human glance provides the same protection.
Sample quality independently
Case studies often report only items that users complained about. That misses silent errors. Draw a quality sample from the full eligible population, including completed, unanswered, escalated, and apparently successful items.
Record:
- sampling frame and extraction query;
- random, stratified, or risk-based method;
- sample size and date;
- strata such as complexity, language, operator, customer group, and exception type;
- blinded or unblinded review;
- rubric and severity definitions;
- reviewer qualifications;
- disagreement and adjudication process;
- automation and reviewer versions;
- confidence intervals or an explicit statement that the sample is descriptive only.
Use risk-based oversampling to find rare severe failures, but weight or report strata separately. Do not present an intentionally enriched risk sample as the overall error rate.
For AI-specific evaluation design, see how to build an LLM evaluation dataset and the self-hosted evaluation gate. A case study is stronger when ordinary operational metrics and a versioned failure-focused evaluation suite agree.
Preserve audit logs that can reconstruct the claim
An audit log should connect each item to:
- stable item and run identifiers;
- timestamps for intake, machine steps, review, completion, and recovery;
- workflow, prompt, model, connector, and policy versions where applicable;
- input category and complexity flags without unnecessary personal data;
- machine outputs and validation results;
- human reviewer decision and material edits;
- retry, timeout, exception, and fallback events;
- final disposition;
- cost and usage records;
- access and change events for the automation itself.
Logs are evidence only if their coverage, retention, and access are known. A dashboard may omit deleted, timed-out, filtered, or pre-trigger items. Reconcile a sample against the source system.
Official vendor documentation shows why a buyer must inspect the actual logging boundary rather than assume that “history” means a complete audit trail. Microsoft documents Power Automate activity logs and per-run fields such as status, duration, error codes, and trigger type (Microsoft activity logging). UiPath documents centralized audit-log views and exports for platform activity (UiPath Automation Cloud logs). Zapier says Zap History records workflow runs, versions, statuses, and task usage, while also documenting storage and display limits (Zapier history).
Those pages describe product capabilities, not proof that a particular case study enabled, retained, exported, or reconciled the logs.
The AI agent observability guide covers trace completeness, retention, and exportability. The production LLM API operations guide adds version, cost, retry, and incident evidence for model-backed workflows.
Design human review as a control, not a slogan
“Human in the loop” is incomplete disclosure. State:
- which cases require review;
- whether review occurs before or after an external effect;
- what information the reviewer sees;
- the review rubric;
- reviewer authority to reject, edit, escalate, or stop;
- target review time and actual review labor;
- queue coverage outside business hours;
- training and calibration;
- override rate and override outcomes;
- what happens when the queue is unavailable.
UiPath’s official Action Center documentation describes workflows that suspend and resume after human input when human validation is required (UiPath Action Center). That is a mechanism. Its effectiveness still depends on the rule, reviewer, evidence, authority, and queue.
Use the human-in-the-loop design guide to define intervention points. Use reviewing AI-generated output as a parallel example of why a named reviewer and an explicit acceptance boundary matter.
Require an error budget and recovery path
A robust automation case study explains what happens when the happy path breaks:
- detect the failure;
- stop or contain unsafe continuation;
- classify retryable versus non-retryable errors;
- retry with bounded attempts and idempotency where appropriate;
- route unresolved items to a visible queue;
- alert a named owner;
- preserve evidence;
- recover the customer or downstream system;
- correct the root cause;
- add the failure to regression evaluation.
Microsoft’s Power Automate guidance recommends grouping actions for error handling, logging errors, using retry policies for transient failures, and monitoring performance (Microsoft error-handling guidance). Zapier documents custom error-handler paths and their history status (Zapier custom error handling). These are implementation options, not evidence that errors were harmless.
Report retry counts, duplicate side effects, recovery labor, and cases that never reached the automation. The timeout, retry, and idempotency guide explains why replaying an uncertain action can create a duplicate rather than repair a failure. The token-budget and circuit-breaker guide provides a pattern for bounding automated work.
Account for training, implementation, cost, and maintenance
A case study should contain two ledgers.
One-time ledger
- discovery and process mapping;
- data cleanup and migration;
- implementation and integration;
- security, privacy, and legal review;
- test-data creation and evaluation;
- operator and reviewer training;
- change management;
- parallel running;
- launch support;
- vendor or consultant fees.
Recurring ledger
- licenses and usage;
- model, API, compute, storage, and logging;
- monitoring and on-call coverage;
- quality sampling;
- exception handling;
- workflow maintenance;
- prompt, rule, and connector updates;
- retraining and reviewer calibration;
- incident and customer recovery;
- audit and compliance work.
State whether software cost is list price, contracted price, allocated platform cost, or an estimate. Include taxes and implementation fees only when the case study can substantiate them. Do not convert labor minutes into money without a disclosed loaded labor rate and currency.
Training can improve both manual and automated performance. If operators received new training only in the after period, AI is not the only plausible cause. OECD workplace research emphasizes training and worker consultation as relevant to outcomes and warns about workplace risks alongside potential benefits (OECD AI and work). Treat those factors as measured conditions, not background decoration.
Trace what happened to the saved capacity
Released capacity can produce several different outcomes:
| Destination of released time | Evidence to request | Interpretation |
|---|---|---|
| More completed work | Throughput, backlog age, service level, and quality | Capacity became output if demand and quality stayed comparable |
| Better customer work | Contact, resolution, satisfaction, or retention measures with caveats | May create value, but causal attribution still needs a design |
| Training or improvement | Attendance, completed practice, resulting competency evidence | Investment in capability, not immediate cost savings |
| Shorter hours or less overtime | Paid hours, overtime, schedule, and staffing records | Stronger evidence of labor reduction |
| Meetings | Calendar categories, duration, attendees, decisions, and follow-up work | Could coordinate value or recreate overhead; measure rather than assume |
| Idle or fragmented time | Work sampling, queue state, and operator interviews | Capacity was released but not necessarily usable |
| Headcount avoidance | Approved staffing plan, workload forecast, vacancies, and service outcomes | A counterfactual claim requiring unusually careful evidence |
Do not label capacity as financial savings unless cash spending actually fell or a defensible counterfactual was avoided. “People had more time” is a capacity claim. “The company saved money” is a financial claim. They need different evidence.
The warning that saved time can return as meetings should be part of the result, not a footnote. Measure calendar load before and after, but do not assume all meetings are waste. Record meeting purpose and whether decisions, escalations, or customer outcomes improved.
A complete worked example with fictional data
The following dataset is entirely fictional. It is an arithmetic example, not an observed deployment, a benchmark, or a fabricated test result. Its purpose is to show the disclosures and calculations a real case study should provide.
Fictional scenario and measurement contract
A six-person operations team processes service notices. The baseline and after windows each contain 20 business days. An eligible item begins when a valid notice enters the case system and ends when the approved notice is delivered or the window closes. The after workflow drafts and validates notices, but a human reviews defined risk categories. All incoming items remain in the denominator.
| Metric | Baseline window | After window | Comparison note |
|---|---|---|---|
| Business days | 20 | 20 | Same duration |
| Incoming and eligible notices | 800 | 820 | After volume is 2.5% higher |
| Accepted completions by cutoff | 760 | 803 | Outcome, not draft count |
| Unanswered at cutoff | 40 | 17 | Retained in denominator |
| Items requiring rework | 96 | 49 | Any material correction |
| First-pass labor minutes | 4,800 | 1,230 | Includes preparation and review in the defined step |
| Rework labor minutes | 480 | 294 | More minutes per reworked after item |
| Incident and recovery minutes | 240 | 175 | Includes customer recovery |
| Recurring maintenance minutes | 0 | 240 | Included only after launch |
| Total recurring labor minutes | 5,520 | 1,939 | Sum of four labor rows |
| Existing monthly software cost | $180 | $180 | Fictional allocated cost |
| New monthly software cost | $0 | $900 | Fictional incremental cost |
| One-time implementation labor | 0 | 72 hours | Kept outside recurring operation |
| One-time operator training | 0 | 12 hours | Six people at two hours each |
| One-time external implementation fee | $0 | $4,500 | Fictional fee |
Fictional quality sample
An independent reviewer draws 200 items from each full eligible population using a documented random query. These counts are descriptive; no confidence interval or causal claim is supplied.
| Sample result | Baseline sample, n=200 | After sample, n=200 | Severity weight used for illustration |
|---|---|---|---|
| Critical error | 4 | 2 | 10 |
| Major error | 12 | 5 | 3 |
| Minor error | 20 | 14 | 1 |
| Clean item | 164 | 179 | 0 |
| Any-error rate | 18.0% | 10.5% | Not severity-adjusted |
| Weighted severity points | 96 | 49 | Descriptive score only |
Fictional calculations
Completion rates:
baseline completion rate = 760 / 800 = 95.00%
after completion rate = 803 / 820 = 97.93%
Unanswered rates:
baseline unanswered rate = 40 / 800 = 5.00%
after unanswered rate = 17 / 820 = 2.07%
Rework rates:
baseline rework rate = 96 / 800 = 12.00%
after rework rate = 49 / 820 = 5.98%
Recurring labor intensity:
baseline labor minutes per incoming item = 5,520 / 800 = 6.90
after labor minutes per incoming item = 1,939 / 820 = 2.36
relative reduction = (6.90 - 2.36) / 6.90 = 65.7%
Normalize the baseline to the after volume before claiming released labor:
expected baseline labor at 820 items = 820 x 6.90 = 5,658 minutes
gross labor minutes released = 5,658 - 1,939 = 3,719 minutes
gross labor hours released = 3,719 / 60 = 61.98 hours
Assume, only for this fictional financial illustration, a disclosed loaded labor cost of $45 per hour:
gross labor value per 20-day window = 61.98 x $45 = $2,789.10
incremental recurring software cost = $900
net operating value per window = $2,789.10 - $900 = $1,889.10
One-time cost:
implementation and training labor = (72 + 12) x $45 = $3,780
external implementation fee = $4,500
total one-time cost = $8,280
simple payback = $8,280 / $1,889.10 = 4.38 comparison windows
Weighted quality score:
baseline points = (4 x 10) + (12 x 3) + (20 x 1) = 96
after points = (2 x 10) + (5 x 3) + (14 x 1) = 49
points per sampled item: 0.480 before, 0.245 after
What this fictional example supports
It supports a descriptive statement:
In two specified fictional 20-day windows, recurring labor minutes per incoming item were lower after the workflow change, while completion, unanswered, rework, and sampled error measures also moved in favorable directions. The example includes implementation, training, maintenance, and software cost.
It does not support:
AI caused a 65.7% productivity improvement.
The before-and-after comparison is observational. Other changes, sampling error, regression to the mean, operator learning, task mix, and the novelty period could explain some or all of the difference. A stronger design might use phased rollout, matched teams, randomized assignment where ethical and practical, interrupted time series, or another pre-specified comparison. Even then, assumptions and spillovers require disclosure.
Minimum evidence checklist
A reader should be able to mark every item as supplied, partially supplied, or absent. “Confidential” can be a legitimate reason not to publish raw data, but it does not turn missing evidence into verified evidence.
- The claim names the exact workflow and unit of work.
- The baseline window has exact start and end dates.
- The after window has exact start and end dates.
- Both windows disclose operating days, hours, shifts, or seasonality.
- Incoming, eligible, attempted, completed, open, and unanswered volumes reconcile.
- Exclusion rules and excluded counts are disclosed.
- Task categories and complexity mix are comparable or stratified.
- The start event and accepted-completion event are defined.
- Elapsed time and hands-on labor time are reported separately.
- Timing includes review, correction, retries, exceptions, and recovery.
- The statistic is named: mean, median, percentile, total, or another measure.
- The number of timed observations and missing timestamps is disclosed.
- Rework count, rate, and labor minutes are reported.
- Error severity definitions were fixed before grading.
- Critical, major, minor, and cosmetic errors are not blended into one reassuring average.
- Unanswered, abandoned, or silently dropped cases remain visible.
- A quality sample covers apparently successful cases, not only complaints.
- The sampling frame, method, size, strata, and query are documented.
- Reviewer qualifications, rubric, disagreements, and adjudication are disclosed.
- Human-review coverage, timing, authority, and override rates are reported.
- Operator and reviewer training time is included.
- Implementation labor, fees, testing, and parallel running are included.
- Recurring licenses, usage, compute, storage, and logging costs are included.
- Maintenance, monitoring, evaluation, and incident labor are included.
- Audit logs can connect inputs, versions, decisions, retries, and final outcomes.
- Failed triggers and items that never entered the automation are reconciled with the source system.
- Changes in staffing, policy, demand, tooling, and task mix are disclosed as alternative explanations.
- The destination of released capacity is measured.
- Financial savings are separated from capacity, speed, throughput, and quality claims.
- The conclusion uses associative language unless the study design supports causality.
- Raw or controlled evidence is available to an independent reviewer.
- Conflicts of interest, vendor involvement, incentives, and affiliate relationships are disclosed.
- Known limitations and adverse outcomes appear near the headline, not only in fine print.
- The workflow, model, prompt, connector, rule, and policy versions are dated.
- The case study states what would falsify or downgrade its claim.
Claim grading rubric
Grade each material claim, not the article as a whole. One case study may have Grade A timing evidence and Grade D causal language.
| Grade | Evidence standard | Language allowed |
|---|---|---|
| A: Reproducible | Reconciled item-level data, pre-specified measures, comparable windows or stronger design, complete cost and quality accounting, independent review, limitations, and accessible evidence | Precise descriptive claim; causal claim only if the design and assumptions genuinely support it |
| B: Substantiated | Clear windows, denominators, labor, quality, failures, costs, and audit logs; minor gaps do not plausibly reverse the result | “Observed,” “was associated with,” or “coincided with” |
| C: Directional | Some real records and comparable measures, but incomplete sampling, cost, exception, or capacity evidence | “The records suggest,” with prominent limitations |
| D: Anecdotal | Self-report, selected examples, dashboard screenshot, or unverified totals without a reconstructable denominator | “The operator reports”; no generalized performance claim |
| E: Unsubstantiated | Undefined metric, missing window or denominator, contradictory arithmetic, inaccessible evidence, or marketing assertion presented as fact | No performance conclusion |
Use the lowest grade triggered by a material gap. A Grade A average-time calculation cannot rescue absent critical-error disclosure.
Scoring worksheet
| Dimension | 0 points | 1 point | 2 points |
|---|---|---|---|
| Population | Missing denominator | Partial counts | Reconciled full funnel |
| Time | Selected or undefined | Named metric, incomplete labor | Comparable elapsed and labor distributions |
| Quality | Testimonials only | Complaint or convenience sample | Independent full-frame sample with severity |
| Failures | Omitted | Aggregate failure rate | Itemized severity, unanswered cases, recovery |
| Costs | Software price omitted or partial | Recurring cost only | One-time and recurring total-cost ledger |
| Human work | “Human in loop” slogan | Review step named | Coverage, authority, edits, training, labor |
| Traceability | Screenshot | Export or dashboard | Reconciled item-level logs and versions |
| Attribution | Causal language from timing | Caveated association | Design supports the stated inference |
| Capacity | Assumed savings | Intended use named | Observed destination and outcome measured |
| Independence | Vendor-written only | Internal review | Qualified independent review and conflicts disclosed |
A total score can help triage, but hard failures override it. Missing critical incidents, contradictory arithmetic, an unreconciled denominator, or false causal language should cap the grade.
Red flags that deserve an immediate downgrade
- The fastest or best run is presented as the typical result.
- “Time saved” measures machine runtime while ignoring human review.
- Before time is a recollection; after time comes from logs.
- The windows differ in length, volume, staffing, or complexity without adjustment.
- The workflow counts drafts as completions.
- Failed, unanswered, filtered, or open items disappear from the denominator.
- Error rate means only reported complaints.
- A zero-error claim has no sample size or independent review.
- Critical and cosmetic errors are averaged together.
- Rework is described as “minor edits” without count or minutes.
- Training and implementation are excluded from return-on-investment calculations.
- Software cost omits usage, integrations, monitoring, or maintenance.
- The automation vendor wrote the case study and the customer cannot provide evidence.
- The customer name is withheld and no independent reviewer had controlled access.
- A percentage improvement has no absolute values.
- Inquiry reduction is assumed to mean quality improvement.
- Headcount avoidance is claimed without an approved hiring counterfactual.
- Released time is monetized even though paid hours did not change.
- More meetings are treated as either value or waste without measurement.
- A model or prompt changed during the after window without versioned results.
- “AI-powered” receives credit for changes caused by process redesign, training, or data cleanup.
- Correlation is described with causal verbs such as “caused,” “delivered,” or “drove.”
- Social engagement is used as proof that the workflow performed.
A reusable case-study disclosure template
Copy this structure into an internal evidence pack. Blank fields indicate missing evidence, not permission to infer a favorable answer.
# Automation case-study evidence disclosure
## Status and authorship
- Draft status:
- Case-study author:
- Workflow owner:
- Data analyst:
- Independent reviewer:
- Vendor involvement:
- Financial, affiliate, or other conflicts:
- Evidence checked date:
## Claim
- Exact descriptive claim:
- Causal claim, if any:
- Claim grade:
- Decision this evidence supports:
- Known conditions where the claim should not transfer:
## Workflow boundary
- Unit of work:
- Eligible population:
- Start event:
- Accepted-completion event:
- Systems and versions:
- Human roles:
- External effects:
## Measurement windows
- Baseline dates and operating hours:
- After dates and operating hours:
- Seasonality and demand:
- Staffing and training differences:
- Policy, product, or data changes:
## Funnel counts
- Incoming:
- Ineligible with reasons:
- Eligible:
- Attempted:
- Straight-through completed:
- Human-reviewed completed:
- Escalated:
- Failed:
- Unanswered:
- Open at cutoff:
- Reconciliation check:
## Time and labor
- Timing source:
- Missing timestamps:
- Mean:
- Median:
- 90th or 95th percentile:
- Hands-on first-pass labor:
- Review labor:
- Rework labor:
- Exception and recovery labor:
- Maintenance labor:
## Quality
- Success rubric:
- Severity definitions:
- Sampling frame and query:
- Sampling method and size:
- Reviewer qualifications:
- Critical, major, minor, and cosmetic counts:
- Disagreement and adjudication:
- Known blind spots:
## Reliability and recovery
- Trigger coverage:
- Retry and timeout policy:
- Duplicate prevention:
- Failure queue:
- Alert and owner:
- Recovery procedure:
- Incident count and labor:
## Cost
- Currency and checked date:
- One-time implementation labor:
- One-time external fees:
- Training:
- Recurring licenses and usage:
- Compute, storage, and logging:
- Monitoring and evaluation:
- Maintenance:
- Loaded labor rate and basis:
- Payback formula and result:
## Released capacity
- Gross labor hours released:
- Paid hours or overtime changed:
- Backlog or throughput changed:
- Capacity destination:
- Meetings added or removed:
- Outcome of the capacity use:
## Attribution and limitations
- Alternative explanations:
- Comparison or counterfactual:
- Statistical uncertainty:
- Missing evidence:
- What would falsify or downgrade the claim:
- Final language approved for use:
How to write the conclusion without overstating causality
Use language that matches the design.
For a basic before-and-after comparison:
During the measured after window, the team recorded lower labor minutes per eligible item than during the baseline window. The change coincided with the automation rollout, training, and process redesign. This comparison does not isolate the effect of AI.
For self-reported figures without audit access:
The operator reports that the task changed from 30 minutes to 90 seconds. We did not receive the underlying denominator, timing distribution, quality sample, or run logs, so we treat the figure as demand or hypothesis evidence rather than verified performance.
For a quality claim:
In a disclosed sample, the after period had fewer observed major errors. The sample is descriptive and does not establish that the automation caused the difference.
For capacity:
The workflow released an estimated number of hands-on hours under the disclosed method. Calendar and work records show where that capacity went; it should not be described as cash savings unless spending also changed.
Avoid “AI saved,” “AI eliminated,” and “AI drove” when the study only observes a change after rollout. Process mapping, cleaner data, changed staffing, training, and new review rules may deserve some or all of the credit.
Demand evidence is not performance evidence
Public posts can show that people are interested in AI automation, meeting summaries, inbox processing, and workflow delegation. They can help an editor decide whether readers have a question. They cannot verify a customer outcome, error rate, cost, or causal effect.
For this article, the following X posts are treated only as self-reported demand evidence:
- https://x.com/mikami01_ai/status/2039669467295482069
- https://x.com/shota7180/status/2038812375642702310
- https://x.com/shota7180/status/2043513287959265788
- https://x.com/0xfene/status/2042047157767926056
Views, likes, reposts, replies, and author-provided examples can change and may be incomplete. No X metric is used here as verified performance. Engagement does not validate the reported four-hour, 40-minute, 12-to-three, or 30-minute-to-90-second figures.
Frequently asked questions
1. Is a before-and-after table enough?
No. It is a useful summary, but a reader also needs the windows, denominators, task mix, measurement method, exclusions, failures, quality sample, labor, cost, and alternative explanations. The table should be generated from reconcilable records.
2. How long should the baseline window be?
There is no universal duration. It should cover enough volume and operational variation to represent ordinary work, including busy periods and common exceptions. State why the selected window is adequate and disclose seasonality.
3. Should we measure average time or median time?
Usually both, plus a tail statistic. The mean captures total burden but is sensitive to extreme cases. The median describes a typical case but can hide severe delays. A 90th or 95th percentile shows the slow tail.
4. Do failed or unanswered cases count in the denominator?
Yes, when they were eligible for the workflow. Excluding them rewards failure. Report completed, failed, unanswered, open, and ineligible items separately and reconcile them to intake.
5. What counts as rework?
Any additional human or machine work required because the first output was not acceptable under the predefined rubric. Report the item count, correction minutes, severity, and whether the rework affected a customer or downstream system.
6. Can customer complaints stand in for a quality sample?
No. Complaints detect only noticed and reported problems. Sample apparently successful items, unanswered cases, and exceptions from the full population. Use complaints as an additional signal.
7. Does a human approval step make the automation safe?
Not by itself. Safety depends on review coverage, timing, evidence shown, reviewer competence, authority, workload, and what happens when the queue is unavailable. Measure overrides and errors that reviewers missed.
8. When can we say AI caused the improvement?
Only when the study design and assumptions support a causal inference. A simple before-and-after change is usually association. Stronger designs can reduce alternative explanations, but none remove the need to disclose assumptions, spillovers, and concurrent changes.
9. How should software cost be reported?
Separate one-time and recurring cost. Include licenses, usage, compute, storage, logging, integrations, monitoring, evaluation, maintenance, and implementation. State the currency, price basis, allocation method, and checked date.
10. Is released employee time the same as financial savings?
No. Released time is capacity. Financial savings require reduced cash expenditure or a defensible avoided-cost counterfactual. Report both separately.
11. What if the company cannot publish raw logs?
It can provide a data dictionary, extraction query, hashes or manifests, aggregate reconciliation, redacted examples, access controls, and an independent review statement. Confidential evidence can support a claim, but readers should know what they cannot inspect.
12. How should we treat a drop in inquiries from 12 to three?
As ambiguous until the denominator and mechanism are known. It may reflect fewer defects, changed demand, changed classification, missing logging, unanswered customers, or suppressed reporting. Pair it with intake, resolution, abandonment, and quality evidence.
13. What if the automation is faster but errors are more severe?
Do not approve the headline based on speed. Predefined critical or major error thresholds should act as hard gates. Report the speed gain and the adverse quality result separately.
14. What should we do when saved time becomes meetings?
Measure the meeting load, purpose, attendance, decisions, and follow-up work. Meetings may create coordination value or erase usable capacity. The case study should report the destination rather than assuming either conclusion.
15. Can vendor audit logs prove return on investment?
They can support run counts, timing, versions, failures, and usage, but they rarely contain the complete labor, quality, source-intake, cost, or capacity story. Reconcile platform logs with source systems, reviewer records, finance records, and work outcomes.
16. Are social-media engagement metrics useful?
They can indicate audience interest or self-reported demand. They do not verify the workflow’s performance, generalizability, or causal effect. Keep them out of the evidence grade.
Related Context Wire guides
Use these existing internal guides to build the evidence pack:
- Build an LLM evaluation dataset from real failures
- Design human-in-the-loop AI agents
- Measure human-review coverage in AI workflows
- Add machine gates for AI output
- Compare AI agent observability tools
- Handle LLM API timeouts, retries, and idempotency
- Run a self-hosted evaluation gate
- Use token budgets and circuit breakers for agents
- Review AI-generated code
- Operate production LLM APIs
Sources checked 2026-07-24
The sources below are primary official publications or official vendor documentation. They support the measurement and control principles in this article; they do not verify the self-reported case-study figures.
NIST
- NIST AI Risk Management Framework Core: https://airc.nist.gov/airmf-resources/airmf/5-sec-core/
- NIST AI RMF Playbook: https://airc.nist.gov/docs/AI_RMF_Playbook.pdf
- NIST, “Challenges to the Monitoring of Deployed AI Systems”: https://www.nist.gov/news-events/news/2026/03/new-report-challenges-monitoring-deployed-ai-systems
- NIST Privacy Framework 1.1 operating guidance: https://www.nist.gov/privacy-framework/using-privacy-framework-11
- NIST SP 800-92, Guide to Computer Security Log Management: https://csrc.nist.gov/pubs/sp/800/92/final
OECD and labor agencies
- OECD, “Using AI in the workplace”: https://www.oecd.org/en/publications/using-ai-in-the-workplace_73d417f9-en.html
- OECD, “The Adoption of Artificial Intelligence in Firms”: https://www.oecd.org/en/publications/the-adoption-of-artificial-intelligence-in-firms_f9ef33c3-en.html
- OECD, “The impact of AI on the workplace: Evidence from OECD case studies of AI implementation”: https://www.oecd.org/content/dam/oecd/en/publications/reports/2023/03/the-impact-of-ai-on-the-workplace-evidence-from-oecd-case-studies-of-ai-implementation_b4c2c6ee/2247ce58-en.pdf
- OECD, “How widespread is algorithmic management in workplaces?”: https://www.oecd.org/en/publications/how-widespread-is-algorithmic-management-in-workplaces_cda7a114-en/full-report.html
- U.S. Bureau of Labor Statistics, productivity calculation method: https://www.bls.gov/opub/hom/msp/calculation.htm
- U.S. Bureau of Labor Statistics, productivity data sources: https://www.bls.gov/productivity/sources/
- U.S. Department of Labor, Field Operations Handbook Chapter 64 work-measurement guidance: https://www.dol.gov/agencies/whd/field-operations-handbook/Chapter-64
Official automation-vendor documentation
- Microsoft, Power Automate activity logs: https://learn.microsoft.com/en-us/power-platform/admin/activity-logging-auditing/activity-logs-power-automate
- Microsoft, Power Automate error handling: https://learn.microsoft.com/en-us/power-automate/guidance/coding-guidelines/error-handling
- Microsoft, Power Automate monitoring metrics: https://learn.microsoft.com/en-us/power-platform/admin/monitoring/monitor-power-automate
- UiPath, Automation Cloud audit logs: https://docs.uipath.com/automation-cloud/automation-cloud/latest/admin-guide/about-logs
- UiPath, Action Center human intervention: https://docs.uipath.com/action-center/automation-cloud/latest/user-guide/introduction
- Zapier, Zap History: https://help.zapier.com/hc/en-us/articles/8496291148685-View-and-manage-your-Zap-history
- Zapier, troubleshooting workflow errors: https://help.zapier.com/hc/en-us/articles/8496037690637-How-to-troubleshoot-errors-in-Zap-workflows
- Zapier, custom error handling: https://help.zapier.com/hc/en-us/articles/22495436062605-Set-up-custom-error-handling
Demand-evidence URLs
The four X URLs are listed in the earlier demand-evidence section. They were checked on 2026-07-24 and are classified only as self-reported indicators of audience interest, not as primary performance evidence.
Final audit rule
Believe only the narrowest claim that the evidence can carry. A fast automated draft is not an accepted outcome. A lower task time is not automatically lower labor. Released capacity is not automatically cash savings. Fewer inquiries are not automatically better service. A before-and-after association is not automatically causation.
The strongest case study lets a skeptical reader reconstruct the funnel, arithmetic, quality sample, cost ledger, review boundary, failure path, and capacity outcome. If the headline survives that reconstruction, it is useful evidence. If it does not, keep the number as a hypothesis and keep measuring.