Appendix B — A 90-Day Process Pilot
This appendix turns the first pilot into a working cadence. It is a proposed operating sequence, not a universal timetable. Shorten or lengthen it to fit the process, but retain the decision points.
Before Day 1: Choose the Flow
Select one recurring process with a visible result and a consequence the business can recognize. Name the process owner, review owner, knowledge maintainer, and accountable executive. Write the pilot charter: result, case boundary, AI-enabled action, human judgments, review condition, escalation route, and stop authority.
Do not start with a technology inventory. A pilot that begins with tools usually ends with adoption data. A pilot that begins with a result gives the team a way to judge whether the tools mattered. Choose a flow that occurs often enough to generate evidence, has a nameable beginning and end, and contains one bounded action AI can support without obscuring human judgment.
Days 1–15: Make the Current Work Visible
Map one ordinary case from trigger to result. Record inputs, decisions, handoffs, waiting, rework, exceptions, and the context people use but rarely document. Do not automate every visible step. Choose one AI-supported action that can be tested without obscuring the complete result.
Create a small baseline. Do not manufacture precision. It may include turnaround time, repeat work, customer wait, exception volume, quality checks, or the time a new colleague needs to handle a routine case. The baseline is a reference point for review, not a promise that every number will move.
Use the first two weeks to identify the cases the pilot should not take. This is not a failure of ambition. It is part of setting a boundary that the business can trust. List the missing context, unresolved customer commitment, unusual exception, or consequence that would require a different level of authority. Keep the established process responsible for those cases until the pilot can state how it will carry them.
Meet briefly with the people who currently repair the work when it goes wrong. Ask them for the exception that does not appear in the documented process, the signal that tells them a normal case is turning abnormal, and the consequence they protect when they slow the work down. These conversations often reveal the judgment the pilot needs to preserve. Record the contribution, show how it changes the design, and let the contributor see the resulting rule or boundary.
At the first review, ask whether the boundary is clear enough to run a parallel route. If not, narrow the case class or clarify the result before proceeding.
Days 16–45: Run One Bounded Change
Run the AI-enabled route beside the established process. Keep track of the outcome, time and effort, review burden, exceptions, recovery work, and changes to context. Watch for the false signal of high activity. A pilot may generate many prompts, drafts, and meetings while producing no more stable result. It may also show a small change in volume but a large gain in clarity: a repeated exception is now visible, a review owner can reject a doubtful output, or a process owner has a clear route to escalate.
Introduce one defined AI-supported action: retrieve relevant knowledge, organize incoming information, create a first draft, flag routine anomalies, prepare a standard response, or assemble material for human judgment. Keep the change narrow enough to identify what caused an improvement or a failure.
The process owner holds a short weekly operating review with the people doing the work. Compare effort and waiting: where did work genuinely move faster, and where did effort shift into checking or cleanup? Review the result: did quality, timeliness, rework, customer response, or another fit-for-purpose outcome change? Capture useful prompts, source material, judgment standards, error patterns, exceptions, and workflow changes. Finally, test the chain: was it ever unclear who could decide, stop, or escalate? Repair that boundary while the pilot is small. In a randomized experiment on a defined AI-extraction task, correction burden and participants’ attitudes toward AI affected acceptance and correction behavior; that finding is a reminder to examine the design of review work, not a verdict on all human review.7
At the end of each weekly review, update one shared record: case type, result, exception or correction, relevant context, decision made, and rule to test next. Its purpose is to make the next decision better and to give the next team a usable starting point.
Days 46–75: Test the Operating Result
Test ordinary variation, not only a favorable example. Ask whether another qualified person can find the context, understand the review condition, and act on a material exception. Compare the pilot with the established route without hiding rework or repair outside the new flow.
Run the workflow under ordinary conditions, including ordinary volume and the exceptions that make the work consequential. Do not protect the pilot so carefully that it never meets the reality it is supposed to improve. Is the result stable beyond an unusually good week? Is one person still rescuing the flow? Has capacity really been released when inputs—time, labor, cost, waiting, and management attention—are compared with outcomes—volume, quality, rework, customer response, and retained knowledge?
This is also the point to compare the pilot with the established process honestly. Ask whether the old process carried unrecorded judgment that the pilot still depends on. Ask whether a faster step created a new queue elsewhere. Ask whether the people doing the work have begun to route difficult cases around the pilot. A comparison that ignores these effects does not show capacity; it shows only that the visible part of the work changed.
Assess four signals together: the intended outcome; an error or consequence signal; the workload and delay moved elsewhere; and the learning captured for the next case. The measures need not all be numerical. Their purpose is to keep a dashboard from replacing the operating conversation.
Days 76–90: Make the Decision
At the 90-day review, choose one of four outcomes: continue and deepen within the boundary; revise the workflow, context, or review condition; redesign the role or result boundary; or stop. The decision should name the person who carries the next action, the evidence that supports it, and the date of the next review.
Bring the same people back to the same one-page view. The process owner presents operating evidence. The accountable executive decides whether the business result justifies the next move. The CHRO assesses role, capability, and performance implications. The CIO or technology lead assesses whether the workflow and controls can support the decision. The CEO arbitrates the trade-off only when it becomes enterprise-wide or irreversible.
Do not treat the end of the pilot as a presentation moment. Treat it as a decision point. A pilot that stops can still leave a valuable record of what the organization could not yet carry. A pilot that continues should leave a clearer method, a visible accountability chain, and organizational memory that the next process can use.
Before closing the review, state the authority boundary for the next cycle. If the pilot deepens, which new case type may it take? If it revises, which action is being narrowed or redesigned? If it stops, who is responsible for retaining the lesson and informing the people affected? If it redesigns the role or process, which accountable executive will decide the resource and transition questions? A decision without these assignments simply moves uncertainty into the next meeting.
Pilot Scorecard
| Question | Evidence to Review |
|---|---|
| Did the intended result improve? | Output quality, customer response, timeliness, or another stated condition. |
| What work moved or became harder? | Review burden, rework, waiting, coordination, and recovery work. |
| Which exceptions mattered? | Case types, missing context, unclear authority, and failed controls. |
| Is accountability visible? | Named human core, review owner, escalation route, and decision record. |
| What did the organization retain? | Rules, examples, corrections, exception patterns, and changed role expectations. |
Questions for the 90-Day Review
- Which result changed in a way the organization can explain?
- Which result did not change despite faster execution?
- Where did human judgment remain essential, and was that judgment given the authority it needed?
- What error, exception, or customer consequence taught the team the most?
- What did the organization retain so the next team will not have to learn the same lesson again?
- Does the evidence point toward lower cost, output, quality, or redesign as the next capacity destination?
- Is the named process owner still able to carry the complete result, or has the consequence grown beyond the unit’s boundary?
What Not to Do in the First Pilot
Do not begin with the most politically sensitive workflow because it has the largest apparent cost. Do not make the pilot team prove a financial return before it has established a stable result. Do not expand the AI action while the responsible person, review point, or escalation route remains unclear. Do not treat a written plan as proof that the workflow now operates differently.
Do not let a pilot survive by borrowing invisible work from the established process. If experienced colleagues are quietly correcting difficult cases, recording missing context, or repairing customer consequences outside the new workflow, put that work on the review table. It may be necessary support in the early stage, but it is not evidence that the new method is ready.
Do not treat employee caution as a communications problem by default. If people give generic descriptions, withhold the difficult exception, or route important cases around the pilot, inspect the exchange the organization has created. Are contributors visible? Does the process state how their context will be used? Is an error report likely to create punishment? Improve the boundary, return mechanism, and authority design rather than demanding more enthusiasm.
Do not confuse a high-quality demonstration with an operating method. A prepared example often contains clean context, a known expected answer, and a team ready to repair any defect. Ordinary work contains missing information, competing priorities, handoffs, customers, and exceptions. A pilot earns the right to deepen only after it meets enough ordinary variation to show that the result boundary and repair path are real.
Do not let the cadence become a substitute for judgment. Ninety days is a useful proposed operating window because it gives a team time to observe a real cycle, but it is not a magic duration. Some processes will show a stable result sooner; others will need longer because their consequence, volume, or cycle time is different. The right question is whether the organization has evidence to decide, not whether a calendar date has arrived.
When the evidence is mixed, state what is mixed. A result may improve while review burden rises. A team may discover useful context while finding that the current decision boundary is too weak. A capacity gain may be real but not yet large enough to justify a staffing choice. These are not failures of measurement. They are the conditions an accountable executive must weigh. The decision record should preserve the trade-off instead of forcing a premature yes-or-no story.
A One-Page Decision Record
Process and result. State the process, the recipient of the result, and the condition the team was trying to improve.
Boundary. State which cases the workflow handled, which remained in the established process, and what the AI-supported action could not decide.
Evidence. Record the result observed, the most important exception, the relevant input or outcome signal, and any unrecorded work that the comparison exposed.
Judgment and authority. Name the person who carried the result, the review condition, the decision that required human judgment, and the escalation route used or still missing.
Learning. List the rule, context, example, correction, or role change that another person can reuse. If nothing is reusable, explain why the method is still too dependent on local expertise.
Decision. Choose continue and deepen, revise, redesign, or stop. Name the next role, the next boundary, and the date on which the evidence will be reviewed again.