Chapter 15 — From One Pilot to a Repeatable Capability
A pilot matters when the second process does not begin from zero. The goal is not a growing inventory of AI projects. It is a repeatable way of choosing a process, making its result visible, assigning responsibility, retaining judgment, and deciding what to do next.
The first ninety days establish one real process. The next ninety days turn temporary practice into operating rules. By the end of the first year, the company should be able to tell whether a second process starts faster because the first one left behind a usable method. These are proposed operating intervals, not universal performance benchmarks.
Days 0–90: Establish One Real Process
Choose a process whose business impact an accountable executive can see and whose result can be observed inside a reasonable review window. Do not choose the most complicated process because it promises the largest theoretical gain. The first process needs a clear result, a visible pain point, and a named process owner willing to carry the ordinary operating work.
Begin with the three ledgers. The time ledger records recurring effort and the actions that consume it. The waiting ledger records delays in approvals, inputs, information, and handoffs. The knowledge ledger records which judgment standards are reusable and which still live only in individual experience. The figures need not be artificially precise, but they must be traceable to ordinary work.
The three ledgers make different losses visible. A time ledger can show that a task is expensive, but not whether it waits in a queue before anyone begins it. A waiting ledger can show delay, but not whether the delay protects a necessary judgment or merely covers a missing decision. A knowledge ledger can show that a capable employee is carrying a method privately, but not whether that method changes the result. Read the ledgers together. They help the group decide whether the first intervention should reduce routine effort, remove a needless wait, capture a judgment standard, or clarify responsibility before changing execution.
Use the ledgers to prevent a familiar misuse of AI measurement. If the team starts from a request to show time saved, it may discover that the work did not disappear; it changed hands. If it starts from a request to show more output, it may discover that quality or review is now carrying the burden. If it starts from a request to reduce cost, it may discover that the relevant knowledge has not yet been retained. A real baseline does not prove a business case in advance. It helps the group see which business case, if any, the workflow could eventually support.
During the first cycle, make the existing process visible before trying to improve it. Record triggers, inputs, handoffs, decisions, rework, and exceptions. Ask where a person uses judgment, where the same question is answered repeatedly, where a customer waits, and where work disappears into a private spreadsheet or message thread. The map need not be perfect. It needs to be honest enough to show what the new route is changing.
Use the parallel pilot to test one bounded intervention. Do not redesign every handoff at once. The group may begin with intake, retrieval, drafting, classification, quality review, or exception triage. The test should be small enough to reverse and consequential enough to reveal whether the complete result improves.
At the end of each review, update one shared record. Include the case type, the result, the exception or correction, the relevant context, the decision made, and the rule to test next. The purpose is to make the next decision better and to give the next team a usable starting point.
The record should make the real work visible, not create a reporting ritual. A useful entry can fit on a page. It gives the team a way to distinguish an error in the output from a gap in the context, a weak review condition, or an authority boundary that nobody had previously named. If the same exception returns, it should be possible to see whether the earlier decision changed the workflow or merely documented the problem.
Days 91–180: Turn Temporary Practice Into a Method
The next phase is not about adding more pilots. It is about making the first one survivable without its original advocates. Convert recurrent decisions into explicit criteria. Turn successful prompts, examples, and exception patterns into reusable organizational memory.243 Put review standards, recovery actions, and escalation routes where the next team can find them.
Institutionalizing does not mean freezing the method. The point is to make the current best judgment inspectable and changeable. A rule should carry enough context that a future user can understand why it exists. An exception should have a path back to the rule rather than remaining a private workaround. A review standard should tell people what they are protecting, not merely what box to check.
There are four pieces to institutionalization. First, stabilize the result. Define the outcome in language the business recognizes and identify the evidence that shows it is holding. A process can be fast and still unstable. The team needs to know whether comparable cases receive comparable treatment, whether rework is falling, and whether the result survives ordinary pressure.
Second, make the method legible. The next team needs more than the final prompt or configuration. It needs the trigger, the input assumptions, the human judgment boundary, the review condition, common exceptions, and the recovery action. Keep the method short enough to use in work. If it needs a full presentation to explain every time, it has not yet become an operating method.
Third, make learning maintainable. Assign a knowledge maintainer, a place for revised rules, and a recurring check on whether a rule still fits the work. Capture the reason for a meaningful change. Without that reason, a later user may undo a safeguard because it appears arbitrary, or repeat an old error because the organization stored a correction without its context.
Fourth, make the role change visible. When execution changes, people need to know what grows in importance: exception judgment, customer interpretation, source quality, review, coaching, and process improvement. If performance still rewards only volume, the organization will treat the method as a private productivity trick rather than a shared capability.
Ask also whether the work now depends on different capabilities. Some people may spend less time producing first drafts and more time interpreting exceptions, coaching others, maintaining context, or improving the flow. Early generative-AI experiments have reported shifts in some professional and technical work toward supervision, analysis, explanation, interpretation, and troubleshooting.1 Those are not incidental changes. They are the role redesign that makes the method durable. The CHRO and accountable executive should make these changes visible before performance expectations and staffing conversations harden around an obsolete job description.
At this point, review the method against the organizational memory it claims to use. Can a new person distinguish a current rule from an obsolete one? Can the team locate the source of a recurring exception? Does the method record why a decision changed, not only that it changed? These questions are unglamorous, but they distinguish a reusable operating capability from an impressive system maintained by a few people who remember its history.
Use the 180-day review to test dependency. What happens if the original process owner is away? What happens if the usual reviewer challenges a result the system treats as routine? What happens when a context source changes or a new customer situation does not resemble the examples? The organization does not need to simulate every disaster. It needs to know whether ordinary variation exposes a missing role, a missing rule, or an escalation path that exists only in someone’s memory.
Day 365: Test Whether the Organization Can Replicate
By the end of the first year, the question changes. It is no longer whether the original team can run the method. It is whether another team can start a related process with less reinvention and clearer accountability.
Replication should preserve a common method, not force identical workflows. Each new process should still name its result, its human core, its AI-enabled actions, its context, its review condition, and its escalation route. What should travel is the discipline: the way the firm distinguishes a task from a result, a process owner from an accountable executive, and a useful exception from a private workaround.
The test for replication is not whether every team uses the same software or copies the same process map. It is whether the second team can establish a credible operating boundary more quickly because the first team left behind an intelligible method. The second process may need a different review condition, a different source of context, or a different capacity decision. It should not need to rediscover that a result has to be named, a human core has to be empowered, and a recurring exception has to change the next version of the work.
The first replication should be close enough to test the method and different enough to reveal its real limits. Moving from one customer-response flow to another may test whether the review condition and knowledge practice travel. Moving immediately from a customer-response flow to a high-consequence decision may prove only that the original boundary was too vague. Choose the next process as a management decision, not as an opportunity to claim scale.
By the end of the year, leaders should be able to compare two questions. Did the second process set up its result, role map, baseline, and review route faster than the first? And did it preserve the discipline of changing the method when its context differed? If the answer to the first is yes and the second is no, the company has created a template without judgment. If the answer to the second is yes and the first is no, the company may be learning but has not yet made that learning reusable. A repeatable capability needs both.
Make Replication a Management Commitment
Replication needs an explicit commitment from leadership. The CEO decides which enterprise objective the next process should serve and arbitrates enterprise-wide or irreversible trade-offs. The accountable executive supplies resources and resolves material business conflicts. The CHRO ensures role design, incentives, and transitions support the work. The CIO ensures that data, systems, and controls can support it. The AI transformation lead connects the sequence without becoming a permanent substitute for operating ownership.
Leadership should resist two temptations. The first is declaring victory after one successful pilot and asking every unit to adopt the same tool. The second is treating every process as so unique that no method can travel. The first creates scale without accountability; the second preserves local expertise without organizational memory. Replication is the discipline between them: a common way to name results, establish boundaries, learn from exceptions, and make the next decision.
One process may release time that should be reinvested in quality. Another may reveal a constraint that makes growth sensible. A third may make a role redesign unavoidable. Do not aggregate these into one enterprise claim that “AI saved X.” The accountable executive should bring each unit’s evidence to a shared operating conversation, while the CEO decides the enterprise-wide or irreversible trade-offs. That is how local learning becomes an enterprise choice without being stripped of its context.
The portfolio should remain small enough to govern. A company that launches many pilots before it has institutionalized the first method often creates a reporting burden rather than a capability. It becomes difficult to see which result each team owns, which lessons can travel, and which exceptions require enterprise attention. Expand the portfolio as the organization can preserve result boundaries, usable learning, and visible decision rights—not simply as the number of available use cases grows.
That is the right meaning of scale. Scale is not merely more automated actions or more licenses in use. It is the ability to add or adapt accountability units without losing the foundations that make the company one organization: common purpose, usable memory, appropriate permission, and a route for resolving a conflict that crosses the unit. A company scales responsibly when a new unit can inherit those foundations while still surfacing a context-specific exception. It scales badly when every unit quietly invents its own rules or when every local decision has to wait for central approval.
The leadership task is to preserve both learning and selectivity. A second process should not be forced to demonstrate the same result as the first, but it should be able to show why its result matters, what it is borrowing from the earlier method, and where its context makes that method insufficient. This produces a portfolio of accountable results rather than a collection of unrelated demonstrations. It also gives executives a clearer basis for deciding where capacity, investment, and organizational attention should go next.