In June 2024, McDonald’s pulled its IBM drive-through voice-AI partnership after eighteen months of viral order errors. Sweet teas multiplying. Butter and ketchup landing in ice-cream orders. Adjacent-lane orders bleeding into each other.
The technology had been tracking its plan. The natural-language model had hit its accuracy benchmarks under test conditions. The acoustic processing handled drive-through noise at the rates the engineering team had targeted. Deployment across more than a hundred US restaurants matched the rollout schedule. And then the operating environment — thousands of customers per day per location, accents the training data had not weighted, ninety-second pull-up-to-pickup expectations — overwhelmed every benchmark the lab had set.
This was not a technology failure. It was a scaling failure.
Seventy-eight per cent of enterprises have launched AI initiatives. McKinsey’s research shows one per cent have reached true implementation maturity. The difference between the one per cent and everyone else is not talent, capital, or algorithmic sophistication. It is the discipline of scaling — the craft of moving from a controlled pilot to an enterprise-wide capability without mistaking the first for the second.
This article gives the framework for that discipline.
Selecting the First Use Case
The selection of the inaugural agentic AI use case often determines the trajectory of the entire transformation. This decision carries weight far beyond the immediate pilot. It sets expectations, builds or erodes confidence, and establishes patterns that echo throughout the scaling journey.
Four criteria balance ambition with pragmatism.
Bounded Autonomy. Trust in autonomous systems builds gradually. Initial deployment should operate within clearly defined parameters where rules are explicit, outcomes measurable, and boundaries naturally contained. Campaign budget allocation exemplifies the principle. Agents can make sophisticated optimisation decisions while operating within predetermined limits that prevent runaway spending or brand damage.
Data Richness. Agentic AI thrives on continuous information streams. The quality and accessibility of data often determines pilot success more than algorithmic sophistication. Customer journey orchestration succeeds because modern marketing stacks already generate abundant behavioural signals that agents can use for personalisation decisions. Without rich data flows, even the most sophisticated agents operate blindly.
Human Oversight Readiness. Select use cases where the team can meaningfully monitor and guide AI decisions without creating operational bottlenecks. Email send-time optimisation illustrates the balance. Agents autonomously determine optimal sending moments for millions of recipients while humans retain complete control over content and strategic direction.
Value Visibility. Choose applications where success metrics are unambiguous and directly tied to business outcomes that matter to stakeholders. Lead scoring and routing provides this clarity through improved conversion rates, reduced response times, and enhanced sales productivity. When value is visible and quantifiable, scepticism transforms into advocacy.
Different sectors gravitate toward distinct starting points. Retail and e-commerce begin with dynamic pricing agents — IKEA’s three-year demand-sensing programme with Blue Yonder lifted forecast accuracy from 92 per cent to 98 per cent and cut the manual-correction load from eight per cent of orders to two per cent. Financial services begin with compliance-first content adaptation. B2B technology begins with account-based marketing orchestration. Healthcare begins with appointment scheduling — at Mid and South Essex NHS Foundation Trust, an AI-driven appointment-management system cut did-not-attends by almost a third within six months.
The anti-patterns matter as much. Brand voice creation allows agents to autonomously develop messaging, risking inconsistency and organisational resistance. Crisis response involves high-stakes, low-frequency events requiring human judgement. Cross-functional processes face unnecessary complexity from day one. Regulatory grey areas involve compliance requirements that remain undefined.
The Crawl-Walk-Run Methodology
The traditional rollout approach requires significant adaptation when applied to agentic AI. You are not merely rolling out features according to a predetermined schedule. You are introducing autonomous entities that learn, evolve, and reshape organisational dynamics.
Phase 1 — Crawl (Weeks 1-6). Establishing foundations.
The crawl phase creates controlled environments where agents demonstrate capability while building human confidence. This period establishes patterns of interaction that define the human-AI collaboration model going forward.
Weeks 1-2: Shadow Mode Deployment. Deploy the agent in shadow mode, processing real data and making decisions without executing them. This approach serves multiple purposes beyond technical validation. It allows teams to observe agent decision patterns, understanding not just what decisions are made but why. Edge cases emerge naturally, providing learning opportunities without operational risk.
Weeks 3-4: Recommendation Mode. Transition to recommendation mode where agents suggest actions but humans retain execution control. Agents present decisions with confidence scores and detailed reasoning. When humans override agent suggestions, documenting the rationale creates feedback loops that train agents on organisational preferences and unwritten rules. This is the bidirectional learning condition of hybrid intelligence operating in practice.
Weeks 5-6: Supervised Autonomy. Grant limited autonomous execution rights within tightly controlled parameters. Start with low-risk, high-frequency decisions where mistakes create minimal impact. Implement automatic circuit breakers that halt agent operations when anomalous patterns emerge. Require human approval for decisions exceeding defined thresholds — monetary, risk-based, or strategic.
Phase 2 — Walk (Weeks 7-20). Expanding capabilities.
The walk phase systematically expands agent autonomy while maintaining safeguards. This period transforms initial success into operational capability, moving from proof of concept to business value delivery.
Weeks 7-12: Bounded Expansion. Threshold limits for autonomous action increase based on demonstrated performance. Operational hours extend beyond standard business periods. Multi-variate decision-making capabilities come online, enabling agents to balance competing objectives like cost optimisation and customer satisfaction. Cross-functional data sources enrich agent intelligence.
Weeks 13-17: Collaborative Intelligence. True human-AI collaboration patterns emerge. Rather than humans overseeing every decision, clear divisions of labour take shape. Agents handle routine decisions with confidence, escalating only exceptions that fall outside their training. Humans shift focus to strategy, creative direction, and handling complex scenarios requiring judgement and empathy.
Weeks 18-20: Performance Optimisation. Historical data accumulated during earlier stages enables fine-tuning. Optimisation expands beyond efficiency metrics to encompass broader business outcomes like customer lifetime value and brand perception. Infrastructure preparations for scaling begin.
Phase 3 — Run (Weeks 21+). Scaled autonomy.
The run phase transforms successful pilots into enterprise capabilities. This transition requires excellence across technical, operational, and organisational dimensions.
Operational Excellence becomes paramount as agents assume mission-critical responsibilities. Redundancy and failover systems ensure continuous operation. Monitoring and alerting frameworks detect issues before they impact business outcomes. Detailed runbooks address edge cases and exception scenarios.
Continuous Learning distinguishes mature implementations from static deployments. Agent-to-agent knowledge transfer allows insights from one domain to benefit others. Federated learning across business units multiplies the value of each interaction without compromising data privacy.
Strategic Integration completes the transformation. Agent objectives align with business strategy through regular review cycles. Performance reviews for agents mirror human processes, acknowledging their role as team members rather than utilities.
The Failure Modes That Derail Most Programmes
Pattern recognition matters. Five failure modes recur across enterprise transformations.
Pilot purgatory. The pilot succeeds technically and culturally. Nobody scales it. Eighteen months later, the agent is still labelled “pilot” in the budget line and nobody can defend it or kill it. The fix is portfolio discipline applied from day one. If the pilot cannot graduate to production within twelve months, it should be retired.
Scale shock. The pilot worked at one customer journey, one country, one segment. Production exposed it to ten times the variance. The system collapsed. The fix is to stress-test the pilot with adversarial inputs and edge cases that mirror real-world variance before scaling.
Governance lag. The agent scaled. The governance framework did not. By month six the agent is making decisions the legal team has not approved and the brand team has not seen. The fix is to build the governance scaffold during the crawl phase, not after the run phase begins.
Talent vacuum. The pilot ran on a small team of enthusiasts. Scaling required ten times the capability and the talent was not in the building. The fix is to start hiring orchestrators in the walk phase, not the run phase.
Vendor lock-in. The pilot ran on one vendor’s stack. Scaling exposed the cost. By the time the procurement team realised, the operational dependence was eighteen months deep. The fix is to design for portability from week one, even when the early pilot uses a single vendor.
Cultural Architecture for Quick Wins
Successful agentic AI transformation transcends technical excellence. It requires organisational momentum that converts sceptics into advocates. This momentum does not emerge spontaneously. It must be orchestrated through strategic quick wins that demonstrate value across stakeholder groups.
The psychology differs from traditional technology deployments. When introducing autonomous decision-makers, you are asking people to cede control. That requires demonstrated trustworthiness, not promised reliability.
Choose quick wins that produce visible business impact and visible team benefit simultaneously. The agent that frees a team member from a task they hated and produces a fifteen per cent uplift on a metric the CFO tracks is the agent that builds coalition. The agent that produces a forty per cent uplift on a metric nobody reads, achieved by automating a role the team valued, is the agent that produces resistance.
Three quick-win patterns work consistently.
Operational drudgery elimination. Find the work everyone hates. Reporting assembly. Campaign trafficking. Cross-system data reconciliation. Automate it. Free the human time for higher-value work. The visible relief is the political capital that funds the next phase.
Capability-creation showcases. Find the capability the team wished it had but could not afford. Real-time competitive intelligence. Multilingual creative variation. Individual-level journey personalisation. Build it. The team’s experience of the new capability is the demonstration that overrides scepticism.
Customer-visible wins. Find the customer experience the team is proudest of and let the agent amplify it. A renewal moment that became more thoughtful. A complaint that got resolved faster. A piece of content that landed more precisely. The customer reaction is the evidence that travels furthest inside the organisation.
Three Scaling Disciplines, Right Now
Before the next initiative kick-off. The discipline is portable. The agents are not.
Score your single most consequential AI pilot against the four selection criteria. If it fails any of the four, the pilot is sized against the wrong opportunity. Reset before the budget commitment.
Map your current AI deployments to the three crawl-walk-run phases. Most enterprise portfolios have multiple agents stuck in shadow mode that should have graduated months ago, and at least one agent running in production that never completed the walk-phase governance work. The mismatch is where most failures originate.
Identify the single quick win that would change the conversation about AI in your organisation if delivered in the next ninety days. Resource it. The political capital generated by a single visible success funds the rest of the programme.
The journey from agentic AI pilot to enterprise-wide transformation differs sharply from traditional technology rollouts. You are introducing autonomous team members that learn and evolve. The discipline of scaling is what separates the one per cent who achieve maturity from the seventy-eight per cent stuck in experimentation.
The McDonald’s voice-AI partnership did not fail because the model could not handle accents. It failed because the scaling discipline did not exist to translate lab performance into operating reality. The same discipline gap kills most enterprise agentic programmes.
Build the discipline before the technology. The technology will improve. The discipline is what compounds.
Keep Reading
That’s all for this week book chapter summary, come back next Monday for the next chapter summary.
The Agentic CMO - Second Edition is available today in hardcover, paperback and ebook.
Disclaimer: The views and opinions expressed in The Agentic CMO, Chronicles of Change and on my social media accounts are my own and do not necessarily reflect the official policy or position of S&P Global.
