← Back to blog

Pilot Sprawl Breaks AI Operating Models. Run a 90 Day Diagnostic

October 8, 2026
Pilot Sprawl Breaks AI Operating Models. Run a 90 Day Diagnostic

The operating model for AI that works is one where accountable human and AI teams deliver orchestrated outcomes on a governed, shared platform, not one where pilots multiply without a backbone to support them. Before building anything, run an AI operating diagnostic to see where you actually stand. Then centralize high-risk governance immediately while piloting domain embedding in one or two business units, rather than waiting for a perfect enterprise-wide design.


TL;DR:

  • High-risk governance should be centralized immediately, while piloting domain embedding in select business units to ensure scalability.
  • Most organizations fall into siloed, CoE, hub-and-spoke, center for acceleration, or embedded models, chosen based on data maturity and regulation.
  • Clear accountability artifacts such as decision logs and role charters are crucial to prevent decision-rights drift in AI workflows.
  • Building shared infrastructure like model registries and observability tools is essential for scalable, governed AI capability across enterprise units.
  • An AI operating diagnostic, phased lighthouse projects, and robust stakeholder engagement are key steps to effectively scale AI initiatives.

Haiphai
haiphai.com
Turn AI Bottlenecks Into Progress
HaiPhai works with life sciences teams to identify operational bottlenecks and tailor AI solutions to streamline critical processes.
See how HaiPhai helps

Table of Contents

What an AI operating model is and why it decides whether pilots compound

An AI operating model is the combined design of people, governance, data and platform infrastructure, delivery mechanisms, and measurement that determines whether AI initiatives turn into repeatable enterprise value or stay stuck as disconnected experiments. It is not a tool stack or a single team: it is the operating logic that connects a use case to a governed, scalable outcome.

Each component carries weight on its own:

  • People and accountability: who owns a decision when a human and an AI system disagree, and who is accountable for the outcome.
  • Governance: the risk tiers, review gates, and escalation paths that keep deployment speed from outrunning oversight.
  • Data and platform: the shared infrastructure, from data products to model registries, that lets capability compound rather than get rebuilt per team.
  • Delivery: how use cases move from idea to production, and who has authority to approve that move.
  • Measurement: the KPIs that tell leadership whether AI is producing value or just activity.

McKinsey's research on organizational rewiring found that top-performing organizations centralize high-risk functions like data governance and compliance while distributing talent and adoption through hybrid models, a structural choice tied directly to stronger bottom-line impact from AI investment. A company that lets every business unit manage its own model risk independently tends to produce inconsistent controls and duplicate infrastructure. A company that locks all AI development inside a single center of excellence tends to produce technically sound pilots that never reach the business units that need them.

Common AI operating model patterns: catalog, trade-offs, and fit criteria

Most organizations fall into one of five recognizable patterns, and the right one depends less on ambition and more on data maturity, regulatory exposure, and how fast you need to show value.

  1. Siloed: Individual business units build and run their own AI capability with little central coordination. This model moves fast on narrow use cases but produces duplicated infrastructure, inconsistent risk controls, and shadow IT. It tends to suit early-stage experimentation in low-risk, low-regulation functions.
  2. Center of excellence (CoE): A central team owns AI strategy, standards, and often delivery, consulting with business units on demand. This model is strong on consistency and governance but can become a bottleneck when demand for AI capability outpaces the central team's capacity, especially in large, multi-unit enterprises.
  3. Hub-and-spoke: A central hub sets standards, owns shared platform components, and provides specialist talent, while spoke teams embedded in business units own delivery and domain context. Bain & Company's framing of the operating model for the age of AI describes this kind of structure as a way to move from managing hierarchy to orchestrating outcomes, since the hub can hold governance steady while spokes adapt to local workflows. This model fits organizations with moderate data maturity and real regulatory stakes, but it requires disciplined decision rights so spokes do not quietly diverge from hub standards.
  4. Center for acceleration: A time-boxed, cross-functional unit tasked with rapidly scaling a small number of high-value use cases, then handing them to permanent owners. This suits organizations that need fast, visible wins to build internal momentum, though it risks leaving a governance and maintenance gap once the acceleration team disbands if handoff planning is weak.
  5. Embedded or AI-native: AI capability and judgment are built directly into functional teams, with a thin central layer for shared infrastructure and risk oversight. This fits organizations with high data maturity and a strong existing governance culture, where the main constraint is adoption speed rather than control. It is the hardest model to retrofit into a legacy organization because it assumes judgment-sharing skills that most teams have not yet built.

Regulatory sensitivity should weigh heavily in the choice: a life sciences or financial services organization has far less room for a loosely governed siloed model than a retail marketing function does. Speed-to-value considerations usually argue for starting narrower (CoE or center for acceleration) and evolving toward hub-and-spoke or embedded models as governance muscle matures.

Structure, teams, and accountabilities: outcome-oriented design

Org charts show reporting lines. They do not show who is accountable when an AI-assisted decision goes wrong, which is the question that actually matters once agents and humans share work. Bain's research argues that AI shifts the core leadership challenge from managing capacity to creating clarity, which means accountability artifacts have to do work that job titles used to do.

Build these artifacts directly:

  • An accountability chart that maps each workflow outcome, not each role, to a named accountable person.
  • A decision log that records where a human overrode an AI recommendation and why, so you can audit judgment quality over time.
  • Role charters that spell out what a model owner, an agent steward, or a human reviewer can approve without escalation.
  • Escalation paths that specify exactly when a disagreement between a human and an AI system gets kicked upstairs, and to whom.

Defining decision rights matters more than it sounds like it should. Bain's analysis notes that these artifacts function as living governance inputs precisely because they prevent decision-rights drift, the slow process by which nobody remembers who is actually accountable once an AI agent starts handling a task that used to sit clearly with one person.

Pro Tip: Review your accountability chart every quarter for the first year: decision rights drift fastest right after a new AI capability goes live, not months later.

Talent, roles, and capability: designing for judgment and orchestration

The shift AI forces on talent strategy is from executing tasks to exercising judgment, which means hiring and development plans built around task completion need a rework. New and evolving roles are emerging to fill that gap.

  • Agent steward: owns the behavior and performance of a specific AI agent or agent cluster in production.
  • Model owner: accountable for a model's performance, drift, and retraining cadence, distinct from whoever built it.
  • Data product owner: treats a dataset as a product with its own roadmap, quality bar, and users.
  • TEVV engineer: runs the test, evaluation, verification, and validation work that keeps high-risk systems safe to operate.

Building judgment at this speed usually means more than a training module. Scenario-based workshops that walk teams through real decision points, paired with rotational programs that move people between model development and domain delivery, build the kind of applied judgment that no amount of reading can substitute for.

On sourcing, a short checklist helps:

  • Stand up an internal talent marketplace so people with emerging AI skills can move toward the teams that need them.
  • Hire externally only for roles that require deep technical specialization you cannot build in time, like TEVV engineering.
  • Tie incentives to outcome quality and judgment, not just deployment speed, so people are not rewarded for shipping fast and governing loosely.

Governance and responsible AI: operationalizing NIST AI RMF and practical controls

Governance built around outcomes, not just restrictions, turns into a competitive advantage rather than a brake on speed. The World Economic Forum's reporting on AI business governance frames this plainly: organizations that center governance on integrity, accountability, transparency, and resilience scale AI more safely and keep stakeholder trust intact, which is itself a competitiveness factor.

The clearest operational scaffold available is the NIST AI Risk Management Framework, organized into four functions you can map directly onto your operating model:

  • GOVERN: establish the policies, roles, and culture that make risk management a standing practice, not a one-time checklist.
  • MAP: identify the context, use cases, and potential impacts of each AI system before it goes live.
  • MEASURE: run continuous testing and evaluation to quantify risk and performance, not a single pre-launch review.
  • MANAGE: allocate resources to respond to identified risks and track mitigation over time.

Concrete components to stand up now include an AI inventory that tracks every system in production, risk tiers that route high-stakes systems to stricter review, continuous TEVV rather than a pre-deployment gate, and independent reporting lines so AI oversight does not report through the same executive who owns deployment targets.

One senior leader with real independence and board reporting authority, covering AI, data, and security together, is associated with stronger outcomes than splitting that authority across silos, according to WEF and Wipro research on building operating models for the intelligence era, which also found that most organizations expect hybrid human and AI teams but few have formalized the operating model to support them. Our own experience with biotech operations reinforces how much of this is about formal structure, detailed in what AI governance platforms really mean for biotech teams.

Unified AI governance structure illustration

Data, platform, and architecture: building the cognitive layer and scalable tech foundation

A good operating model falls apart on bad infrastructure, which is why the platform layer deserves the same design discipline as governance and talent. McKinsey's work on agentic AI describes the need for a cognitive layer: shared infrastructure that lets AI agents act across processes rather than inside single-task silos, which requires governed autonomy limits, interoperability between systems, and shared memory so agents do not relearn context every time.

Treating data as a managed product, not a one-off extraction, is central to making that layer work. McKinsey's guidance on scaling data products recommends assigning product ownership and tracking data products against business-style KPIs, which keeps governance from becoming a narrow compliance exercise disconnected from the return it is supposed to protect.

Operational checklist items worth building early:

  • A model registry that tracks every model version, owner, and deployment status in one place.
  • Observability tooling that flags drift, latency, and failure patterns before they become incidents.
  • Cost tagging down to the use case level, so finance can see which AI investments are paying off.
  • A shared prompt and agent library so teams are not rebuilding the same capability from scratch.

Pro Tip: Fund the platform layer as shared infrastructure, not as a line item inside the first pilot's budget: pilots that carry the full weight of platform costs almost always look like failures on paper.

Decision framework to choose and evolve your AI operating model

Choosing a model is a trade-off exercise, and getting it wrong usually shows up within the first year as either runaway risk or stalled adoption.

A simple criteria matrix helps structure the decision:

  • Data maturity: low maturity favors a CoE or center for acceleration; high maturity supports hub-and-spoke or embedded models.
  • Regulatory risk: high-risk sectors need centralized controls regardless of which delivery model you pick.
  • Speed-to-value need: urgent, narrow wins favor a center for acceleration; durable, broad capability favors hub-and-spoke.
  • Talent availability: scarce specialized talent (TEVV, data product ownership) argues for a central hub that can be shared.
  • Budget structure: centralized budgets support CoE models; distributed budgets support embedded models, with hub-and-spoke sitting between the two.

Track a small set of KPIs once you have chosen a direction:

  1. Value capture, measured against the outcome hypothesis set at the start of each use case.
  2. Deployment velocity, the time from approved use case to production.
  3. TEVV coverage, the share of production systems under continuous testing and evaluation.
  4. Incident rate, tracking both model failures and operational handoff failures between humans and AI systems.

WEF and Wipro's research found that operational failures, not model accuracy problems, are the more common cause of production issues, usually traced to unclear human-AI handoffs or weak TEVV discipline rather than the model itself. The most common transition path moves organizations from CoE toward hub-and-spoke as data maturity and governance culture mature, with the central team shrinking into a platform and standards function as spokes take on more delivery autonomy.

Implementation roadmap: diagnostic, lighthouse projects, platform build, and scale mechanics

A staged rollout beats a big-bang redesign because it lets you test governance assumptions on real use cases before committing enterprise-wide budget.

  1. Run an AI operating diagnostic across people, governance, data, and delivery to find your actual starting point rather than an assumed one. Resources like biotech workstream prioritization strategies are useful at this stage for picking where to focus first.
  2. Design one or two lighthouse projects that deliberately stress both value delivery and governance, not just technical accuracy. McKinsey's data leader's operating guide recommends this approach specifically because it validates whether your operating model can manage risk at scale, not only whether the use case works in isolation.
  3. Build the shared platform and governance backbone in parallel with the lighthouse projects, since waiting until after the pilots succeed usually means rebuilding infrastructure under time pressure.
  4. Scale with funding guardrails tied to KPI performance, not a fixed annual budget, so teams that hit value and governance targets get resourced ahead of teams that only hit deployment targets.

Change management runs alongside every stage: a regular communication cadence, visible incentives for judgment quality rather than just speed, and funding rules that reward teams for flagging governance gaps early rather than hiding them. A documented example of this staged approach in practice is in our biotech playbook for redesigning clinical workflow with governed AI.

Integration with existing enterprise architecture and legacy systems

An AI operating model has to work alongside systems that were never designed for it, and pretending otherwise is one of the most common causes of stalled rollouts. Legacy enterprise resource planning, clinical, and compliance systems rarely expose the clean APIs that a modern agentic architecture assumes, which means integration work often takes longer than model development itself.

The practical approach is to treat legacy integration as its own workstream with its own owner, not an afterthought inside a use case team's sprint. A thin integration layer, built once and shared across use cases, prevents every new AI initiative from reinventing its own connection to the same legacy database or document system.

Data quality is usually the real blocker, not the connection itself: legacy systems often hold inconsistent, duplicated, or poorly structured data that needs cleanup before any AI system can use it reliably. Budgeting time and ownership for that cleanup, rather than assuming it will happen organically during a pilot, keeps timelines honest. For regulated environments specifically, a useful reference point is what it takes for an AI tool to actually meet compliance obligations, since the same review applies to any integration touching protected data.

Cross-functional collaboration mechanisms and stakeholder engagement strategies

AI operating models fail quietly when IT, legal, compliance, and business units each assume someone else owns coordination. The fix is structural, not aspirational: standing cross-functional forums with real decision authority, not just information-sharing meetings.

A practical mechanism is a recurring review board that includes a business sponsor, a technical owner, a compliance or legal representative, and someone accountable for data quality, meeting on a fixed cadence to approve use cases moving between stages. This group should have authority to pause a deployment, not just flag concerns.

Stakeholder engagement works best when it starts before a use case is built, not after. Business unit leaders who are consulted on problem framing early tend to adopt the resulting tool faster than those who receive a finished system and are asked to use it. Documenting outcome hypotheses jointly with business stakeholders, a practice covered in aligning AI tools with clinical goals, keeps technical and business teams working from the same definition of success.

Scalability challenges and methods to address scaling AI across diverse business units

What works in one business unit rarely transfers cleanly to another, because data maturity, regulatory exposure, and existing workflows differ sharply across an enterprise. A model trained on one unit's data often underperforms elsewhere, and a governance process built for a low-risk function can be dangerously thin for a high-risk one.

The most reliable fix is building shared infrastructure once, at the platform layer, while letting each business unit customize the workflow layer around its own context. A shared model registry, a common data product framework, and a single risk-tiering system prevent every unit from building its own version of the same control, while local teams still get to decide how AI fits into their specific process.

Hub and spoke AI scaling model

Staffing is often the quieter scaling constraint: specialized roles like TEVV engineers and data product owners are scarce, and spreading them thin across every business unit dilutes their effectiveness. A hub-and-spoke structure, where these specialists sit in a central hub and rotate support across units, tends to scale further than hiring a dedicated specialist into every team. Our executive biotech workflow optimization guide walks through how this plays out when scaling AI-enabled workflows across clinical, regulatory, and operational functions at once.

Change management specific to cultural and organizational shifts caused by AI adoption

The hardest part of an AI operating model is rarely the technology: it is convincing people that their judgment is still valuable when a system can do part of their job faster. Change management for AI adoption has to address that directly, not paper over it with generic communication plans.

Transparency about what is changing and what is not builds more trust than reassurance alone. Telling a team exactly which tasks an AI system will take over, which decisions still require human sign-off, and how performance will be measured going forward removes the ambiguity that fuels resistance.

Incentive structures need to shift alongside the work itself. Rewarding people for catching an AI system's mistake or for escalating a governance concern, rather than only for shipping deployments quickly, signals that judgment quality matters as much as speed. Rotational exposure, where team members spend time inside both the technical build and the domain delivery side, tends to build the kind of comfort with AI-assisted work that a single training session cannot. Our life sciences AI integration benefits overview documents how that cultural shift tends to unfold once teams move past the initial adjustment period.

Practitioner perspective: common pitfalls and leadership trade-offs

The most common mistake is treating AI governance as a separate workstream from the business unit that owns the risk, which guarantees the two drift apart within months. A close second is treating AI as a bolt-on to existing processes rather than redesigning the workflow around it, which produces tools nobody actually uses. The third, and most damaging long-term, is underfunding TEVV and data product work because it does not show up as a visible deliverable the way a new agent or dashboard does.

The real trade-off every leader faces is speed against control, and centralization against autonomy. There is no universally correct answer: it depends on regulatory exposure and data maturity. What is consistent is that a diagnostic run in the first 90 days, before committing budget to a specific structure, reduces the risk of locking into the wrong model for the next three years.

— John

The HaiPhai operating partnership and AI Velocity Diagnostic

Building this kind of operating model from scratch takes specialized talent that is hard to hire fast, especially in regulated life sciences environments where regulatory drafting and clinical site activation carry real timeline risk. We work as an operational partner, starting from a client's strategic goals and backtracking to find where AI can remove operational bottlenecks.

Haiphai

Our work spans:

  • The operating partnership, embedding expert teams directly into live operations.
  • The AI Velocity Diagnostic, an initial assessment of where operational time is being lost.
  • Clinical, scientific, and regulatory operations redesign, tailored to regulatory contexts.
  • Governed automation, to balance speed gains with oversight.

A partner model tends to make sense when internal teams lack the specialized bandwidth to redesign workflows and build governance at the same time; building entirely in-house can work when that capacity already exists, especially with expertise in revenue cycle management AI. If you are weighing which route fits your situation, our decision framework for choosing a fractional AI partner versus an internal team lays out the considerations. To see where your own operation stands, start with our AI Operating Maturity Diagnostic.

FAQ

What is the AI operating model?

An AI operating model is the combined design of people, governance, data and platform infrastructure, delivery processes, and measurement that determines whether AI initiatives scale into reliable enterprise value. It covers who is accountable for outcomes, how risk is managed, and how capability is shared across business units rather than rebuilt in silos.

What are the main types of AI operating models?

The most common patterns are siloed, center of excellence, hub-and-spoke, center for acceleration, and embedded or AI-native models. Each fits a different combination of data maturity, regulatory risk, and speed-to-value need, and most organizations evolve from a more centralized model toward a more distributed one as governance capability matures.

What are the four types of operating models?

Definitions vary across sources, but a common framing used for enterprise AI groups models into siloed, centralized (center of excellence), hybrid (hub-and-spoke), and fully embedded or distributed structures. The right choice depends on an organization's regulatory exposure, data maturity, and how fast it needs to show results.

What is the 30% rule in AI?

Readers encountering that phrase should check the specific source using it, since the term is not standardized across the frameworks referenced here.

How do I know when to evolve our AI operating model?

Signals that it is time to evolve include rising incident rates between human and AI handoffs, deployment velocity stalling despite available budget, or business units building shadow AI capability outside governance. Tracking value capture, deployment velocity, and TEVV coverage against baseline targets is the clearest way to catch these signals early.

Sources