Enterprise AI Strategy and Architecture: The Three-Pillar Framework
- Last updated
- 17 min read
Key takeaways
- Three pillars divide the work by responsibility: the evidence layer (what systems may read), the action layer (what they may change), and the execution layer (what decides the next step).
- Three interface contracts hold the pillars together: retrieval, action and telemetry. Each needs a named owner on both sides and a test that can fail. In 19 of 31 programmes reviewed across 2024 and 2025, a quarter was lost to work no single team owned.
- Data readiness is the most common obstacle. In Deloitte's 2025 Tech Value Survey, 48 percent of organizations named data searchability and 47 percent named reusability as barriers to their AI automation strategy.
- Sequencing by workload puts two systems live in twelve months. Pick the first workload against six criteria, capture its baseline, and build each pillar only as far as that workload's contract requires.
An enterprise AI strategy is an architecture document that happens to be read by executives. It has to answer what the systems may read, what they may change, and what decides between those two, and it has to answer each one concretely enough that an engineer can implement it and an auditor can inspect it. Most strategy decks answer none of the three. They name ambitions, list use cases, and leave every interface between the moving parts undefined, which is why the funding survives contact with the board and the architecture does not survive contact with the existing environment.
The framework below organizes the work into three pillars and, more usefully, into the contracts between them. Deloitte’s 2026 Tech Trends research measures the gap that framework is trying to close: 30 percent of organizations are exploring agentic options, 38 percent are piloting, and only 11 percent are running these systems in production.
Why enterprise AI strategy fails at the architecture layer rather than the ambition layer
The failure is rarely a bad idea. It is three funded programs running in parallel with no agreement on what passes between them.
A data team builds a lakehouse and declares success on ingestion volume. An application team modernizes a service and declares success on latency. An AI team builds an agent and declares success on a demonstration. Each program hits its own targets. The agent still cannot answer a question correctly, because the data it needs is present but not filterable by entitlement, and it still cannot complete a task, because the system of record exposes a read API and no transactional write path. Nobody failed. The interfaces were never specified, so nobody was accountable for them.
Gartner puts more than 40 percent of agentic AI projects on track to be canceled by the end of 2027, and names cost escalation, unproven value, and weak risk controls as the reasons. Treat all three as downstream effects rather than causes.
Costs escalate when an execution layer retries against an unreliable action layer. Value stays unclear when no baseline was captured before the work started. Risk controls are inadequate when authorization was designed inside one pillar and assumed by the other two.
An enterprise AI strategy earns its name when it specifies the contracts. The pillars are the easy part.
The three pillars of an enterprise AI strategy
Separate the environment by what each layer is responsible for, rather than by which team currently owns it.

Pillar one: the evidence layer
This is the data foundation, defined by what an AI system can actually do with it rather than by how much of it exists. Volume is irrelevant. Five properties decide the ceiling.
- Searchability. Content has to come back both by meaning and by exact token, because enterprise questions hinge on part numbers, policy codes, and case references that vector similarity handles poorly. Deloitte’s 2025 Tech Value Survey puts numbers on how widely this bites: 48 percent of organizations named data searchability and 47 percent named reusability as obstacles to their AI automation strategy. Those two numbers describe a data problem, not a model problem.
- Freshness, declared per source. Every source needs a stated staleness tolerance, and that tolerance determines the ingestion architecture rather than the other way around. Minutes means change data capture. Overnight means batch. Deciding this late means rebuilding the pipeline.
- Entitlement propagation. Access rights have to travel with the content into the index and constrain the query itself. Filtering results after retrieval lets forbidden material occupy the candidate set, which quietly degrades answer quality for exactly the users with the least access.
- Provenance. Every retrieved passage needs a stable identifier that survives reindexing, so an answer can cite something a human can open and a reviewer can trace.
- Deletion. A removal request has to propagate to the index, the caches, the derived summaries, and the evaluation fixtures. Designing this after launch is one of the least pleasant retrofits in the field.
Pillar two: the action layer
Reading is half a system. The action layer covers every system of record an AI system may change, and it is where enterprise programs discover their real timeline.
- A transactional interface, not a reporting one. Most systems of record expose something that was built for extracts. An AI system needs typed operations with defined preconditions and explicit failure semantics.
- Idempotency on every mutating operation. Retries are normal, not exceptional. Each mutating call therefore needs a key the caller can reproduce from the same request, so that a repeat arriving after a network timeout resolves to the original outcome rather than a second one. Duplicated payments are the version of this failure that ends up in a board deck.
- Business rules that live somewhere readable. Where the rules sit in stored procedures and triggers written by people who have left, no interface in front of them is trustworthy until the rules are extracted and named.
- A compensating path for anything that cannot be rolled back. Classify operations by the cost of undoing them and implement the compensation before launch, because the business process on the far side frequently never had one.
- An access model that expresses delegation. The action layer has to distinguish an action taken by a person, an action taken on a person’s behalf, and an action taken autonomously by a system. Environments that cannot express that distinction cannot pass a controls review once agents are involved.
Pillar three: the execution layer
This is where models, orchestration, and tools live, and it is the thinnest of the three despite receiving most of the attention and most of the budget.
- A declared authority envelope. An enumerated set of permitted operations with parameter limits, enforced where the effect occurs rather than requested in a prompt.
- Bounded execution. A ceiling on iterations, wall-clock time, and token spend per task, applied by the orchestrator. Unbounded loops are the most common cause of a cost incident.
- Defined behavior under uncertainty. Escalate with context, narrow and retry, refuse and record, or proceed and flag. Each choice loads a different team, so the choice belongs to whoever carries that load.
- Emitted decision evidence. A structured record of what was read, what was called, in what order, and why the loop stopped, written at the moment of the decision.
- A named operator. Somebody accountable for the escalation policy, the release gate, and the pager, identified during design rather than after go-live.
The interface contracts that hold the three pillars together
This is the part a strategy document usually omits and the part that determines whether the program works. Each contract is a short, testable specification owned jointly by two pillars.

Write these before build starts. Three properties make them work.
- They are testable rather than aspirational. “Retrieval will be entitlement-aware” is a sentence. “Recall at ten stays above the agreed threshold for every entitlement class, measured weekly against a labeled set” is a contract, because it can fail.
- They have a named owner on each side. A contract owned by everyone is owned by nobody, and interface defects are exactly the class that goes unclaimed between teams.
- They version independently of the pillars. The evidence layer can be rebuilt without renegotiating the contract, provided the guarantees hold. That is the point of specifying guarantees rather than implementations.
in the model version explicitly and treat a version change as a code change that must clear the evaluation suite. Provider versions move, and behaviour drift traced to an unpinned model has cost more than one team a fortnight of debugging that found nothing in their own code.
Four cross-cutting concerns that no single pillar owns
Four things span all three pillars, and each one fails badly when assigned to a single team.
- Identity and delegated authorization. A person’s authority has to propagate through the execution layer into the action layer without being widened, and an autonomous system needs its own scoped identity with short-lived credentials. A shared service account with standing privilege is convenient during a pilot and indefensible in review.
- Evaluation. A labeled set of real tasks with known correct outcomes, owned by the business function rather than the engineering team, wired into the release path so that no change ships without a scored comparison.
- Observability. Operational telemetry answers how the system behaved. A decision record answers what it relied on and why. Both are needed, and only the second satisfies a reviewer.
- Unit economics. Spend per unit of completed work, held as a distribution rather than an average, and set against what the current process costs once people, tooling, and rework are counted. Watch the upper percentiles, since a thin band of tasks that loop before failing tends to account for a startling share of the invoice.
How to sequence an enterprise AI strategy without stalling
The common mistake is treating the pillars as a maturity model and finishing one before starting the next. That produces a two-year data program with no demonstrated value and a canceled AI budget. The correct unit of sequencing is a workload, not a layer.

The minimum viable contract per pillar
For a single first workload, each pillar needs only enough to satisfy its contract for that workload.
- The evidence layer needs the sources that workload touches, entitlement-filtered and within a stated freshness bound. It does not need the whole environment governed.
- The action layer needs the specific operations that workload performs, typed, idempotent, and compensable. It does not need the legacy system retired.
- The execution layer needs an authority envelope covering those operations, a labeled task set, and a decision record. It does not need a platform.
Scoping this way turns a multi-year program into a sequence of workloads, each of which extends the contracts a little further. The second workload is materially cheaper than the first, because the contracts, the evaluation harness, and the operating rhythm already exist.
Where parallel work is safe, and where it is not
Parallel work is safe inside a pillar and dangerous across an unspecified interface. Ingestion pipelines, index construction, and entitlement synchronization can proceed alongside tool development, provided the retrieval contract is agreed first. Building an execution layer against an action layer whose semantics are still being negotiated produces work that gets discarded.
The reliable signal that sequencing has gone wrong is an integration date that keeps moving while every individual team reports being on track.
“Nineteen of the 31 enterprise programmes we reviewed across 2024 and 2025 lost a quarter to something no single team owned. An entitlement rule nobody could express inside a query. An operation that could be read and never written. In each case the layer that stalled the schedule was the one with the smallest budget.”
03
How to choose the first workload
Selection matters more than any technology decision on this list, because the first workload buys the organizational assets that every subsequent one reuses. Score candidates against six criteria.
| Criterion | Strong candidate | Weak candidate |
|---|---|---|
| Baseline |
Volume, handling time, and error rate already measured |
No instrumentation, and a baseline that would need constructing |
| Failure tolerance |
Errors are recoverable and visible |
Errors are irreversible or externally visible on first contact |
| Evidence readiness |
Sources are identifiable and entitlements expressible |
Knowledge lives in people, email threads, and undocumented spreadsheets |
| Action surface |
Few operations, on a system with a usable interface |
Many operations across systems with no write path |
| Owner |
A named business owner who wants the change | An owner who was volunteered |
| Volume | Enough to matter, small enough to shadow safely |
Trivial, or too critical to run in shadow |
A candidate that scores strongly on baseline, failure tolerance, and owner will teach the organization more than a technically elegant one that scores strongly on nothing else.
04
Operating model: who owns each pillar and which decisions they hold
Ambiguous decision rights produce the interface defects described above. Name the holder of each decision explicitly.
How to measure an enterprise AI strategy
Measure each pillar on its own terms, and measure the program on business outcome. Mixing the two produces dashboards nobody trusts.
| Decision | Leading indicator | Lagging indicator | Warning sign |
|---|---|---|---|
| Evidence |
Recall against a labeled set, per entitlement class |
Share of answers with a verifiable citation |
Recall healthy in aggregate, poor for restricted users |
| Action | Share of required operations exposed as typed, idempotent tools |
Reversal rate on completed |
Falling reversal rate alongside rising manual correction |
| Execution |
Unassisted completion rate |
Cost per completed unit of work at p50 and p95 | Completion rising while escalation falls and complaints rise |
| Program |
Number of workloads with all three contracts satisfied |
Change in the business baseline that was captured before launch |
Activity metrics substituting for outcome metrics |
Capture every baseline before the first workload goes live. A baseline reconstructed afterward is a number the business will not accept, and the argument about attribution will consume more time than the measurement would have.
Failure modes in enterprise AI strategy
Five patterns account for most of the wreckage.
- Three parallel programs with no interface contract. Every team succeeds and the system does not work. This is the dominant failure and the reason the contracts sit at the center of this framework.
- The maturity-model trap. A sequential program that finishes the data layer before starting anything else, and loses its funding before demonstrating value.
- Authority designed in one pillar and assumed in the others. Security models the access, the AI team assumes it holds, the application team exposes a broad operation, and the review finds the gap.
- Evaluation owned by engineering. The team building the system also grades it, so the grading drifts toward what the system does well.
- A pilot chosen for demonstrability. An impressive use case with no baseline, no owner, and no reusable assets, which proves the technology works and teaches the organization nothing it can apply again.
A twelve-month sequencing plan
Here’s how a practical plan can turn out to be:
- Months 1 to 2. Select the first workload against the six criteria. Capture its baseline. Draft the three interface contracts and name an owner on each side of each one.
- Months 3 to 5. Build to the minimum viable contract in each pillar. Entitlement-aware retrieval over the sources that workload touches. Typed, idempotent operations for the actions it performs. An authority envelope, a labeled task set, and a decision record in the execution layer.
- Months 6 to 7. Shadow production. Compare against the captured baseline. Close the controls review using real decision records rather than a design document. Hand the named operator a runbook and a stop mechanism they have exercised themselves.
- Months 8 to 10. Second workload, deliberately chosen to extend one contract rather than all three. This is where the cost curve of the program becomes visible, and where the reusability of the first workload’s assets gets tested.
- Months 11 to 12. Consolidate. Fold what the two workloads proved into a standing operating model: decision rights, release gates, cost ceilings, and the evaluation cadence. Publish the contracts as internal standards so the third workload does not renegotiate them.
That plan produces two systems in production and a repeatable method inside a year. A plan that spends the same twelve months building a platform produces neither.
FAQs
Who should own an enterprise AI strategy?
A single accountable executive with authority across all three pillars, which in practice means someone who can direct data engineering, application engineering, and the receiving business function. Strategies owned inside one of the three consistently optimize that pillar and under-specify the interfaces. Where no such role exists, the strategy needs a standing forum with named decision rights rather than a coordinating committee with none.
How long before an enterprise AI strategy produces measurable value?
For a well-chosen first workload, one to two quarters to production and one more to a defensible measurement against the baseline. Programs that report nothing for a year are usually building a platform rather than delivering a workload, and platform-first sequencing is the most reliable way to lose executive patience before the first outcome lands.
How does this framework apply if the environment is already mostly cloud-native?
The pillars are unchanged and the effort redistributes. A modern environment usually satisfies more of the action contract on day one, because transactional interfaces exist and access models are expressible.
The evidence layer is rarely as ready as teams expect, since being in the cloud says nothing about entitlement propagation, freshness declarations or provenance. Run the same contract review and expect the surprises to concentrate in the first pillar.
What belongs in the strategy document itself, as opposed to the architecture?
Three things: the sequence of workloads with the business outcome each targets, the decision rights table, and the three interface contracts stated as guarantees. Everything else is architecture and should live where architecture lives, because a strategy document that restates implementation detail ages within a quarter and stops being read.
How should an enterprise AI strategy handle model and platform choice?
As a reversible decision made late. Keep tool contracts, prompts, the labeled task set, and the evaluation harness in your own repository, and treat the model and the managed platform as replaceable execution. Programs that decide the platform first tend to arrange the architecture around a vendor’s abstractions and discover the cost of that arrangement at renewal.
What is the right level of central control?
Central ownership of the contracts, the decision rights, and the evaluation standard. Local ownership of implementation. Centralising implementation produces a queue, and a platform team that becomes a bottleneck is the second most common way these programs stall. Federating the contracts produces incompatible interfaces, which is the first.
How should an enterprise AI strategy be funded?
Per workload, with a cost ceiling attached and a business owner who has signed the target. Program-level funding with no per-workload accountability disconnects spend from outcome, which is precisely the condition under which cost escalation and unclear value appear together in a cancellation review.
What changes when regulators are involved?
The sequence hardens rather than changes. Evidence requirements move earlier, the decision record becomes a design input rather than an operational nicety, and the authority envelope needs sign-off before build rather than before launch.
Bringing risk and audit into the contract-writing sessions costs a few meetings and removes most of the exposure that would otherwise surface during the review.
AI agents for apps
Ready to move past the pilot?
Talk to delivery