Skip to content

Agentic AI: The Production Path Every Enterprise Will Need by 2027

Eleven percent of organisations run agentic systems in production. The attrition between piloting and production is not a model capability problem — it is a distributed systems problem wearing a machine learning label.

Four agent stages in a closed loop with one stage elevated above the others

Key takeaways

  • Production is rare, not impossible. 30 percent are exploring agentic options, 38 percent piloting, 14 percent ready to deploy, 11 percent running in production.
  • The blocker is engineering discipline, not model quality. Transaction design, authorization, observability and cost engineering close the gap.
  • Five decisions gate the build: authority boundary, decision trace, uncertainty routing, reversibility classification, operational ownership.
  • Three platform dependencies cap performance: retrieval readiness, a transactional path to the legacy system of record, and agent identity with short-lived credentials.

Definition

Agentic AI is software that selects its own next action against a goal, invokes tools to take that action, and operates inside a permission envelope defined by someone other than the model. That last clause carries the engineering weight.

Deloitte's 2026 Tech Trends research measures how few organisations have built it: 30 percent are exploring agentic options, 38 percent are piloting, 14 percent have something ready to deploy, and 11 percent are running these systems in production.

The attrition between piloting and production is not a model capability problem. It is a distributed systems problem wearing a machine learning label, and the disciplines that close it are transaction design, authorization, observability and cost engineering.

Agentic AI architecture: what separates an agent from a copilot

The term covers five distinct architectures. They differ in who selects the next action, and that single variable determines the testing strategy, the failure surface and the governance model.

Agent autonomy levels, from assistant to multi-agent orchestration

Five architectures, ordered by who decides the next step
Level Control flow Who selects the next action Dominant failure mode
Assistant Single request, single response A person, on every turn An incorrect answer that nobody validates
Scripted automation Static DAG, fixed branches The author of the workflow Silent breakage when an input schema drifts
Copilot Generate, then human approval gate A person, at every step Approval fatigue degrading into rubber-stamping
Agent Closed loop: observe, decide, act, observe The model, inside an enforced envelope Confident action on a wrong premise
Multi-agent Delegated subtasks with a coordinator Several models plus routing logic Compounding error with no single trace to follow

Only the final two rows are agentic in a way that changes engineering practice. The discriminator is whether the control flow is decided at runtime by the model. A copilot has a static call graph.

An agent does not, which means the set of reachable states cannot be enumerated in advance and the test strategy has to shift from path coverage to invariant checking.

Why runtime control flow changes the failure surface

An assistant returning a poor answer costs a minute of attention. An agent taking a wrong action emits a side effect: a written record, a dispatched message, a posted transaction. Four engineering properties become mandatory at that point.

  • Idempotency on every tool that mutates state. Agent loops retry. Without an idempotency key derived from the task and the intended effect, a retry after a timeout produces a duplicate side effect, and duplicate refunds are the canonical example that reaches an executive.
  • Bounded execution. A step budget, a wall-clock timeout and a token ceiling per task, enforced by the orchestrator rather than requested in a prompt. Unbounded loops are the most common cause of an agent cost incident.
  • Schema validation at the tool boundary. Tool arguments arrive as model output, which means they are untrusted input. Validate against a strict schema, reject on failure, and never pass a model-generated string into an interpreter, a query planner or a shell.
  • A decision trace. Every tool call, its arguments, its result, and the model's stated basis for the next step, captured synchronously.

Fig. 01 The four gates that turn a loop into a controlled system. Every badge on the loop is enforced by the orchestrator, never requested in the prompt.Mactores delivery reference architecture, 2026.

02

The agentic AI production gap: 11 percent, and the reasons behind it

The same Deloitte research reports 42 percent of organisations still developing an agentic strategy roadmap and 35 percent with no formal strategy at all. Those figures explain the funnel shape. A pilot establishes that a capability exists. Production establishes that the capability holds under conditions the pilot never created.


Fig. 02 Adoption concentrates in the pilot stage. The stages are separate states, so the figures do not sum to 100.Source: Deloitte Insights, Tech Trends 2026.

Pilot conditions against production conditions

What changes between a pilot and a production system
Dimension Pilot Production
Input distribution Curated, representative of the happy path Long tail, adversarial, malformed
Concurrency One user, one session Hundreds of sessions, shared rate limits, contention
Failure handling A developer notices and reruns Automatic, bounded, observable, on-call
Evidence A screen recording A queryable trace with retention and access control
Cost Absorbed in an innovation budget A per-task figure that survives a finance review
Model version Whatever was current that week Pinned, with a regression gate before any change

That last row is routinely missed. Provider model versions move under running systems. Pin the version, treat a version bump as a code change, and gate it behind the evaluation suite. Programmes that skip this experience unexplained behaviour drift and spend weeks looking for a code cause that does not exist.

Agentic AI adoption forecasts through 2028

Gartner predicts over 40 percent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. The same firm expects 15 percent of day-to-day work decisions to be taken autonomously by 2028, against none in 2024, and 33 percent of enterprise applications to embed agentic capability, against under 1 percent today.

Both forecasts hold simultaneously. Adoption climbs and the cancellation rate stays high, which describes a market separating on execution rather than on ambition.

03

Five architecture decisions that gate agentic AI production readiness

These consume the design phase on real engagements. None is a model selection question. Each one changes the system architecture, which is why each has to be resolved before build rather than during it.

Decision 1 — Scoping the authority boundary at the tool layer

Express the boundary as an enumerated capability set with parameter constraints, then enforce it where the side effect occurs. "Refund up to a stated ceiling against an order the agent has verified as delivered, with no write access to the customer record" is enforceable. "Operates within appropriate limits" is a sentence that every reader interprets differently.

Enforcement belongs in the tool implementation and in the identity layer, never in the system prompt. A prompt instruction is a preference that any sufficiently unusual input can dislodge. A permission check inside the tool, backed by a token whose scope does not include the forbidden operation, is a control that holds when the model behaves unexpectedly. Design the boundary so that a fully compromised prompt still cannot produce an unauthorised side effect.

Decision 2 — Designing the decision trace for audit

Emit a span per step with a stable schema: task identifier, step index, tool name, argument hash, the retrieved evidence identifiers, the result status, latency, token counts, and the termination reason. OpenTelemetry semantics work well here and give the platform team something their existing stack can already ingest.

Capture it synchronously at decision time. A trace assembled afterwards from application logs answers the question the logs were designed for, which is rarely the question an auditor asks. Decide retention and access control at design time as well, because a trace store containing model inputs is a repository of the same sensitive data the agent reads, and it inherits the same obligations.

Decision 3 — Uncertainty routing and escalation policy

Uncertainty is the steady state, so the design question is where it is routed. Four destinations exist: escalate to a human with the accumulated context attached, narrow the scope and retry, refuse and record the refusal, or proceed and flag for asynchronous review. Each routes operational load to a different team, and the choice belongs to whoever absorbs that load.

Confidence signals from the model itself are weak. Stronger triggers are structural: a tool returned an empty result set, retrieved evidence failed a freshness check, two sources disagreed, the step budget is nearly exhausted, or the requested action falls into a high-reversibility-cost tier. Route on those. Agents without an explicit policy default to proceeding quietly, which is the worst available option.

Decision 4 — Reversibility classification and compensating actions

Classify every tool by the cost of undoing its effect, and let that classification drive the control model rather than sitting beside it in a document.


Fig. 03 The cost of undo rises down the table, and the control model has to rise with it.Mactores delivery reference architecture, 2026.

Anything below the first tier needs a written compensating action implemented and tested before launch, in the manner of a saga. Teams that plan to implement compensation later discover that the business process on the other side never had one.

Decision 5 — Operational ownership and the on-call model

The team that builds an agent is rarely the team paged at three in the morning. Name the operating owner during design and give that owner authority over the escalation policy, the evaluation set and the release gate.

Define what an agent incident is, who is paged, and what the kill switch does: pausing new task admission while allowing in-flight tasks to drain is usually correct, and a hard stop mid-task can leave compensable actions uncompensated.

04

Three platform dependencies that cap agentic AI performance

An agent inherits the properties of the systems it reads from and writes to. Three of them set the ceiling, and the readiness work on each is unglamorous, expensive to defer, and almost never inside the pilot budget.

Retrieval readiness: searchability, freshness and entitlement propagation

A Tech Value Survey found nearly half of organisations naming searchability of data at 48 percent and reusability at 47 percent as obstacles to their AI automation strategy. Both translate directly into agent behaviour.

Three properties matter at the engineering level. Retrieval has to be entitlement-aware inside the query rather than filtered afterwards, so that authorisation constrains the candidate set instead of trimming the result set.

Freshness has to be explicit per source, with an ingestion path matched to the tolerance: change data capture for minutes, batch for overnight. And the index has to expose the structure the agent can navigate, since an agent that cannot filter by effective date or document status will retrieve a superseded policy and act on it with complete confidence.

In the agent programmes reviewed, the technical blocker was almost never the model. Plenty stalled on one of three things: an authority boundary nobody had written down, an evaluation set nobody had built, or a source system nobody had assessed. The programmes that shipped fastest spent their first fortnight on the paperwork most teams treat as a distraction.

Dan Marks VP, Business, Mactores

FAQs

Who should sponsor an agentic AI programme?
The executive who owns the process being changed, rather than the one who owns the technology. Agentic work redesigns how a business function operates, and the sponsor needs authority over the operating model, the headcount plan and the escalation policy. Programmes sponsored purely from technology leadership tend to deliver a working system the receiving function never adopts.
How long does a first agent take to reach production?
For a bounded process with a queryable source system and a defined authority boundary, weeks rather than quarters. The dominant variable is platform readiness rather than the agent. Where data has to be made searchable or a legacy system has to expose a write path, that work sets the schedule and belongs in a separate scope so it cannot hide inside an AI timeline.
What does an agentic AI team look like?
Smaller than most expect. A senior engineer owning the orchestration and the tool contracts, a data engineer owning ingestion and retrieval, a domain expert from the receiving function who validates behaviour and owns the golden set, and a named operating owner from the first design session. Machine learning specialists help and are rarely the constraint, because the hard problems are transactional and operational.
How do agents change the compliance conversation?
They move it earlier and shift its object. With a copilot, compliance reviews a human decision. With an agent, compliance reviews the system's authority model and its evidence, which puts control design in the architecture rather than in a policy written afterwards. Bringing risk and audit into the design phase costs a few sessions and removes most of the later exposure.
What happens to the people whose work the agent takes over?
Answer this before the pilot, because the people who understand the process are the same people whose cooperation the build depends on. Programmes that go well move those people into ownership of the exception queue and the golden set, and say so early. Where the conversation is avoided, knowledge transfer slows in ways that are hard to attribute and expensive to repair.
Can agentic AI run on data that cannot leave our environment?
Usually. Inference can be kept inside a private network and every tool the agent calls can stay behind the same boundary, so sensitive content never transits a public endpoint. The harder questions are procedural. Which team can read the trace store, how long traces are retained, and whether the observability stack the platform team already runs sits inside the boundary or outside it. That last one catches more programmes than the model deployment ever does.
How many agents should a first programme attempt?
One, against a process with a measurable baseline and a tolerant failure mode. The first agent buys the organisational assets: the escalation policy, the trace schema, the golden set, the release gate and the operating rhythm. Every subsequent agent reuses them. Programmes starting with five deliver none of those well and learn nothing transferable.
What belongs in an agentic AI statement of work?
Four terms carry most of the weight. First, the authority boundary written as an enumerated list of permitted actions with parameter constraints. Second, acceptance criteria per phase exit with the signer named. Third, the commercial consequence when a vendor-caused delay moves the date, since a commitment without one is a preference. Fourth, ownership of the golden set and the trace data, both of which belong to the customer whoever writes the code.