Agentic AI: The Production Path Every Enterprise Will Need by 2027
Eleven percent of organisations run agentic systems in production. The attrition between piloting and production is not a model capability problem — it is a distributed systems problem wearing a machine learning label.
- Last updated
- 17 min read
Key takeaways
- Production is rare, not impossible. 30 percent are exploring agentic options, 38 percent piloting, 14 percent ready to deploy, 11 percent running in production.
- The blocker is engineering discipline, not model quality. Transaction design, authorization, observability and cost engineering close the gap.
- Five decisions gate the build: authority boundary, decision trace, uncertainty routing, reversibility classification, operational ownership.
- Three platform dependencies cap performance: retrieval readiness, a transactional path to the legacy system of record, and agent identity with short-lived credentials.
Definition
Agentic AI is software that selects its own next action against a goal, invokes tools to take that action, and operates inside a permission envelope defined by someone other than the model. That last clause carries the engineering weight.
Deloitte's 2026 Tech Trends research measures how few organisations have built it: 30 percent are exploring agentic options, 38 percent are piloting, 14 percent have something ready to deploy, and 11 percent are running these systems in production.
The attrition between piloting and production is not a model capability problem. It is a distributed systems problem wearing a machine learning label, and the disciplines that close it are transaction design, authorization, observability and cost engineering.
Agentic AI architecture: what separates an agent from a copilot
The term covers five distinct architectures. They differ in who selects the next action, and that single variable determines the testing strategy, the failure surface and the governance model.
Agent autonomy levels, from assistant to multi-agent orchestration
| Level | Control flow | Who selects the next action | Dominant failure mode |
|---|---|---|---|
| Assistant | Single request, single response | A person, on every turn | An incorrect answer that nobody validates |
| Scripted automation | Static DAG, fixed branches | The author of the workflow | Silent breakage when an input schema drifts |
| Copilot | Generate, then human approval gate | A person, at every step | Approval fatigue degrading into rubber-stamping |
| Agent | Closed loop: observe, decide, act, observe | The model, inside an enforced envelope | Confident action on a wrong premise |
| Multi-agent | Delegated subtasks with a coordinator | Several models plus routing logic | Compounding error with no single trace to follow |
Only the final two rows are agentic in a way that changes engineering practice. The discriminator is whether the control flow is decided at runtime by the model. A copilot has a static call graph.
An agent does not, which means the set of reachable states cannot be enumerated in advance and the test strategy has to shift from path coverage to invariant checking.
Why runtime control flow changes the failure surface
An assistant returning a poor answer costs a minute of attention. An agent taking a wrong action emits a side effect: a written record, a dispatched message, a posted transaction. Four engineering properties become mandatory at that point.
- Idempotency on every tool that mutates state. Agent loops retry. Without an idempotency key derived from the task and the intended effect, a retry after a timeout produces a duplicate side effect, and duplicate refunds are the canonical example that reaches an executive.
- Bounded execution. A step budget, a wall-clock timeout and a token ceiling per task, enforced by the orchestrator rather than requested in a prompt. Unbounded loops are the most common cause of an agent cost incident.
- Schema validation at the tool boundary. Tool arguments arrive as model output, which means they are untrusted input. Validate against a strict schema, reject on failure, and never pass a model-generated string into an interpreter, a query planner or a shell.
- A decision trace. Every tool call, its arguments, its result, and the model's stated basis for the next step, captured synchronously.

02
The agentic AI production gap: 11 percent, and the reasons behind it
The same Deloitte research reports 42 percent of organisations still developing an agentic strategy roadmap and 35 percent with no formal strategy at all. Those figures explain the funnel shape. A pilot establishes that a capability exists. Production establishes that the capability holds under conditions the pilot never created.

Pilot conditions against production conditions
| Dimension | Pilot | Production |
|---|---|---|
| Input distribution | Curated, representative of the happy path | Long tail, adversarial, malformed |
| Concurrency | One user, one session | Hundreds of sessions, shared rate limits, contention |
| Failure handling | A developer notices and reruns | Automatic, bounded, observable, on-call |
| Evidence | A screen recording | A queryable trace with retention and access control |
| Cost | Absorbed in an innovation budget | A per-task figure that survives a finance review |
| Model version | Whatever was current that week | Pinned, with a regression gate before any change |
That last row is routinely missed. Provider model versions move under running systems. Pin the version, treat a version bump as a code change, and gate it behind the evaluation suite. Programmes that skip this experience unexplained behaviour drift and spend weeks looking for a code cause that does not exist.
Agentic AI adoption forecasts through 2028
Gartner predicts over 40 percent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. The same firm expects 15 percent of day-to-day work decisions to be taken autonomously by 2028, against none in 2024, and 33 percent of enterprise applications to embed agentic capability, against under 1 percent today.
Both forecasts hold simultaneously. Adoption climbs and the cancellation rate stays high, which describes a market separating on execution rather than on ambition.
03
Five architecture decisions that gate agentic AI production readiness
These consume the design phase on real engagements. None is a model selection question. Each one changes the system architecture, which is why each has to be resolved before build rather than during it.
Decision 1 — Scoping the authority boundary at the tool layer
Express the boundary as an enumerated capability set with parameter constraints, then enforce it where the side effect occurs. "Refund up to a stated ceiling against an order the agent has verified as delivered, with no write access to the customer record" is enforceable. "Operates within appropriate limits" is a sentence that every reader interprets differently.
Enforcement belongs in the tool implementation and in the identity layer, never in the system prompt. A prompt instruction is a preference that any sufficiently unusual input can dislodge. A permission check inside the tool, backed by a token whose scope does not include the forbidden operation, is a control that holds when the model behaves unexpectedly. Design the boundary so that a fully compromised prompt still cannot produce an unauthorised side effect.
Decision 2 — Designing the decision trace for audit
Emit a span per step with a stable schema: task identifier, step index, tool name, argument hash, the retrieved evidence identifiers, the result status, latency, token counts, and the termination reason. OpenTelemetry semantics work well here and give the platform team something their existing stack can already ingest.
Capture it synchronously at decision time. A trace assembled afterwards from application logs answers the question the logs were designed for, which is rarely the question an auditor asks. Decide retention and access control at design time as well, because a trace store containing model inputs is a repository of the same sensitive data the agent reads, and it inherits the same obligations.
Decision 3 — Uncertainty routing and escalation policy
Uncertainty is the steady state, so the design question is where it is routed. Four destinations exist: escalate to a human with the accumulated context attached, narrow the scope and retry, refuse and record the refusal, or proceed and flag for asynchronous review. Each routes operational load to a different team, and the choice belongs to whoever absorbs that load.
Confidence signals from the model itself are weak. Stronger triggers are structural: a tool returned an empty result set, retrieved evidence failed a freshness check, two sources disagreed, the step budget is nearly exhausted, or the requested action falls into a high-reversibility-cost tier. Route on those. Agents without an explicit policy default to proceeding quietly, which is the worst available option.
Decision 4 — Reversibility classification and compensating actions
Classify every tool by the cost of undoing its effect, and let that classification drive the control model rather than sitting beside it in a document.

Anything below the first tier needs a written compensating action implemented and tested before launch, in the manner of a saga. Teams that plan to implement compensation later discover that the business process on the other side never had one.
Decision 5 — Operational ownership and the on-call model
The team that builds an agent is rarely the team paged at three in the morning. Name the operating owner during design and give that owner authority over the escalation policy, the evaluation set and the release gate.
Define what an agent incident is, who is paged, and what the kill switch does: pausing new task admission while allowing in-flight tasks to drain is usually correct, and a hard stop mid-task can leave compensable actions uncompensated.
04
Three platform dependencies that cap agentic AI performance
An agent inherits the properties of the systems it reads from and writes to. Three of them set the ceiling, and the readiness work on each is unglamorous, expensive to defer, and almost never inside the pilot budget.
Retrieval readiness: searchability, freshness and entitlement propagation
A Tech Value Survey found nearly half of organisations naming searchability of data at 48 percent and reusability at 47 percent as obstacles to their AI automation strategy. Both translate directly into agent behaviour.
Three properties matter at the engineering level. Retrieval has to be entitlement-aware inside the query rather than filtered afterwards, so that authorisation constrains the candidate set instead of trimming the result set.
Freshness has to be explicit per source, with an ingestion path matched to the tolerance: change data capture for minutes, batch for overnight. And the index has to expose the structure the agent can navigate, since an agent that cannot filter by effective date or document status will retrieve a superseded policy and act on it with complete confidence.
Dan Marks · VP, Business, MactoresIn the agent programmes reviewed, the technical blocker was almost never the model. Plenty stalled on one of three things: an authority boundary nobody had written down, an evaluation set nobody had built, or a source system nobody had assessed. The programmes that shipped fastest spent their first fortnight on the paperwork most teams treat as a distraction.
FAQs
Who should sponsor an agentic AI programme?
How long does a first agent take to reach production?
What does an agentic AI team look like?
How do agents change the compliance conversation?
What happens to the people whose work the agent takes over?
Can agentic AI run on data that cannot leave our environment?
How many agents should a first programme attempt?
What belongs in an agentic AI statement of work?
AI agents for apps
Ready to move past the pilot?
Talk to delivery