Mactores POV

Test: Agentic AI: The Production Path Every Enterprise Needs by 2027

Written by Ammar Nizami | Sep 15, 2026, 4:44:43 PM

Key takeaways

  • Production is rare, not impossible. 30 percent are exploring agentic options, 38 percent piloting, 14 percent ready to deploy, 11 percent running in production.
  • The blocker is engineering discipline, not model quality. Transaction design, authorization, observability and cost engineering close the gap.
  • Five decisions gate the build: authority boundary, decision trace, uncertainty routing, reversibility classification, operational ownership.
  • Three platform dependencies cap performance: retrieval readiness, a transactional path to the legacy system of record, and agent identity with short-lived credentials.

Deloitte's 2026 Tech Trends research measures how few organisations have built it: 30 percent are exploring agentic options, 38 percent are piloting, 14 percent have something ready to deploy, and 11 percent are running these systems in production.

The attrition between piloting and production is not a model capability problem. It is a distributed systems problem wearing a machine learning label, and the disciplines that close it are transaction design, authorization, observability and cost engineering.

Agentic AI architecture: what separates an agent from a copilot

The term covers five distinct architectures. They differ in who selects the next action, and that single variable determines the testing strategy, the failure surface and the governance model.

Agent autonomy levels, from assistant to multi-agent orchestration

Five architectures, ordered by who decides the next step
Level Control flow Who selects the next action Dominant failure mode
Assistant Single request, single response A person, on every turn An incorrect answer that nobody validates
Scripted automation Static DAG, fixed branches The author of the workflow Silent breakage when an input schema drifts
Copilot Generate, then human approval gate A person, at every step Approval fatigue degrading into rubber-stamping
Agent Closed loop: observe, decide, act, observe The model, inside an enforced envelope Confident action on a wrong premise
Multi-agent Delegated subtasks with a coordinator Several models plus routing logic Compounding error with no single trace to follow

Only the final two rows are agentic in a way that changes engineering practice. The discriminator is whether the control flow is decided at runtime by the model. A copilot has a static call graph.

An agent does not, which means the set of reachable states cannot be enumerated in advance and the test strategy has to shift from path coverage to invariant checking.

Why runtime control flow changes the failure surface

An assistant returning a poor answer costs a minute of attention. An agent taking a wrong action emits a side effect: a written record, a dispatched message, a posted transaction. Four engineering properties become mandatory at that point.

  • Idempotency on every tool that mutates state. Agent loops retry. Without an idempotency key derived from the task and the intended effect, a retry after a timeout produces a duplicate side effect, and duplicate refunds are the canonical example that reaches an executive.
  • Bounded execution. A step budget, a wall-clock timeout and a token ceiling per task, enforced by the orchestrator rather than requested in a prompt. Unbounded loops are the most common cause of an agent cost incident.
  • Schema validation at the tool boundary. Tool arguments arrive as model output, which means they are untrusted input. Validate against a strict schema, reject on failure, and never pass a model-generated string into an interpreter, a query planner or a shell.
  • A decision trace. Every tool call, its arguments, its result, and the model's stated basis for the next step, captured synchronously.

Fig. 01 The four gates that turn a loop into a controlled system. Every badge on the loop is enforced by the orchestrator, never requested in the prompt.Mactores delivery reference architecture, 2026.

02

The agentic AI production gap: 11 percent, and the reasons behind it

The same Deloitte research reports 42 percent of organisations still developing an agentic strategy roadmap and 35 percent with no formal strategy at all. Those figures explain the funnel shape. A pilot establishes that a capability exists. Production establishes that the capability holds under conditions the pilot never created.


Fig. 02 Adoption concentrates in the pilot stage. The stages are separate states, so the figures do not sum to 100.Source: Deloitte Insights, Tech Trends 2026.

Pilot conditions against production conditions

What changes between a pilot and a production system
Dimension Pilot Production
Input distribution Curated, representative of the happy path Long tail, adversarial, malformed
Concurrency One user, one session Hundreds of sessions, shared rate limits, contention
Failure handling A developer notices and reruns Automatic, bounded, observable, on-call
Evidence A screen recording A queryable trace with retention and access control
Cost Absorbed in an innovation budget A per-task figure that survives a finance review
Model version Whatever was current that week Pinned, with a regression gate before any change

That last row is routinely missed. Provider model versions move under running systems. Pin the version, treat a version bump as a code change, and gate it behind the evaluation suite. Programmes that skip this experience unexplained behaviour drift and spend weeks looking for a code cause that does not exist.

Agentic AI adoption forecasts through 2028

Gartner predicts over 40 percent of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. The same firm expects 15 percent of day-to-day work decisions to be taken autonomously by 2028, against none in 2024, and 33 percent of enterprise applications to embed agentic capability, against under 1 percent today.

Both forecasts hold simultaneously. Adoption climbs and the cancellation rate stays high, which describes a market separating on execution rather than on ambition.

03

Five architecture decisions that gate agentic AI production readiness

These consume the design phase on real engagements. None is a model selection question. Each one changes the system architecture, which is why each has to be resolved before build rather than during it.

Decision 1 — Scoping the authority boundary at the tool layer

Express the boundary as an enumerated capability set with parameter constraints, then enforce it where the side effect occurs. "Refund up to a stated ceiling against an order the agent has verified as delivered, with no write access to the customer record" is enforceable. "Operates within appropriate limits" is a sentence that every reader interprets differently.

Enforcement belongs in the tool implementation and in the identity layer, never in the system prompt. A prompt instruction is a preference that any sufficiently unusual input can dislodge. A permission check inside the tool, backed by a token whose scope does not include the forbidden operation, is a control that holds when the model behaves unexpectedly. Design the boundary so that a fully compromised prompt still cannot produce an unauthorised side effect.

Decision 2 — Designing the decision trace for audit

Emit a span per step with a stable schema: task identifier, step index, tool name, argument hash, the retrieved evidence identifiers, the result status, latency, token counts, and the termination reason. OpenTelemetry semantics work well here and give the platform team something their existing stack can already ingest.

Capture it synchronously at decision time. A trace assembled afterwards from application logs answers the question the logs were designed for, which is rarely the question an auditor asks. Decide retention and access control at design time as well, because a trace store containing model inputs is a repository of the same sensitive data the agent reads, and it inherits the same obligations.

Decision 3 — Uncertainty routing and escalation policy

Uncertainty is the steady state, so the design question is where it is routed. Four destinations exist: escalate to a human with the accumulated context attached, narrow the scope and retry, refuse and record the refusal, or proceed and flag for asynchronous review. Each routes operational load to a different team, and the choice belongs to whoever absorbs that load.

Confidence signals from the model itself are weak. Stronger triggers are structural: a tool returned an empty result set, retrieved evidence failed a freshness check, two sources disagreed, the step budget is nearly exhausted, or the requested action falls into a high-reversibility-cost tier. Route on those. Agents without an explicit policy default to proceeding quietly, which is the worst available option.

Decision 4 — Reversibility classification and compensating actions

Classify every tool by the cost of undoing its effect, and let that classification drive the control model rather than sitting beside it in a document.


Fig. 03 The cost of undo rises down the table, and the control model has to rise with it.Mactores delivery reference architecture, 2026.

Anything below the first tier needs a written compensating action implemented and tested before launch, in the manner of a saga. Teams that plan to implement compensation later discover that the business process on the other side never had one.

Decision 5 — Operational ownership and the on-call model

The team that builds an agent is rarely the team paged at three in the morning. Name the operating owner during design and give that owner authority over the escalation policy, the evaluation set and the release gate.

Define what an agent incident is, who is paged, and what the kill switch does: pausing new task admission while allowing in-flight tasks to drain is usually correct, and a hard stop mid-task can leave compensable actions uncompensated.

04

Three platform dependencies that cap agentic AI performance

An agent inherits the properties of the systems it reads from and writes to. Three of them set the ceiling, and the readiness work on each is unglamorous, expensive to defer, and almost never inside the pilot budget.

Retrieval readiness: searchability, freshness and entitlement propagation

A Tech Value Survey found nearly half of organisations naming searchability of data at 48 percent and reusability at 47 percent as obstacles to their AI automation strategy. Both translate directly into agent behaviour.

Three properties matter at the engineering level. Retrieval has to be entitlement-aware inside the query rather than filtered afterwards, so that authorisation constrains the candidate set instead of trimming the result set.

Freshness has to be explicit per source, with an ingestion path matched to the tolerance: change data capture for minutes, batch for overnight. And the index has to expose the structure the agent can navigate, since an agent that cannot filter by effective date or document status will retrieve a superseded policy and act on it with complete confidence.

Legacy application and database access: APIs, CDC and stored-procedure logic

Concretely, the agent needs a system of record that offers no transactional API, an access model built for a reporting tool, and business rules living in stored procedures and triggers that predate everyone on the current team.

The available routes carry different risks. A read path via change data capture into a queryable store is usually safe and gets the agent current state without touching the source. A write path is harder, and the honest options are a service façade in front of the legacy system that enforces the business rules explicitly, or modernising the component that owns the process.

Screen-level automation against a terminal or a browser session is fragile enough that it belongs in the plan as an interim measure with an expiry date rather than a destination.

This is application and database engineering with an agentic outcome attached. Scoping it as an AI project guarantees the difficult work surfaces in the phase with the least remaining budget flexibility.

 

jAgent identity, delegated authorization and short-lived credentials

An agent acting for a user should carry that user's authority through a token exchange, so downstream systems see a delegated principal and their existing authorization logic keeps working unchanged.

An agent acting autonomously needs a workload identity of its own with a narrowly scoped role, credentials measured in minutes rather than months, and an audit trail that distinguishes agent-initiated actions from human ones.

Two anti-patterns dominate. The first is a shared service account with broad standing privilege, convenient in a pilot and indefensible under review, where remediation means a rebuild rather than a patch.

The second is a long-lived API key embedded in agent configuration, which converts a prompt injection into credential exfiltration. Assume any string the agent can read may reach the model context, and keep secrets out of that path entirely.

05

Agentic AI cost model: why spend scales with step count

Chatbot economics are per-response. Agent economics are per-step, and step count grows with task ambiguity. A task requiring twelve tool calls with full history each time costs far more than twelve times a single call, because context accumulates quadratically across the loop.

Six cost drivers and the engineering control for each
Cost driver Mechanism Engineering control
Steps per completed task Ambiguous goals and vague tool descriptions cause exploratory wandering Sharper task decomposition, tool descriptions with explicit preconditions
Context growth across steps Full transcript resent on every call, so tokens grow with the square of steps Summarise between steps, scope context per tool, prune resolved subtasks
Retries after failed calls Flaky dependencies and unhandled errors produce loops Step budgets, exponential backoff, a defined give-up state
Cache miss rate Stable prefixes rebuilt on every call Prompt caching with a deliberately stable system prefix
Tool latency Slow dependencies stall the loop and multiply timeout cost Parallel independent calls, caching, realistic timeout policy
Human review time Low unassisted completion returns people to the loop Raise autonomy only where the reversibility tier permits

Track cost per resolved task against the fully loaded cost of the process being replaced. Cost per token is an input to that figure and says nothing useful on its own. Watch the distribution rather than the mean, since agent cost is long-tailed and a small fraction of pathological tasks frequently dominates the bill.

06

Agentic AI evaluation: metrics, golden sets and regression gates

Demonstrations are judged by whether the audience is impressed. Production systems need numbers a sceptical operations lead will accept.

Seven production metrics and the warning pattern for each
Metric Definition Warning pattern
Unassisted completion rate Tasks finished with no human touch Rising alongside a rising reversal rate
Escalation precision Share of escalations a human agreed were necessary Falling escalation volume while quality complaints climb
Reversal rate Completed actions later undone Anything above the human baseline for the same process
Cost per resolved task Total spend divided by tasks completed A figure that only holds at pilot volume
Latency at p95 Wall clock to resolution Acceptable mean with an unacceptable tail
Trace completeness Actions with a full, queryable decision trace Anything below 100 percent for consequential actions
Audit pass rate on first review Reviews cleared without engineering explanation Reviewers needing a walkthrough of raw logs

Build a golden set of real tasks with known correct outcomes before tuning anything, and run it as a gate in continuous integration so that no prompt, tool or model changes ship without a scored comparison.

Where a model is used as a judge, calibrate it against human labels on a sample and re-calibrate whenever the judge model changes, because an uncalibrated judge produces confident scores with no relationship to quality.

Capture the human baseline for every metric before go-live. A baseline reconstructed afterwards produces a comparison the business will not accept.

07

Build, buy or partner for agentic AI delivery

Most enterprises will use all three across a portfolio, and deciding per use case beats adopting a house position.

Three delivery routes, with the real cost of each
Route Best fit Real cost Evidence to weigh
Build in house Core differentiation, deep domain knowledge, spare senior capacity The product roadmap stalls while the team learns transaction design and evaluation engineering Strong where the process is proprietary and stable
Buy a product Common processes with an established vendor category Configuration debt and dependence on another roadmap Reasonable where the process can bend to the product
Partner on delivery A contracted date with a capability gap in between The handover has to be designed rather than assumed MIT research found pilots built through partnerships roughly twice as likely to reach full deployment

That MIT finding, from the 2025 study on the divide between AI pilots and AI in production, also recorded employee usage of externally built tools at nearly double the internal equivalent. The mechanism is unglamorous.

Teams that have shipped this class of system arrive with the failure modes already catalogued, and they are contractually accountable for a date.

08

Agentic AI vendor evaluation: separating agents from agent washing

Gartner applies the term "agent washing" to assistants, chatbots and robotic process automation rebranded as agentic products, and estimates roughly 130 of the thousands of vendors claiming agentic capability are genuine. Five technical questions sort them quickly.

Where is the authority boundary enforced

Ask to see the code path. An answer pointing at a system prompt is describing a preference. The correct answer names the tool layer and the identity layer, and can show a token whose scope excludes the forbidden operation.

What does the decision trace contain

Ask for the span schema and a real trace from a production task. A vendor without a stable trace schema has not been through a compliance review.

How does the evaluation gate work

Ask what runs before a release and what blocks it. The answer should describe a labelled task set, a scored comparison and a threshold. A person trying the system and forming an impression is not a gate.

What is the cost per resolved task

Vendors confident in their unit economics answer immediately, with a distribution rather than a mean. Redirecting to model pricing answers a different question.

Who operates it on day one after go-live

If the answer is a managed service with no defined path to internal ownership, the dependency has moved rather than resolved.

09

Four failure modes in enterprise agentic AI programmes

The failures worth studying cluster tightly.

  • The demonstration was optimised instead of the system. Effort concentrated on the happy path, and the long tail was discovered by users in week one.
  • The authority boundary lived in a document. Nothing enforced it, so behaviour drifted quietly as prompts were edited by people who never saw the original design intent.
  • Nobody owned the numbers. With no baseline and no golden set, the question of whether the agent works never resolves, and the programme loses its sponsor before it loses its funding.
  • The platform dependencies were assumed rather than assessed. The data was not searchable, the legacy system had no write path, and both facts arrived after the architecture was committed.