Most enterprise AI clears every technical bar and still never reaches the people meant to use it. The reason is rarely the model.
By Martin Papy, CTO APAC, CBTW
Here is a distinction most enterprise AI programs never make explicit, and it is one that often decides whether they survive contact with production: a use case defines what the AI should do. A workflow defines how the output gets used, reviewed, challenged, corrected and acted on. A pilot can satisfy the first completely and still fail the second, and when it does, the failure rarely looks like a model problem. It looks like adoption stalling, sign-off taking months, or a system quietly being routed around.
I see this often enough that it no longer surprises me. The pilot works in controlled conditions. The demo is persuasive. Executive support follows. Then progress slows. The model still produces a usable answer. The answer simply does not fit cleanly into the operating environment it was meant to serve. The model clears its own bar. For enterprise AI, the last mile, the handoff to the people who have to act on what it produces, is where many programs actually get stuck.
Deloitte’s 2026 State of AI in the Enterprise report found that only 25 per cent of respondents globally, and 28 per cent in Australia, have moved at least 40 per cent of their AI pilots into production. The market has largely settled on three explanations for numbers like this: weak governance, poor data quality, unclear ROI. All three can genuinely stall a program. In my experience, there is also a fourth cause that gets less attention than the other three, and it sits upstream of them: many of the requirements a workflow surfaces are invisible until people use the system on live work.
Why workflow requirements don’t show up until production
A use case can be fully specified on paper: input, model, output. A workflow cannot, because it depends on where human judgment enters the process, what exceptions occur in live conditions, what data can actually be trusted, and what happens when the system is wrong. None of that exists in a requirements document. It exists in the gap between what engineers, planners, compliance teams and operational users expected the system to do and what they discover when they actually use it. That gap is the last mile, and it is fundamentally a human one.
Forward deployed engineering has become the recognized label for closing it. The label matters less than the operating mechanism underneath it: proximity to the live workflow. Placing engineers on site does not close the gap on its own. Proximity to the live workflow does, because it changes the question being asked from what can the AI do, to what must be true for someone to trust this output and act on it. Can the output be traced to source data? Can a user override it? Who owns the workflow after deployment? Those questions only get answered when the delivery team is close enough to observe the workflow as it actually happens.
What this looks like close to live work
A recent example from CBTW‘s delivery work with Twintech on Konnect xD, an AI agentic platform for EPC execution, makes the distinction concrete. Konnect Planning has been live on a 28,000-activity schedule across 16 users and six departments since March. The team ran structured review sessions, shipped five releases, and tied each one to specific feedback from the people using the product daily. The process moved the roadmap away from speculative feature planning and toward requirements shaped by live execution: day-based float units, Monte Carlo risk export and import for bulk editing, manual overrides for factors automated analysis could not capture, and schedule version comparison. These are the kinds of requirements that only surface once a planner is managing a live 28,000-activity schedule under deadline pressure, rather than reviewing a brief about one.
The same pattern showed up in a different module. Konnect Engineering processed 750 wiring diagram pages and 6,515 cable tags in under 60 seconds on a live offshore construction project. Speed alone left a gap. A flagged discrepancy still requires the engineer to work out what to do about it, and at 6,515 cable tags that translates into hours of manual interpretation. The workflow needed each discrepancy mapped to a specific, per-platform action item, so the output arrives ready to act on rather than ready to interpret.

Architecture constraints surface the same way. Initially, Konnect 3D was cloud-dependent, which blocked on-premise deployment in data-sovereign regions. It capped out at 500MB models and roughly two hours of processing. The team reworked it for on-premise use. The new version handles 15GB-plus models in about five minutes. That constraint never appeared in a benchmark, it appeared in the operating environment.
For enterprise AI, this proximity to the operating environment is what turns technical capability into something people can actually use. Production exposes requirements around architecture, accountability and human decision-making that controlled pilots often cannot surface.
What makes an output trustworthy enough to act on
This is also where the trust question gets resolved, or left unresolved. NIST’s AI Risk Management Framework identifies validity, reliability, accountability, and transparency among the characteristics of trustworthy AI. In my experience working across regulated environments, those characteristics hold up when they are built into the system and workflow from the start, rather than assessed after the fact.
CBTW’s work on Konnect Cost Estimation, currently in customer pilots, shows what that looks like in practice. The module produces a cost figure alongside class transition reporting that explains why a number moved between revisions, giving the output enough traceability to withstand scrutiny from the people whose job is to challenge it. A model can be statistically sound and still fail this test when it produces an answer without a clear trail back to why.
This becomes particularly important as enterprise AI moves into workflows where outputs influence operational, commercial or compliance decisions. Trust depends on whether people can understand, challenge and act on what the system produces, not simply whether the model performs well in isolation.
Five questions worth asking before scaling
Production readiness comes down to five practical questions. Each one shifts attention from model performance alone to the conditions required for people to trust, use, and sustain the system in a live workflow:
- Is the team close enough to the live workflow to understand how the work actually happens beyond the documented process?
- Can users review, challenge, and override the output when needed?
- Can the output be traced back to source data and decision logic?
- Does the architecture match the operating environment, including its sovereignty and security constraints?
- Is feedback changing the product quickly enough to keep pace with the workflow, instead of waiting for a quarterly steering cycle?
Strong models, clean data, and governance still matter. Their value shows up when they translate into something a workflow will actually accept. The five questions above are the practical test for whether that translation can happen.
The last mile is where the competitive question is changing
As enterprise AI capability becomes more accessible, competitive advantage depends less on running the most experiments or adopting the newest model first, and more on closing the distance between the people building the system and the people whose work it is meant to change. That distance is harder to productize than a delivery model name, yet it is where production AI succeeds or fails.
Martin Papy is CTO APAC at CBTW, a global technology solutions company operating in 20+ countries, with 650+ people across eight offices in the Asia-Pacific region.








