Chevron left
RESEARCH

The production gap: 12 lessons from a year of enterprise AI delivery

A field report from seven engagements between December 2025 and September 2026. Most of what decided whether AI reached production was not the model.
September 25, 2026
Research

Finding: in our recent enterprise work, the things that decided whether an AI system reached production and stayed there were rarely the model. They were the process, organizational, and design choices around it: how a benchmark was agreed, when legal review started, whether mobile was tested, how scores were calibrated, who would run the system after launch, and how adoption was handled.

Summary

  • We reviewed seven engagements delivered between December 2025 and September 2026 for two client organizations, one of them a large enterprise software company.
  • The work covered voice AI training and certification, enterprise retrieval, conversation intelligence, a financial business case, an AI-native operations platform, and an architecture exploration for an AI-native product.
  • Four systems are in production or in active delivery. One was delivered as a benchmarked proof. One was a strategy and modeling deliverable. One stayed exploratory.
  • The twelve lessons below fall into four groups: proving value, shipping, scoring and trust, and operating after launch.

Method

This is a qualitative field report, not a survey. For each engagement we reviewed the engagement record written at wrap-up, the project tracking history, and recorded meeting summaries. We then grouped the recurring causes of delay, rework, and success. Client identities and identifying details are removed. The limitations are at the end.

Proving value

1. Agree on the benchmark before you deliver the proof

We rebuilt a stalled help assistant in three weeks and cut answer time from tens of seconds to 3 to 5 seconds. The formal head-to-head comparison kept getting postponed. Put the method, question set, and date in writing before the proof is delivered. See Rebuild or tune?

2. Build proofs with the internal team, not against it

A fast rebuild that beats a long internal effort creates a political cost for the team that owns it. Share the code, the evaluation set, and the credit.

3. Check the numbers in a business case against the recordings

While building a financial model, we found that a sales-cycle figure repeated across several documents came from an AI meeting summary, not from anything a stakeholder said. If a number drives a decision, find the moment someone actually said it. See the business case study.

4. Strategic exploration needs its own paid, fixed-scope structure

One engagement produced useful architecture work for a client's multi-year AI-native product plans, but it never became a paid project. Open-ended strategy sessions with no commercial structure behave like unpaid architecture work. A small, fixed-fee exploration tier fixes that for both sides.

Shipping

5. Test mobile before launch

On two separate voice AI programs, the first real users found mobile bugs. Make a mobile test pass a release requirement.

6. Scope cohort management up front

Bulk invitations with learning-track assignment, role changes from trainee to manager, and team-level reporting were each requested after launch and rebuilt later. For any system with managed user groups, how cohorts are managed needs its own scoping questions.

7. Speed and polish have to be balanced

Shipping more than 150 updates in a single week got a training platform into daily use fast. It also left navigation and role-switching rough edges that users raised for months. Set aside time for polish in every release cycle.

8. Start legal review early for anything that records people

Conversation recording raises consent rules that vary by jurisdiction, plus customer disclosure questions. Legal review set the contract timeline more than engineering did. Start review early, pilot in one-party-consent jurisdictions, and support manual entry for customers who decline to be recorded.

Scoring and trust

9. AI scores need calibration, not just a rubric

Both voice AI programs showed early score inflation, with at least one transcript about ten points too high. Calibration needs reference sets, a quote behind every point, and drift checks. See Calibrating AI graders.

10. Keep a human override while trust builds

A trainer override for AI scores, and approved answers served ahead of generated ones in retrieval, have the same purpose: people can trust the system while it improves. See Premium and fallback answers.

Operating after launch

11. Decide who runs it, and plan the handoff

One client kept a production platform vendor-run instead of taking it in-house on the original date, then moved it in stages once usage was clearer. Who runs the system is a decision, not a default. See Vendor-operated or in-house.

12. For operations platforms, adoption is the project

Replacing tools a team already knows succeeds or fails on adoption. What we are doing: access that starts with the process owner, role-specific surveys, walkthrough videos, weekly check-ins, and prioritizing speed as soon as users mention it. See the operations platform case study.

What these lessons have in common

Most of these twelve lessons are not about models. The model decided whether a system could work. Commercial agreements, legal review, user groups, calibration, and operating decisions decided whether it did.

Limitations

  • Seven engagements and two clients is a small sample. Most engagements are with one enterprise client, so the lessons may reflect that organization.
  • This is qualitative. We have not measured these effects against a control group.
  • Several systems are still in delivery, so long-term outcomes are not yet known.

We plan to update this report as more engagements finish. For the broader transition from prototype to production, see From prototype to production AI.

About ALLTIPLY Labs

ALLTIPLY Labs publishes what we learn from delivering AI systems for real organizations. If your AI work is stuck between demo and production, talk to us.

Read Our Research