Executive summary
- Client: an enterprise software company certifying its field sales force on a new go-to-market message.
- Problem: the existing certification was a video course. It checked that reps had watched the material, not that they could deliver the message to a skeptical buyer.
- What we built: a tiered voice AI certification. Reps deliver the message to an AI buyer in competitive scenarios and are scored against 16 weighted criteria, with an 80 percent pass mark and a manual override for trainers.
- Status: in production since April 2026, rolling out to roughly 80 field reps and senior sales leaders in controlled waves.
- Key lesson: the rubric was the easy part. Keeping AI scores accurate and defensible took ongoing calibration.
The problem
The sales organization had rebuilt its go-to-market story around four core messages. Leadership wanted proof that every field rep could deliver all four, clearly, under pressure, against the competitors they actually face.
The old certification could not show that. A rep could finish every video and still fumble the message in front of a buyer. Leaders also needed certification to hold up: if a rep failed, the result had to be explainable.
What we built
The certification runs on the voice AI roleplay platform ALLTIPLY had already built for the same company. The platform stayed the same. The scenarios, rubric, and reporting were new.
Three tiers
- Beginner: open book, with heavy prompting. Practice mode, no stakes.
- Intermediate: hints can be turned on or off.
- Advanced: the rep must pass without hints. This is the certification.
The rubric
- 16 weighted criteria covering all four messages, supporting examples, and the use of open-ended questions.
- Pass mark of 80 percent, meaning a rep can miss up to three criteria and still pass.
- Bonus points for concise, clear delivery, to reward the behavior leaders wanted to see more of.
- A 10-minute target per roleplay, with a 12-minute grace period.
The scenarios
- Three competitive scenarios based on the real displacement situations reps face, each with a buyer who already uses a competitor.
- The AI buyer raises realistic objections and steers the conversation back when a rep drifts off topic.
- Versions of each scenario were written for different sales roles, after early feedback showed a single sales-focused version did not fit every team.
Feedback and reporting
- Reps see results on a simple grid showing each of the four messages. The design deliberately leaves out the full criterion-by-criterion score, to reduce cognitive load and keep attention on the four things that matter.
- Positive feedback appears during the roleplay, not only after it.
- Trainers can override any AI score. That keeps the process trusted while the scoring model is still being tuned.
- Sales leaders see certification progress by team hierarchy, so they can follow up with the specific reps who have not started or who are stuck.
How it rolled out
The rollout plan had sales leaders take the certification themselves, as reps, before the field did, so wording problems and scoring disagreements would surface while the stakes were still low.
Rollout then moved in waves. The first wave assigned around 40 users with a fixed deadline, then scaled toward a target of 50 to 60 active users before widening further. Scaling slowly gave the team time to fix invitation problems, clarify roles, and improve reporting between waves.
Completion varies a lot by division, ranging from about 15 percent to 75 percent at the time of writing. That spread is why team-level reporting became a priority: aggregate completion hides which managers need help.
What was hard, and what we would do differently
- Early scores ran high. In early testing, at least one transcript scored about ten points above what a human grader gave it. The same pattern showed up on the sister training platform. Scoring calibration needs to be its own engineering discipline, with reference sets and drift checks, not tuning done separately for each program. Our recommended method is written up in Calibrating AI graders.
- Bulk invitations with track assignment should have been in scope from the start. It was requested after the first wave and rebuilt later. Cohort management deserves its own scoping questions.
- Role changes are messy. People move from trainee to manager mid-program. Account roles, default views, and reporting all have to handle the transition cleanly.
- Mobile needs its own test pass. The first test user found mobile bugs. Same lesson as the training platform.
What was multiplied
Certification stopped measuring attendance and started measuring performance. Every rep is scored against the same 16 criteria, sales leaders can see exactly who needs help and on which message, and the organization has a repeatable way to certify the next message it launches.
Related
Planning a certification program? Talk to us.
Related service: Voice AI Training and Certification.



