Executive summary
- Client: an enterprise software company with a large field and inside sales organization.
- Problem: sales practice depended on live roleplay with managers, which was subjective, hard to schedule, and impossible to scale to every new hire.
- What we built: a voice AI roleplay platform where reps practice realistic buyer conversations, get scored against a configurable rubric, and receive specific feedback. Managers build and assign scenarios without engineering help.
- Status: in production since December 2025. More than 2,000 recent roleplays, running across several sales divisions.
- Why it matters: practice became something a rep can do any time, scored the same way every time, instead of something that happens when a manager has a free hour.
The problem
New sales hires at this company had to learn a complex product line and a long, multi-stakeholder sales cycle. The standard way to practice was roleplay with a manager or senior rep. That worked when it happened, but it rarely happened often enough.
Three things made the old approach hard to scale:
- Scoring was subjective. Two managers could watch the same roleplay and grade it differently. Reps had no consistent picture of what "good" sounded like.
- Practice was scarce. Every session needed a person on the other side. Managers had quotas of their own.
- Training content was passive. Much of the existing enablement lived in video modules that reps watched but never had to perform.
What we built
ALLTIPLY designed, built, and ran a web platform that replaces the human buyer in a roleplay with a voice agent, and replaces ad hoc grading with a structured rubric.
The practice loop
- A rep picks or is assigned a scenario, then talks to an AI buyer by voice in the browser or on a phone.
- Speech is transcribed in real time. The buyer responds in character, raises objections, and pushes back when the rep is vague.
- When the call ends, the rep gets a score on a 4.0 scale and written feedback tied to specific moments in the conversation. The feedback flags habits like filler words and closed questions where an open question would have surfaced more.
The scenario engine
- Scenarios are generated from uploaded files, stated goals, objection lists, and competitive positioning, grounded in a corpus of more than 70,000 company documents.
- Documents are tagged with metadata for relevance, compliance review, and regional variation, so a scenario for one market does not borrow facts from another.
- A knowledge graph of buyer personas, pain points, and competitors, with tens of thousands of nodes, gives each AI buyer a consistent background and set of concerns.
- Scenarios use dynamic variables, so one scenario template can produce many realistic variations.
The manager portal
- Managers build scenarios, set scoring rubrics, and assign practice to cohorts and learning tracks.
- Role-based access separates learners, managers, and administrators. Edits are reversible and every change is recorded in an audit log.
- Bulk invitations, team hierarchy reporting, and progress dashboards let sales leaders see who is practicing and where the team is weak.
How we delivered it
The platform launched on the client's single sign-on and moved into daily use quickly. During the most intense refinement period the team shipped more than 150 platform updates in a single week, driven by feedback from the first cohorts.
One design decision paid off later: the platform was built to serve more than one population from the start. When the sales organization wanted to certify its field reps on a new go-to-market message, that program launched on the same platform without rebuilding the core. See the sales certification case study.
Billing followed the shape of the cost: a flat fee for the platform, plus usage-based billing for voice minutes, which spike during certification windows.
Operating model
The original plan was to hand the platform to the client's own IT team in mid-2026. The client chose to keep it vendor-run while usage grew. The platform is now moving onto client infrastructure in stages: separate QA and production environments, the corporate identity provider for sign-on and role-based access, and pay-as-you-go tiers on third-party services while usage is still hard to predict. We cover how to make that decision in Vendor-operated or in-house.
What was hard, and what we would do differently
- User experience rough edges lasted too long. Navigation, role switching, and a few reporting inconsistencies came up repeatedly in feedback. Feature velocity outpaced polish.
- Mobile testing has to happen before launch, not after. The first real users found mobile bugs that a dedicated mobile pass would have caught.
- Cohorts drop off. One workshop cohort lost 15 to 20 percent of participants. Dropout is better handled in the product itself, with onboarding nudges and at-risk flags visible to managers, than by follow-up emails.
- AI scoring needs calibration, not just a rubric. At least one transcript was graded about ten points too high. Scoring accuracy became a standing discipline. See Calibrating AI graders.
What was multiplied
The company did not need more trainers. It needed practice that was available on demand and graded consistently. The platform turned a scarce, manager-dependent activity into one that scales with the size of the sales team, and it created a record of how every rep performs against the same standard.
Related
- Business voice AI
- Custom AI development
- The production gap: a field report from a year of enterprise AI delivery
Building something similar? Talk to us.
Related services: Voice AI Training and Certification and Private AI Deployment.



