Finding: AI graders tend to score generously, and a good rubric does not stop them. Accurate, defensible AI scoring needs its own calibration process: a human-graded reference set, a quote from the transcript behind every point awarded, a logged human override, and ongoing checks for drift.
Summary
- On two separate voice AI sales platforms ALLTIPLY built and operates, early AI scores ran high. In at least one case a transcript scored about ten points above the human grade.
- The same inflation on two platforms with different rubrics suggests the cause is how language models grade, not a flaw in one rubric.
- Rubric wording helps but does not solve it. The controls that address it are a transcript quote behind every point, grading criteria separately, a human-graded reference set, logged trainer overrides, and drift checks. This paper separates what is already in place on these programs from what we recommend.
- When scores affect certification or pay, leave a human override in place until measured agreement justifies removing it.
Where this comes from
ALLTIPLY built and runs two voice AI programs for an enterprise software company's sales organization: a roleplay training platform scored on a 4.0 scale, and a field sales certification scored against 16 weighted criteria with an 80 percent pass mark. Both use a language model to grade a transcribed voice conversation against a rubric. Both showed score inflation early on.
Why AI graders inflate scores
Three mechanisms are likely at work, and all three apply to grading conversations against a rubric.
- Leniency. Models tuned to be helpful tend to give credit when a response is close enough. A rep who mentions a topic in passing gets credit for "clearly presenting" it.
- Fluency bias. Confident, well-structured speech reads as competent even when required content is missing.
- Criteria bleeding together. When all 16 criteria are graded in one pass, a strong impression on some raises the scores on others. The overall impression leaks into individual items.
What is in place today
- A human override. Trainers can override any AI score. That keeps the program trusted while the grader is being tuned.
- Leaders go first. The rollout plan had sales leaders take the certification as reps before the field did, to surface scoring disagreements and wording problems while the stakes were low.
- Tuning before field rollout. The certification grader was tuned after early transcripts scored high, and the field rollout moved in controlled waves so scoring problems surfaced with small groups.
The controls we recommend
The following controls address the causes above. They are the standard we recommend for any AI grading that affects people, including these programs as they mature.
1. Require evidence for every point
For each criterion, the grader must quote the exact part of the transcript that satisfies it, or state that nothing does. A point without a quote is not awarded. This goes straight at leniency, because the model can no longer give credit for an impression.
2. Keep criteria separate
Grade criteria individually or in small related groups, not all 16 in one pass. This costs more model calls and returns cleaner scores, and it makes each score easier to explain to the rep who received it.
3. Build a human-graded reference set
Have at least two people grade a fixed set of transcripts, including clear passes, clear fails, and borderline cases. Transcripts from leaders who take the program first are a good starting point. Any change to the grader has to match the human grades on that set before it ships.
4. Treat overrides as data
Record every trainer override with a reason. A pattern of overrides on one criterion shows where the grader needs work next.
5. Check for drift
Model updates, prompt changes, and new scenarios all shift scores. Re-run the reference set after any change and compare score distributions over time. A sudden rise in pass rate is more likely a grader change than a sudden improvement in reps.
Design choices that help reps
Accuracy is not enough if reps cannot act on the result. Two design choices from the certification program:
- Show less detail. Reps see a simple grid by core message instead of a full criterion-by-criterion breakdown, to keep their attention on what matters most. Trainers can still see the detail.
- Reward the behavior you want. Bonus points for concise, clear delivery make it visible that leadership values it.
A checklist before AI scores count for anything
- A reference set graded by at least two humans, with their level of agreement known.
- The grader's agreement with humans measured on that set, not assumed.
- A transcript quote required for every point awarded.
- A human override with a recorded reason.
- A drift check that runs on every change to the prompt, the model, or the scenarios.
- Pass thresholds set after calibration, not before.
Method and limitations
This paper describes what we observed operating two production programs in 2025 and 2026, the controls already in place, and the controls we recommend. The ten-point figure is one clear, documented case, not a measured average across all transcripts. We have not yet published agreement statistics for these programs. Teams should measure grader-human agreement on their own reference sets. For a broader look at evaluation evidence, see Evidentiary evaluation for AI certification.
About ALLTIPLY Labs
ALLTIPLY Labs publishes what we learn from running AI systems in production. If you are putting AI scores in front of people whose certification or pay depends on them, talk to us.





