Calibrate AI scores against your own reviewers
QA calibration is checking that everyone who scores conversations, people and software, applies the scorecard the same way. Before you rely on AI scores, score the same conversations with your reviewers and measure how closely they agree, overall and criterion by criterion. Merivex has this built in.
Step by step
How to calibrate AI scoring
Choose a calibration set
Pick conversations that cover your main contact reasons and include good, average and poor examples. A set drawn only from easy conversations will overstate agreement.
Score it with your reviewers
Have reviewers score each conversation against the current scorecard, overall and per criterion, without seeing the AI scores first.
Import the reviewer scores
In Merivex, upload a CSV with a transcript column, an overall reviewer score from 0 to 100, and a ref_<criterion> column for each criterion you scored, for example ref_empathy. Up to 500 rows per batch.
Read agreement overall
The report shows the average gap between AI and reviewer scores, whether the AI tends to score higher or lower, the share of conversations within 5 and within 10 points, and how closely the two sets of scores move together.
Read agreement per criterion
Overall agreement can hide a criterion that is consistently off. The per-criterion view shows the average gap and direction for each one, and the report lists the conversations where the two disagree most.
Fix, then re-run
Where a criterion disagrees, reword its definition, adjust the scorecard or keep that criterion under human review. Run a new batch after each change and compare.
What the report measures
Calibration metrics in Merivex
Average difference
The mean absolute gap between AI and reviewer scores, on a 0 to 100 scale.
Bias
The average signed gap, which shows whether AI scores run higher or lower than your reviewers'.
Within 5 and 10 points
The share of conversations where AI and reviewer scores are close.
Correlation
Whether AI and reviewer scores rise and fall together across the set.
Per criterion
Average gap and bias for each criterion you provided reviewer scores for.
Largest disagreements
The conversations where the scores differ most, to review side by side.
Questions
QA calibration, answered
What is QA calibration in a contact center?
QA calibration is the practice of having several evaluators score the same conversations and comparing the results, so everyone applies the scorecard the same way. With AI scoring, the AI is one of the evaluators being calibrated.
How accurate is Merivex compared with human reviewers?
No accuracy figure is published, because none has been independently established, and agreement depends on your scorecard and your conversations. Calibration measures it on your own data instead.
How many conversations do I need to calibrate?
Enough to cover your main contact reasons and a range of quality levels. A few dozen show large gaps; a few hundred give a steadier per-criterion picture. Each batch can hold up to 500 conversations.
How often should we recalibrate?
Whenever the scorecard changes, and on a regular schedule as conversations and policies drift. A new batch can be run at any time.
Keep reading
Measure agreement on your own conversations
Export conversations your reviewers have scored, import them with their scores, and read the agreement criterion by criterion.