Methodology
How we measure what we publish
Every number on this site traces to a query and a date. That is the whole standard, and it is a deliberately awkward one to hold: it rules out the round figures the category runs on, and it means publishing the measurements that do not flatter us alongside the ones that do.
Two things are published here in full. The rubric an interview is actually graded against, including where that grading is weak. And an evaluation of our own scorer whose headline result is that it does not discriminate well.
What is published
- Published 2 September 2026 · measured 18 August 2026How we evaluated the scorer, and what it cost us to sayWe committed publicly to publishing a blind-grader disagreement rate. That measurement was never produced, and this page leads with that rather than burying it. What we did measure is worse for us: 124 of 141 per-dimension scores are exactly 90 out of 100, on a corpus a label-blind independent rater separates into distinct behavioural groups.
- The rubric in full · figures as of 28 August 2026What is graded, how the bands are set, and where the score is weakestThe whole rubric, the dimension sets actually emitted, the recommendation mix across 177 graded interviews, and 7 stated limits. The limits are on the page itself, not in a footnote.
The sample behind the published figures
177 graded interviews by 71 people, between 26 June 2026 and 28 August 2026. That is a real denominator and a small one. We would rather publish a small true number than a large round one.
The recommendation each of those interviews ended on:
- YES43Overall score of 7.0 or above.
- MAYBE73Overall score from 4.0 up to 7.0.
- NO61Overall score below 4.0.
The query those figures come from, verbatim, so the count is reproducible:
SELECT … FROM ai_bot_interview_sessions WHERE status = 'COMPLETED' AND feedback ? 'overallEvaluation'What we know is weak
These are stated on the rubric page itself rather than in a footnote, and they are restated here because a methodology index that omits the limits is marketing.
The sample is small, and recent
Every figure on this page comes from 177 graded interviews by 71 people between 26 June 2026 and 28 August 2026. That is a real denominator, and it is a small one. Weigh it accordingly — we would rather publish a small true number than a large round one.
The score you see is calibrated, not raw
A published score is not the grader’s raw output. A fixed adjustment is applied before the report is served. We have measured a discontinuity in that adjustment near the top of the range, which means the gap between two already-high scores carries less information than the same gap in the middle of the range. Read a high score as "high", not as a ranking against another high score. The adjustment is being replaced.
The stated seniority label moves the score
We re-graded identical transcripts while changing only the seniority label attached to them. The score moved, and it moved in the direction that flatters the label — a "senior" label should raise the bar, and instead it slightly raised the grade. The effect is small but it is real and repeatable, and we are removing the label from the grading prompt rather than describing it away.
Not every dimension is observed in every interview
An interview that never produced evidence for a dimension has that dimension excluded from the roll-up and the remainder renormalised. It is never scored zero, because a zero would be a measurement we did not take. Behavioral evidence in particular appears in a minority of interviews, so a card can be backed by fewer than five dimensions.
Our own evaluation found the scorer does not discriminate well
We built a corpus of deliberately different candidate behaviours and had a label-blind, different-vendor rater separate them. It did. Our production scorer largely did not — it returned nearly the same number across strata a rater could tell apart, and the archetype built to be confidently wrong was not penalised for it. The full method, the raw corpus and the numbers are published at /methodology/scorer-compression, including the human grading round we promised and did not produce. That page is the source of record for the evaluation; this one is the source of record for the rubric.
Scores are comparable within a rubric, less so across them
The eight dimension sets above are not interchangeable. Two candidates graded on the same set are directly comparable; two graded on different sets are comparable only after the roll-up onto the five canonical dimensions, and that roll-up is a lossy step.
A score is evidence, not a verdict
The transcript is published alongside the score, so a recruiter can check the grade against what was actually said instead of trusting the number. If the two disagree, the transcript is the one that is true.
What we deliberately do not do
- No review scores or star ratings, anywhere — including in this page’s structured data. There is no verifiable review corpus behind them, and a fabricated rating is exactly the thing this page exists to argue against.
- No self-ratings and no keyword or resume matching. A score is produced from an interview that happened, or it is not produced.
- No zero for an unmeasured dimension, and no rating at all on a card with no graded assessments behind it — it shows as unrated, not as a zero.
- No number on this site that does not trace to a query and a date. This page names both.