TheInterviews Logo
TheInterviews
Scoring

How an Interview
Actually Gets Scored

An AI interviewer runs a live voice interview for your target role. The session is graded against a rubric chosen for that interview type — four to six named dimensions, an overall score out of 10, and a YES / MAYBE / NO recommendation. Those dimensions are then rolled up onto five canonical ones, which is what makes two different interviews comparable on one Profile Card. The transcript is published next to the score.

Every figure on this page is drawn from 177 graded interviews by 71 people, between 26 June 2026 and 28 August 2026. As of 28 August 2026.

The Path From Answer to Score

Four steps, in this order, every time.

  1. Step 1

    The interview happens

    A real-time voice interview with an AI interviewer, in the browser, calibrated to the role. Questions adapt to the answers; nothing is pre-scripted and nothing is graded from a resume.

  2. Step 2

    The session is graded against a rubric

    The rubric is chosen by interview type, not per candidate. It names four to six dimensions, each scored, and produces an overall score out of 10 with written evidence for every judgement.

  3. Step 3

    The score is calibrated and banded

    A fixed adjustment is applied, then the recommendation is derived from the adjusted score. The adjustment is the same for every candidate; it is not tuned per person, per customer, or per plan.

  4. Step 4

    It rolls up onto the Profile Card

    Per-interview dimensions map onto the five canonical dimensions, weighted by role family, so a backend engineer and an account executive are not measured against one rubric.

The Rubrics, and How Often Each Is Used

These are the dimension sets our grader actually emitted across the sample — not a set designed on paper. The counts sum to all 177 graded interviews.

Scoring rubrics in use, their dimensions, and the number of graded interviews each covered between 26 June 2026 and 28 August 2026.
RubricDimensions gradedInterviews
Technical screenTechnical knowledge · Problem solving · Coding · Communication72
Applied depthTechnical knowledge · Conceptual clarity · Trade-off analysis · Edge cases52
Hands-on codingCode correctness · Code quality · Efficiency · Problem-solving approach20
Experience screenCoding proficiency · Conceptual depth · Project experience · Scenario judgment12
BehavioralCollaboration · Communication · Self-awareness · STAR examples11
LeadershipLeadership · Decision making · Accountability · Stakeholder management5
Role and market fitProfessional communication · Expectations clarity · Market awareness · Flexibility4
System designRequirements scoping · Topology fit · Data-storage choices · Trade-off justification · Bottlenecks and failure modes · Design communication1
All graded interviews, 26 June 202628 August 2026177

What the Bands Mean

The overall score runs 0–10 and the recommendation is derived from it by fixed thresholds — no human moves a candidate between bands.

YES

43

of 177 graded interviews

Overall score of 7.0 or above.

MAYBE

73

of 177 graded interviews

Overall score from 4.0 up to 7.0.

NO

61

of 177 graded interviews

Overall score below 4.0.

We publish this mix rather than a success rate because the mix is the honest shape of it: most interviews do not end on YES.

The Five Dimensions a Profile Card Rolls Up To

Each is scored 0–100 and weighted by role family. A dimension an interview never observed is excluded and the rest renormalised — never scored zero.

Technical

Correctness and depth of knowledge, including the design-level equivalents — trade-off analysis, topology fit, data-storage choices.

Problem-Solving

How an unfamiliar problem is approached: scoping, decomposition, and what happens when the first idea does not work.

Communication

Clarity and structure, and whether the reasoning behind an answer is made visible rather than left implied.

Code Quality

For interviews where code is actually produced — readability, structure, and handling of the cases the happy path misses.

Behavioral

Collaboration, ownership, and how past situations are described when the detail is pressed for.

Where This Score Is Weakest

Findings from our own measurements. They are here because a rubric page that hides them is not a rubric page, and because you will find them anyway. The evaluation of the scorer itself is published separately, in full, at our scorer-compression methodology.

The sample is small, and recent

Every figure on this page comes from 177 graded interviews by 71 people between 26 June 2026 and 28 August 2026. That is a real denominator, and it is a small one. Weigh it accordingly — we would rather publish a small true number than a large round one.

The score you see is calibrated, not raw

A published score is not the grader’s raw output. A fixed adjustment is applied before the report is served. We have measured a discontinuity in that adjustment near the top of the range, which means the gap between two already-high scores carries less information than the same gap in the middle of the range. Read a high score as "high", not as a ranking against another high score. The adjustment is being replaced.

The stated seniority label moves the score

We re-graded identical transcripts while changing only the seniority label attached to them. The score moved, and it moved in the direction that flatters the label — a "senior" label should raise the bar, and instead it slightly raised the grade. The effect is small but it is real and repeatable, and we are removing the label from the grading prompt rather than describing it away.

Not every dimension is observed in every interview

An interview that never produced evidence for a dimension has that dimension excluded from the roll-up and the remainder renormalised. It is never scored zero, because a zero would be a measurement we did not take. Behavioral evidence in particular appears in a minority of interviews, so a card can be backed by fewer than five dimensions.

Our own evaluation found the scorer does not discriminate well

We built a corpus of deliberately different candidate behaviours and had a label-blind, different-vendor rater separate them. It did. Our production scorer largely did not — it returned nearly the same number across strata a rater could tell apart, and the archetype built to be confidently wrong was not penalised for it. The full method, the raw corpus and the numbers are published at /methodology/scorer-compression, including the human grading round we promised and did not produce. That page is the source of record for the evaluation; this one is the source of record for the rubric.

Scores are comparable within a rubric, less so across them

The eight dimension sets above are not interchangeable. Two candidates graded on the same set are directly comparable; two graded on different sets are comparable only after the roll-up onto the five canonical dimensions, and that roll-up is a lossy step.

A score is evidence, not a verdict

The transcript is published alongside the score, so a recruiter can check the grade against what was actually said instead of trusting the number. If the two disagree, the transcript is the one that is true.

What We Deliberately Do Not Do

  • No review scores or star ratings, anywhere — including in this page’s structured data. There is no verifiable review corpus behind them, and a fabricated rating is exactly the thing this page exists to argue against.
  • No self-ratings and no keyword or resume matching. A score is produced from an interview that happened, or it is not produced.
  • No zero for an unmeasured dimension, and no rating at all on a card with no graded assessments behind it — it shows as unrated, not as a zero.
  • No number on this site that does not trace to a query and a date. This page names both.

Read the Rubric.
Then Take the Interview.