TheInterviews Logo
TheInterviews

Methodology · published 2026-09-02

What we promised, and what we found instead

On 6 August 2026 we committed publicly to publishing one number by 15 September: the rate at which blind graders and our AI disagree, broken out by role type. That measurement was not produced, and it is not on this page. The paid grading round it required was deferred by owner decision on 18 August 2026, before a single grade was collected. No person has graded this corpus. There is no disagreement rate here and no role-type breakdown.

We are publishing anyway, because the same work produced a different result and the commitment was to return what we found either way. That result is worse for us than the one we went looking for: on a corpus that an independent rater sorts into distinct behavioural groups, our production scorer returns very nearly the same number for all of them. One correction to how we described this in August: our wording then implied a person had separated these transcripts. That was wrong. The rater is a different vendor’s model, run label-blind, with no sight of our scorer’s output.

The scorer compresses the corpus to one number

47 interview cases carry a frozen score vector across three dimensions — 141 observations in total. 124 of them are exactly 90 out of 100. 32 of the 47 cases carry the identical (90, 90, 90) triple. Technical is 90-or-100 on 45 of 47.

DimensionEvery value it returned (n = 47)DistinctMost common value covers
Technical80 ×1 · 90 ×36 · 95 ×1 · 100 ×9477%
Problem solving80 ×1 · 85 ×1 · 90 ×44 · 93 ×1494%
Communication80 ×1 · 85 ×1 · 90 ×44 · 100 ×1494%

Standard deviations, for anyone recomputing: 4.31, 1.67 and 2.19 respectively — population standard deviations, not sample.

The corpus carries signal the scorer does not read

A flat distribution could mean the cases really are alike. Two independent checks say otherwise, which is what makes this a property of the scorer rather than of the corpus.

First, the defects were planted. Each case was generated to exhibit a named behaviour, and the ones built to be bad score at or above the control group on precisely the dimension they were built to be bad at.

Confidently wrong

n = 9 · attacks technical

92.2

this archetype

89.2

control

All nine score 90 or above on the dimension they are wrong in.

Strong technically, weak at communicating

n = 9 · attacks communication

88.3

this archetype

90.0

control

Seven of the nine are handed exactly 90 — the control constant.

Second, a label-blind rater from a different vendor — with no sight of the archetype labels or of our scorer’s output — sorts the same 47 cases into legible groups: 15 control, 11 rambling-but-correct, 8 confidently wrong, 7 strong-but-poorly- communicated, 6 terse-and-correct. Every one of the 8 it reads as confidently wrong carries a technical score of 90 or above from us.

Its disagreements are informative in their own right. Across all 60 generated cases, 4 of the confidently-wrong candidates read as ordinary control cases — confident wrongness is genuinely hard to spot without checking the domain claims, which is rather the point.

What does move the score is the seniority label

Take one transcript, change nothing in it, and relabel the candidate from Junior to Senior. On our powered arm — 30 sessions, permutation test — the score moves +0.333 on a 0–10 scale going up, and −0.130 going down: a gap of +0.463, p = 0.00042. No session changed band. On this corpus’s smaller arm the movement is technical only: exactly +10 points on 5 of 7 pairs, nothing on the other two, and never any movement at all on problem solving or communication.

We are not going to overclaim what that means. A good interviewer should grade against the level a candidate is being considered for, so a model that shifts with the label is only wrong if it shifts more than an expert would. The direction is established. The magnitude against a human baseline is exactly the question the deferred round was going to answer.

Proxies inside the transcript: what we controlled for, and what we did not

When we published the August post, a reader raised a specific objection: a grader blind to name, school and score can still be anchored by proxies inside the transcript — where the candidate says they worked, how long they say they have been doing this. We committed to reporting which of those controls ran and which did not. That commitment never depended on the grading round, so its deferral does not excuse it.

Controlled for

Employer, product, city, named third parties

Every case passes a seven-check de-identification pass, signed on 17 August by a reader who did not run the generator: no real company, no real product, no city tied to the candidate, no named third party, no project detail specific enough to search for. No employer name survives in any of the 47 published cases. The one case the sweep flagged for a company name is not among them — it is one of the 13 the admission gate rejected on separate grounds, and it ships in the archive still flagged and unsigned.

Controlled for

The stated seniority level

This is the control that found something, and it is the section above: relabel the same transcript and the score moves. It is the reason we treat the untested proxies below as an open risk rather than a formality.

Not controlled for

Explicit tenure claims in the answers

“I have been doing backend engineering for about eight years” and its variants appear in 15 of the 47 published cases. They were never varied and never removed. The de-identification checks are aimed at re-identification, not at anchoring, so a tenure claim passes all seven of them by design. Nothing on this page tells you whether it moves a score.

Not controlled for

Redacting the proxies and re-scoring

Floated rather than promised in the original reply, and gated on scoping the grading round that was deferred on 18 August. The seniority arm shows this scorer is movable by an attribute that is not an answer, so this is owed work rather than an impossible measurement.

The same reply committed to reporting grader-to-grader agreement alongside the disagreement rate, on the reasoning that a bias two raters share shows up as suspiciously high agreement rather than as disagreement. Neither number exists. There is one rater here, and agreement between raters needs at least two; the round that would have supplied the others was deferred. So that objection is unaddressed, not answered — and it is not a remote worry, because our scorer and the rater are both language models and may share the very prior that would produce it.

What is not published here, and why

Stated in full rather than summarised. An evaluation page that argues for rigour while trimming its own limitations would be making the same move it is criticising.

The graded round is deferred, not cancelled

Everything that round needs is built and sitting on our main branch: the frozen and signed corpus, the grading tool with blindness enforced server-side, and the grading instrument generated from the served objects with a parity test against the page. What is missing is the round itself — recruiting three engineers, running the adversarial pilot, and paying for roughly 17.5 hours of grading.

The condition for running it is a stimulus rebuild that produces variance on the scorer side. Against the near-constant on this page, an agreement statistic is undefined — so buying grades today would buy a number that cannot exist, rather than a check on our own work. Fix the compression first, then the comparison becomes answerable.

The pre-registration for that round was fixed in writing before any result was visible, and it contains this clause, which is why the page you are reading exists in the shape it does:

“If the result is unflattering, it is published unchanged. We committed publicly on 2026-08-06 to running these cases and returning raw output either way. A low alpha, poor agreement, or a model label-sensitivity exceeding the human baseline will be reported with the same prominence as a favourable result would have been, in the same report, on the same date.”

Check it yourself

Every number above was recomputed from the frozen corpus on 2026-08-18 and independently re-derived from the raw case files before publication. The raw archive is a single package: all 60 generated cases — including the ones our own admission gate rejected — the seed, the analysis code, the pre-registration verbatim, and a checksum for every file. Re-running the generator with the pinned seed reproduces the same case identities and the same group assignments; it does not reproduce byte-identical transcripts, and never claimed to, because those are sampled from a vendor model.

Ask us for it at support@theinterviews.ai and we will send the package. If you find an error in it, we would rather hear it from you than not hear it.