Grading the Grader: Reading the Duolingo English Test's Own Validity Evidence

📍 Psychometrics 📅 August 12, 2026

STATUSPublished ENVresearch-methods AUDITv1.0 VIEWS

Reading a vendor's own validity evidence taught me to separate what a study proved from what its summary is used to claim.

Someone has handed you a PDF with a chart in it and asked you to approve something. It was professionally typeset, it cited real studies, and it was written by the party who wanted the decision to go one way. You had to decide what to believe.

The Duolingo English Test is an unusually good specimen of that document, because the evidence behind it is real and the sentence it gets compressed into is not.

Reported concurrent validity

DET scores against the incumbent tests

vs TOEFL iBT

r = .77

N = 2,319 · LaFlair & Settles, 2020

vs IELTS

r = .78

N = 991 · LaFlair & Settles, 2020

Shared variance at r = .77

≈ 59%

The other 41% is not noise

Earlier work (Brenzel & Settles, 2017) reported r = .71 vs TOEFL iBT and r = .70 vs IELTS on smaller samples.

Those are strong numbers. They are also published, replicated across studies, and accompanied by a technical manual that states its methods. That is more transparency than most enterprise software vendors offer for products that make bigger decisions about people’s lives.

And the sentence they turn into — “the DET is equivalent to IELTS” — is not something any of them established.

What r = .77 does and doesn’t buy you

A correlation of .77 means roughly 59% shared variance. Which means about 41% of the variation in how people score is not shared between the two tests.

That remainder is invisible in aggregate and decisive individually. At the population level, the two tests rank candidates similarly. At the level of one applicant sitting on an admissions threshold, the residual is the whole story: it is the difference between a place and a rejection, and it is exactly the person about whom the aggregate statistic says the least.

This is the pattern that recurs in every vendor evidence document I have read since:

What the research establishedWhat it gets quoted as
Scores correlate strongly with an established measureThe two tests are interchangeable
It held in the sample studiedIt holds for your applicants
The instrument measures the construct reliablyIt predicts the outcome you actually care about
Evidence supports one specific stated useThe test is valid, full stop
Every claim on the left is defensible. Every claim on the right is what survives into a procurement meeting.

The bottom row is the one that matters. Validity is a property of an interpretation for a specific use — not a badge a test earns once and carries around. “This test is valid” has no truth value. “These scores support this decision, for this population, at this level of consequence” does.

The detail I’d want a buyer to notice

Buried in the methodology: the concordance sample did not come entirely from official score reports. Only a few hundred official records were available for each comparison — one write-up puts it at 304 — so the set was topped up with self-reported TOEFL and IELTS results, volunteered by test takers at the end of a session, to reach the recommended minimum sample size of around 1,500.

So the headline N is mostly people telling Duolingo what they scored elsewhere. Disclosing that is good practice; many vendors would not have. But it changes who is in the sample. Volunteering a second score requires having taken a second test, remembering it, and choosing to share it — a self-selected group, and not obviously the marginal applicant an admissions office is actually worried about.

The strongest evidence that this mattered comes from Duolingo’s own research team. Later work (Cardwell et al., 2024) treats the combination of self-reported and official scores as a methodological problem needing a purpose-built solution, and states plainly that the two sources might yield different concordance outcomes. The current subscore concordance tables, built from 2024–25 data, use official score reports only — 1,943 test takers for IELTS Academic, 948 for TOEFL iBT.

That is the vendor quietly correcting itself, four years on. Which is what good research organisations do, and also exactly why the earlier figure should never have hardened into “equivalent to IELTS.” A published limitation and an acted-upon limitation are different things, and the gap between them is where evidence stops constraining decisions.

The five questions I ask now

Valid for what decision?

If the paper never names one, it cannot have supported it.

Who is in the sample — and how did they get there?

Recruitment method is a filter. Absence never announces itself.

What is the benchmark inheriting?

Correlating against an incumbent imports its flaws with its authority.

What does the unexplained variance cost, and who pays it?

Aggregate agreement says least about the person at the threshold.

Does the headline survive the methods section?

Read the summary last. It is the most edited sentence in the document.

Reading it backwards — conclusions first, then checking each against the methods — found everything a forward read had waved through. Going forward, you agree paragraph by paragraph because each step follows the last. The document’s own momentum does the persuading.

Why I don’t want this read as a takedown

I am a non-native English speaker who has been on the wrong end of a language score. A test that is cheap, takeable at home, and finished in an hour is not a small thing; it removes a barrier that had nothing to do with anyone’s English.

And the identical access argument, made without the evidence underneath it, is exactly how a bad test gets adopted at scale.

Both stay true. Refusing to collapse them into one opinion is the discipline — holding “this is genuinely good work” and “that claim is bigger than its evidence” in the same paragraph, without needing either to cancel the other.

The strongest evidence I read that year still did not answer the only question that mattered: valid for what?


Sources: Validity, Reliability, and Concordance of the Duolingo English Test · Duolingo English Test Technical Manual · Practical considerations when building concordances between English tests (Cardwell et al., 2024) · Additional details for score comparison tables · Examining the predictive validity of the Duolingo English Test (Isaacs et al., 2023)

← All articles

Engage Comments & discussion

Comments