The Telephone Survey Is Dying: Reading the Task Force Report on What Replaces It

📍 Survey Research Methodologies 📅 August 23, 2026

STATUSPublished ENVsurvey-mode AUDITv1.0 VIEWS

A task force report on a dying method taught me what an honest failure post-mortem looks like when it's written by the profession that failed.

Report of the AAPOR Task Force on Transitions from Telephone Surveys to Self-Administered and Mixed-Mode Surveys

16

authors

11

chapters

~87k

words

0

recommended replacements

Census Bureau, Pew, Gallup, NORC, Westat, SSRS, four universities and two more federal statistical agencies, on one author list.

Somewhere in the last decade, a number you rely on stopped being produced the way you think it is. Poll results, consumer confidence, childhood vaccination rates, the health statistics your employer’s insurance is priced against — a great many of them descend from a method that worked beautifully for fifty years and then quietly stopped working.

The interesting part is not that it broke. It’s that the people who broke it wrote the report.

The frame collapsed underneath the method

A telephone survey works on a simple premise: households have phone numbers, phone numbers can be enumerated by area code and exchange, and you can therefore draw a random sample of the population by drawing a random sample of numbers. Random digit dialling. The sample frame — the list you draw from — is a proxy for the country.

Every part of that premise has now failed, and the report walks through the failures in order.

FAILED

Households have phone numbers.

By late 2018, 56.7% of adults and 67.5% of children lived in cell-only households.

FAILED

Numbers can be enumerated by area code and exchange.

Cellular numbers have no listings, and people keep their number when they move.

UNSUPPORTED

So a random sample of numbers is a random sample of the population.

The conclusion the whole method rests on.

The conclusion was never stronger than the two premises underneath it.

Early 2000s

~3%

Late 2018, adults

56.7%

Late 2018, children

67.5%

Meanwhile the share with no home phone service has stayed roughly flat since 2003. This is not a coverage gap. It is a migration.

The frame didn't shrink. The population walked off it.

Then the secondary failures, each of which removes a property the method depended on:

Efficiency shortcut

List-assisted designs dropped number blocks with no listed residential numbers. In the 1990s that excluded under 4% of people, similar to everyone else — except for being more mobile.

No directory

Cellular numbers have no listings, so the stratification trick that made landline sampling affordable has nothing to work with.

Geography detached

People keep their numbers when they move. An area code used to be a place. Now it's a biography.

Nobody answers

Declining response rates, which the report calls a "perfect storm" in combination with the above.

Four separate failures. Each one individually survivable. Together, a different method.

The one I keep returning to is the mobility footnote. The 1990s efficiency shortcut was validated by checking that the excluded group looked like everyone else on the characteristics anyone thought to check. They differed on one: they moved more. That was noted, judged immaterial, and the shortcut shipped. Thirty years later the entire population became the mobile group. A caveat that was true and small at the time of writing does not stay small, and nobody re-runs the check.

1990s — when the shortcut was validated

Dropping the unlisted blocks excluded under 4% of people, similar to everyone else — except for being more mobile.

Noted, judged immaterial, shipped.

Thirty years later

The entire population became the mobile group.

Nobody re-ran the check.

The caveat never changed. The population it described did, and the decision stayed on the shelf.

What the report does that I did not expect

I opened this expecting a transition guide. It is not one. It is a catalogue of unresolved problems, and the honesty is structural rather than decorative.

The most striking artifact is the survey the task force ran on organisations that had already made the jump. Data quality was the top motivation — a large majority said the response rates on their old telephone survey were extremely or very important in the decision. And then:

UP · 7
DOWN · 5
SAME · 5

Costs, by contrast, mostly did fall — 13 of the 19 who answered on cost said the mode change reduced them, and only one said the new mode cost more.

The stated reason for moving was data quality. The reliable result was cheaper data collection.

A profession reporting that its members switched methods for quality and got a coin flip on quality plus a discount is not the finding anyone set out to publish. It is in there anyway, in a table, with the sample size attached.

I want to be careful with that chart, because it is easy to over-read. Seventeen organisations answering that particular question, self-selected, self-reporting, and a response rate is not data quality — a higher rate can hide worse bias, which is the point I keep having to relearn. The report says as much. But the gap between the motivation and the outcome is real, and nobody buried it.

Answering on response rates

17

How they were recruited

Self-selected

Where the numbers came from

Self-reported

And a response rate is not data quality — a higher rate can hide worse bias.

The gap between the stated motive and the outcome is real. The sample is small enough that it has to be said out loud.

The part that should worry anyone who quotes a number

Chapter 11 is titled Communicating the Impact of the Change of Modes, and its first question is: how do you talk to the public and data users about a break in the time series?

That question is the whole problem in eight words.

If a survey has run for thirty years by telephone and now runs by web and mail, the new numbers are not comparable to the old ones. People answer differently when a human is asking — more agreeably, less candidly on sensitive topics, differently again on a phone screen versus a laptop. Those are mode effects, and they are not noise you can average away. They are a systematic shift that arrives on the same date as the method change.

So every long-running indicator that made this transition has a discontinuity in it. Somewhere in the series there is a step, and part of that step is the world changing and part of it is the instrument changing, and separating the two is a research project rather than a footnote.

A long-running indicator with a step at the point where its collection mode changed Schematic, not data. A series runs at one level while collected by telephone, then jumps to a higher level at the moment collection switches to web and mail, and continues at the new level. The size of the jump is marked as two unknown parts: real change in the world, and the effect of the instrument. collected by telephone collected by web and mail the mode change

Part of the step is

The thing the survey exists to measure, actually moving.

And part of the step is

People answer differently when a human is asking, and differently again on a phone screen versus a laptop.

Schematic, not data. Separating the two parts is a research project, not a footnote.

The report has a whole chapter on diagnosing and adjusting for exactly this. It does not claim the problem is solved.

What I took from it

Two things, and only one of them is about surveys.

A caveat has a shelf life. The 1990s decision to drop unlisted number blocks was correct, documented, and validated against the evidence available. It became wrong later, without anybody doing anything, because the population moved. Every assumption I’ve ever written into a test plan has the same property: it was true about a system that has since been changed by people who never read my plan. I don’t re-validate assumptions on a schedule. I should.

And this is what an honest post-mortem looks like. Sixteen authors from institutions whose entire business is the credibility of survey data, publishing that their flagship method has structurally failed, that the replacements introduce new errors, that the transition breaks their own historical series, and that they cannot yet tell you which mode to use. No recommendation, because the evidence doesn’t support one. That is a much harder document to write than a transition guide, and it is worth vastly more.

I have written the other kind of report — the one that closes cleanly because closing cleanly was the objective. This is the standard.

What it changes for me practically is small: when I see a poll or an index quoted, my first question is no longer what’s the margin of error. It’s what mode, and did it change. The margin of error describes sampling noise, and nothing else. The frame is the thing that broke.

What I used to ask first

“What’s the margin of error?”

Describes sampling noise, and nothing else.

What I ask first now

“What mode, and did it change?”

The frame is the thing that broke.

The margin of error was never measuring the failure. It was measuring around it.

Thanks to a course that assigned the profession’s own bad news rather than a textbook chapter about how it used to work.

If you rely on a long-running published indicator: do you know whether it changed collection mode in the last decade? I’d be interested in who can answer that without looking it up.

← All articles

Engage Comments & discussion

Comments