Report of the AAPOR Task Force on Transitions from Telephone Surveys to Self-Administered and Mixed-Mode Surveys
16
authors
11
chapters
~87k
words
0
recommended replacements
Somewhere in the last decade, a number you rely on stopped being produced the way you think it is. Poll results, consumer confidence, childhood vaccination rates, the health statistics your employer’s insurance is priced against — a great many of them descend from a method that worked beautifully for fifty years and then quietly stopped working.
The interesting part is not that it broke. It’s that the people who broke it wrote the report.
The frame collapsed underneath the method
A telephone survey works on a simple premise: households have phone numbers, phone numbers can be enumerated by area code and exchange, and you can therefore draw a random sample of the population by drawing a random sample of numbers. Random digit dialling. The sample frame — the list you draw from — is a proxy for the country.
Every part of that premise has now failed, and the report walks through the failures in order.
FAILED
Households have phone numbers.
By late 2018, 56.7% of adults and 67.5% of children lived in cell-only households.
FAILED
Numbers can be enumerated by area code and exchange.
Cellular numbers have no listings, and people keep their number when they move.
UNSUPPORTED
So a random sample of numbers is a random sample of the population.
The conclusion the whole method rests on.
Early 2000s
~3%
Late 2018, adults
56.7%
Late 2018, children
67.5%
Meanwhile the share with no home phone service has stayed roughly flat since 2003. This is not a coverage gap. It is a migration.
Then the secondary failures, each of which removes a property the method depended on:
Efficiency shortcut
List-assisted designs dropped number blocks with no listed residential numbers. In the 1990s that excluded under 4% of people, similar to everyone else — except for being more mobile.
No directory
Cellular numbers have no listings, so the stratification trick that made landline sampling affordable has nothing to work with.
Geography detached
People keep their numbers when they move. An area code used to be a place. Now it's a biography.
Nobody answers
Declining response rates, which the report calls a "perfect storm" in combination with the above.
The one I keep returning to is the mobility footnote. The 1990s efficiency shortcut was validated by checking that the excluded group looked like everyone else on the characteristics anyone thought to check. They differed on one: they moved more. That was noted, judged immaterial, and the shortcut shipped. Thirty years later the entire population became the mobile group. A caveat that was true and small at the time of writing does not stay small, and nobody re-runs the check.
1990s — when the shortcut was validated
Dropping the unlisted blocks excluded under 4% of people, similar to everyone else — except for being more mobile.
Noted, judged immaterial, shipped.
Thirty years later
The entire population became the mobile group.
Nobody re-ran the check.
What the report does that I did not expect
I opened this expecting a transition guide. It is not one. It is a catalogue of unresolved problems, and the honesty is structural rather than decorative.
The most striking artifact is the survey the task force ran on organisations that had already made the jump. Data quality was the top motivation — a large majority said the response rates on their old telephone survey were extremely or very important in the decision. And then:
Costs, by contrast, mostly did fall — 13 of the 19 who answered on cost said the mode change reduced them, and only one said the new mode cost more.
A profession reporting that its members switched methods for quality and got a coin flip on quality plus a discount is not the finding anyone set out to publish. It is in there anyway, in a table, with the sample size attached.
I want to be careful with that chart, because it is easy to over-read. Seventeen organisations answering that particular question, self-selected, self-reporting, and a response rate is not data quality — a higher rate can hide worse bias, which is the point I keep having to relearn. The report says as much. But the gap between the motivation and the outcome is real, and nobody buried it.
Answering on response rates
17
How they were recruited
Self-selected
Where the numbers came from
Self-reported
And a response rate is not data quality — a higher rate can hide worse bias.
The part that should worry anyone who quotes a number
Chapter 11 is titled Communicating the Impact of the Change of Modes, and its first question is: how do you talk to the public and data users about a break in the time series?
That question is the whole problem in eight words.
If a survey has run for thirty years by telephone and now runs by web and mail, the new numbers are not comparable to the old ones. People answer differently when a human is asking — more agreeably, less candidly on sensitive topics, differently again on a phone screen versus a laptop. Those are mode effects, and they are not noise you can average away. They are a systematic shift that arrives on the same date as the method change.
So every long-running indicator that made this transition has a discontinuity in it. Somewhere in the series there is a step, and part of that step is the world changing and part of it is the instrument changing, and separating the two is a research project rather than a footnote.
Part of the step is
The thing the survey exists to measure, actually moving.
And part of the step is
People answer differently when a human is asking, and differently again on a phone screen versus a laptop.
The report has a whole chapter on diagnosing and adjusting for exactly this. It does not claim the problem is solved.
What I took from it
Two things, and only one of them is about surveys.
A caveat has a shelf life. The 1990s decision to drop unlisted number blocks was correct, documented, and validated against the evidence available. It became wrong later, without anybody doing anything, because the population moved. Every assumption I’ve ever written into a test plan has the same property: it was true about a system that has since been changed by people who never read my plan. I don’t re-validate assumptions on a schedule. I should.
And this is what an honest post-mortem looks like. Sixteen authors from institutions whose entire business is the credibility of survey data, publishing that their flagship method has structurally failed, that the replacements introduce new errors, that the transition breaks their own historical series, and that they cannot yet tell you which mode to use. No recommendation, because the evidence doesn’t support one. That is a much harder document to write than a transition guide, and it is worth vastly more.
I have written the other kind of report — the one that closes cleanly because closing cleanly was the objective. This is the standard.
What it changes for me practically is small: when I see a poll or an index quoted, my first question is no longer what’s the margin of error. It’s what mode, and did it change. The margin of error describes sampling noise, and nothing else. The frame is the thing that broke.
What I used to ask first
“What’s the margin of error?”
Describes sampling noise, and nothing else.
What I ask first now
“What mode, and did it change?”
The frame is the thing that broke.
Thanks to a course that assigned the profession’s own bad news rather than a textbook chapter about how it used to work.
If you rely on a long-running published indicator: do you know whether it changed collection mode in the last decade? I’d be interested in who can answer that without looking it up.
Comments