A workshop activity worth two and a half points, graded mostly on effort. Search two survey databases, find questions that measure something like your construct, write down what you found. I expected to spend the time copying items into a document. I spent it discovering that the thing I want to measure has never been asked about by anyone, in any national survey, ever — and that the one good item I did find was published by an organisation that wasn’t the one whose archive I found it in.
Everyone has searched for something using the only word they had for it, and been handed a pile of results that share the word and none of the meaning. Search for stability and get furniture. Search for delivery and get logistics. You refine, you synonym, you eventually find the thing — or you conclude it doesn’t exist, which is sometimes right and sometimes just means your vocabulary and the index never agreed.
I write acceptance criteria for a living. A criterion is a sentence written by one person and applied by another, and the entire craft is finding where those two readings diverge before it costs anything. A survey item is the same object pointed at strangers. So I went into this workshop with a comfortable assumption: professionals have already written my question better than I would, and my job is to go and collect it.
The assignment
The instruction was specific, and the specificity is the lesson:
1. Identify one concept or construct that you would like to measure in your questionnaire for this class.
3. Search the Roper Center iPoll database to find an example of a survey question that measures something similar to your concept or construct. Note the question in your submission. (If you can’t find anything relevant, please explain briefly what you did).
4. Search the Gallup Analytics database to find an example of a survey question that measures something similar to your concept or construct.
5. Create and submit a document that includes: the questions you identified … and the results. A concrete example of how you would use one or several of these tools to help answer a specific policy question.
Read step 3 again. If you can’t find anything relevant, please explain briefly what you did. That parenthesis is doing more work than the rest of the assignment. Somebody who has run this workshop many times knew that a meaningful share of students would search honestly and come back with nothing, and built a route for reporting a null result instead of manufacturing a hit.
I did not appreciate that clause until I needed it.
What I was trying to measure
My project studies a distributed software delivery team and the recurring meetings it runs — the pre-deployment readiness review, the morning call where yesterday’s deployments get discussed with leadership in the room. Three constructs:
- Techno-emotional Authenticity — whether a practitioner’s stated confidence matches their actual confidence. Whether someone says I don’t think this is ready when their director is listening.
- Leadership Inquiry Style — whether the team reads leadership as curious about failure or as looking for who caused it.
- True Change Failure Rate — official failures plus the near-misses quietly fixed before anyone logged them.
That third one is the whole reason for the study. The industry-standard metric counts recorded failures. If a team resolves problems silently, the metric improves while the system does not, and the number becomes a measure of reporting behaviour wearing the costume of a measure of stability.
I searched three archives: Roper Center iPoll, Gallup Analytics, and Pew’s own published questions.
Before the results, it is worth knowing what these three organisations actually are, because I had been treating them as interchangeable brand names on a search page. They are not. One is a pollster, one is an archive, one is a fact tank that fields its own surveys — and two of them were founded by men who made their reputations on the same election and lost them on the same one twelve years later.
Three names on one search page · three different jobs
Who fielded the question, who kept it, and who published the method.
Gallup
the pollster · founded 1935
George Gallup founded the American Institute of Public Opinion in 1935. In 1936 he publicly challenged The Literary Digest — which had a vastly larger sample and predicted Alf Landon — and called the Roosevelt landslide correctly with a far smaller, deliberately structured one. That killed the idea that volume alone was the answer. It did not settle what should replace it: Gallup's method then was quota sampling, which is exactly what failed in 1948.
Roper Center
the archive · founded 1947
Elmo Roper ran the Fortune Survey from 1935 to 1950 and also called 1936 correctly. In 1947 he donated roughly 250,000 interviews to Williams College, creating the first social science data archive — on the belief that polls were a record future historians would need. It became a formal centre in 1957, with George Gallup on its board; moved to the University of Connecticut in 1977 and to Cornell in 2015.
Pew Research Center
the fact tank · 1990 → 2004
The youngest by decades — 43 years after the archive, 55 after Gallup. It designs and fields its own probability-based surveys, including the American Trends Panel. It began in 1990 as the Times Mirror Center for The People & The Press, was taken over and renamed by The Pew Charitable Trusts in 1996, and was established as a subsidiary under its current name in 2004. Its distinguishing habit is documentation: publishing exact wording, toplines, and the split-ballot experiments where its own phrasing moved the result.
There is a detail in that history I have not been able to stop thinking about. In 1948, Gallup, Roper and the rest of the field predicted Thomas Dewey would beat Harry Truman. Gallup had stopped polling roughly two weeks out, on the strength of an early lead. They were all wrong, publicly, in a way that nearly ended the industry twelve years after it had been made.
After the failure, CBS handed Elmo Roper an abacus on air. It was a joke at his expense. He kept it, and then he gave it to his own archive — explicitly so the field would be reminded that polls falter and the work is to keep improving the instrument.
A man who founded an archive to preserve his profession’s results also chose to preserve the object mocking its worst failure, and filed it alongside the data. I have written a lot of incident reports. I have never once archived the joke someone made at my expense as part of the permanent record.
What I searched for, and what came back
Three constructs · three national archives · no usable item
Two searches returned a keyword match with a construct mismatch. The third returned nothing at all.
| My construct | Closest fielded item | Why it doesn't transfer |
|---|---|---|
| Techno-emotional Authenticity honesty about technical readiness | "Is it important that the people you get news from display each of these personal traits in their work? … Authenticity (being their true selves)" Pew American Trends Panel, Wave 168, April 2024 — found via Roper iPoll. 9,397 answered: Yes 82% · No 10% · Not sure 8%. | Same word, different object. It measures what a stranger wants from a broadcaster — not whether an engineer will say "I don't think this is ready" with their director on the call. And 82% agreement leaves almost no variance. There is nothing left to correlate against. |
| Leadership Inquiry Style how leaders respond to failure | "Are you currently responsible for managing others at your workplace — that is, does at least one person answer or report directly to you?" Gallup Poll Social Series, August 2024, Work & Education. 597 answered: Yes 43.26% · No 56.74%. | Doesn't measure my construct at all — searching "leadership" in the World Poll returned political leaders, so I moved to a different collection and found this instead. What it is good for is routing. |
| True Change Failure Rate logged failures + silent recoveries | Nothing. No item, in any of the three archives. | Public-opinion archives hold what populations think, not what organisations do. Asking them for firm-level operational behaviour is asking the wrong instrument — the pointer was to the World Management Survey, an entirely different literature. |
Keyword match, construct mismatch
The first row is the one I want to sit with, because it is the failure I am professionally supposed to catch.
I searched authenticity. I got back a well-built, nationally fielded, enormous-sample item about authenticity. Every surface signal said hit. And it is measuring whether people want news presenters to seem like their true selves — an attitude a respondent holds about strangers on a screen, at no cost to themselves.
My construct is whether a specific engineer, in a specific meeting, with their reporting line present and a deployment window closing, will volunteer that they are not confident. Those two things share a word and share nothing else. One is a preference; the other is a behaviour under professional risk.
This has an exact analogue in my working life, and it is embarrassing to write down. A test that asserts on a string rather than on the behaviour passes for years, gives a green tick every night, and is measuring the label instead of the thing. Everyone has one. Mine was in a survey archive, and I nearly copied it into an assignment because the search results told me I had succeeded.
The 82% is the second tell, and I have to be careful about it, because the obvious version of this point is wrong. At Pew’s sample size the minority is roughly 1,700 people — ample for subgroup analysis, and the item retains most of its theoretical variance. Nothing is inert about it there.
The ceiling problem is mine, not Pew’s. My target population is a census of about forty engineers. Eighty-two per cent agreement in a national sample is a finding; in a sample of forty it is seven dissenters, and any relationship I claimed to detect would be resting on them. Item selection is not a property of the item — it is a property of the item and the sample you intend to field it to, and I had been evaluating candidates on relevance alone without once asking whether it could still move at my n.
The Pew question I found in Roper’s archive
Now the part that gave this piece its title, and the thing I didn’t know before the workshop.
The authenticity item is a Pew question. It comes from Pew’s American Trends Panel, Wave 168, fielded April 2024. I did not find it on Pew’s site. I found it inside the Roper Center’s archive, in a collection Roper maintains, having gone to Roper because that is where the assignment sent me.
I had been treating “the archive” and “the publisher” as one thing, the way you might treat a package registry and a package author as one thing until the day the distinction costs you something. They are separate roles: somebody designed and fielded the instrument, and somebody else curates, stores, indexes and re-serves it. The chain has at least two links, and each has its own conventions, its own search behaviour, and its own way of deciding what metadata survives.
That matters practically. Search Roper and you reach across many organisations at once, which is why my search worked at all — but you inherit Roper’s indexing of Pew’s item. Search Pew directly and you get their full topline, their methodology statement, and the trend where the item has been repeated. The provenance and the discovery surface are different systems, and the one you searched determines what you can see about what you found.
What the publisher gives you that the index doesn’t
This is where Pew’s own material earned its place in my notes, because the Center documents its instruments unusually well — including the experiments where the wording, not the public, produced the result.
Pew split-ballot results · two forms, one difference
| The change | What moved |
|---|---|
| One appended clause January 2003 | Support for military action in Iraq ran 68% favor / 25% oppose. Adding "even if it meant that U.S. forces might suffer thousands of casualties" produced 43% / 48% — a majority inverted by one subordinate clause. |
| A narrower label January 2002 | Domestic policy versus foreign policy: 52% / 34%. Replace "foreign policy" with "the war on terrorism" and it becomes 33% / 52%. The option didn't change — its name did. |
| Listing the answer Post-election 2008 | 58% named the economy the most important issue when it appeared on a list; 35% volunteered it when it didn't. And off-list answers ran 8% closed against 43% open — most of the public's answer lived outside the options. |
| Context from the item before December 2008 | 88% said they were dissatisfied with the country's direction when asked right after a presidential approval question, against 78% without it. The respondent didn't change; the question before them did. |
Two things follow, and the second is the one I keep.
The obvious one: wording is not packaging around the measurement, it is the instrument, and I had been treating item drafting as a copy-editing pass. The last row is worse than the others — nothing about the item changed at all, only its neighbour, and the answer moved ten points. That is state leaking between test cases, in a suite that runs against a human being in a fixed order with no way to reset.
The one I keep: none of this is visible from the index. The archive gave me the item, the sample size and the topline. It did not tell me that this organisation has spent forty years demonstrating that items like it are fragile. If I had taken the authenticity question at face value from the search results, I would have inherited an instrument without inheriting the warnings that come with it — which is the definition of a dependency you have not read.
The item I kept, for a job it wasn’t built for
The Gallup management question doesn’t measure my construct and never will. But do you have at least one direct report is a clean, unambiguous, nationally validated binary — and my questionnaire needs to route people-leads away from sections that only make sense to engineers.
So I kept it as a screener. Not as a measure of anything; as a branch condition. Its value to me is entirely in its precision as a sorting question, which is the property the original authors also needed and got right.
I like this outcome more than a construct match would have given me. It is the realistic version of reuse: you rarely find the thing you were looking for, and you occasionally find a well-made component that solves an adjacent problem you had stopped thinking of as a problem.
Where the construct actually lives
The null row had an answer, and it wasn’t in a public-opinion archive at all. Firm-level management practice — how organisations respond to poor performance, whether failure triggers inquiry or blame — is measured by the World Management Survey, in a literature I hadn’t been reading because I had been searching by my words instead of by my field.
The general form of my mistake has two layers, and only one of them is mine. The assignment sent me to Roper and Gallup, so the domain was chosen for me — and public-opinion archives were never going to hold firm-level operational data. That is a domain mismatch, and noticing it is the assignment’s point.
What is mine is the second layer: inside that domain I searched for the words I already owned, took a keyword hit as a construct match, and treated the absence as evidence that the measurement didn’t exist anywhere. A vocabulary problem reported as an evidence problem.
What I would search first, next time
- Search the respondent’s words, not the construct’s name. The abstract term is for reading literature; the plain phrase is for querying archives.
- Test the match on the object, not the word. Ask what the respondent is being asked to do, to whom, at what cost. If the answer differs from your setting, a shared keyword is a false positive.
- Check the marginals against your own sample size, not the published one. A skewed item that works nationally can leave you a handful of dissenters at the n you can actually field.
- Know whether you are searching a publisher or an index, and go to the publisher for the methodology, topline and trend before you rely on anything.
- Read what precedes the item in the fielded instrument. After a ten-point context effect, position is not a formatting detail.
- Report the null. If the construct isn’t there, that is a finding about the fit between your field and the archive, and it is more useful than a keyword match dressed up as a hit.
What I took away
I went in to collect questions and came out having learned that my central measure has never been asked about by anyone, anywhere, in a national survey — because it belongs to a setting those surveys were never built for. That is a real result, and the assignment had a clause reserved for it, which tells you the people who wrote it expected it.
The habit that changed is small. When something matches my search, I now ask what the respondent was actually being asked to do, and compare that to my own situation, before I look at anything else. Relevance signals lie. They lie in the same direction every time, because search rewards the vocabulary you brought.
And I stopped assuming that finding the item means inheriting the knowledge behind it. The archive hands you a sentence and a percentage. The organisation that wrote the sentence may have spent decades documenting how easily sentences like it break — and none of that travels with the copy you found.
Thank you for reading. If you have ever searched for prior art and concluded there wasn’t any: how confident are you that you were searching in the language of the people who would have written it?
Comments