This is the construction log for a psychometric measure my group built from scratch for a graduate psychometrics course — the actual steps, in order, the wrong turns, and what the results looked like once we stopped being able to round anything up.
Everyone has had the thought: this feeling has a name, and somebody should be able to measure it. Most of us never find out how hard that actually is. This project put three of us on the other side of that question for one semester — pick a psychological construct nobody had measured cleanly, define it precisely enough to write questions about it, defend every one of those questions when someone else read them closely, and then do the same critical reading to someone else’s work in return. We chose catastrophizing: the tendency to respond to a bad event by generalizing and ruminating on the worst possible version of it.
The goal
Design an original psychometric instrument — the Catastrophizing Scale (CS) — from a literature review through item-writing, peer review (given and received), data collection, and a reliability/validity analysis, in one semester, as a team of three.
Simple as a rubric line. Negotiated in practice, the same way most builds are.
The setup: what we were working with
- Three of us, no prior overlap in how we write test items
- One semester, split across ten formal project steps
- Qualtrics for survey deployment, SPSS v29 for analysis
- An existing literature that had already built scales for pain catastrophizing, insomnia catastrophizing, even breathlessness catastrophizing — but nothing general-purpose, independent of which body system it was attached to (Sullivan et al., 1995; Correia et al., 2022)
- A working definition we had to lock in before writing a single question: catastrophizing is a pattern of responding to negative events that leads an individual to generalize and ruminate on the worst possible outcomes of those events
The word doing the real work in that sentence is pattern — a standing tendency, not a one-off reaction. Every item written afterward had to ask about a disposition, not a moment. That choice would get tested directly a few steps later, when a reviewer read our items the opposite way.
The steps, as they actually happened
1. Define the construct, in writing, before writing a single item. The assignment’s Step 1 rubric was blunt about sequencing: the name of your proposed construct of study and its definition; a one to three sentence definition/operationalization; a brief description of how your group’s definition is similar to and different from how your construct might be defined in the literature. No items yet — just a claim about what a “yes” and a “no” would each mean, on the record, before anyone got to be clever about wording.
2. Draft items independently, in three different formats. Each of us brought a different approach to the same construct: Likert-type agreement scales, forced-choice ambiguous-scenario items, and true/false statements. One of my own forced-choice items dropped the test-taker into a scenario and asked which interpretation felt closest to instinct:
You overhear your colleagues whispering and laughing in the office, and you catch them looking in your direction. What do you think they're saying?
— They're probably gossiping about how my incompetence is affecting the entire team's performance.
— They must be ridiculing me for not dealing with pending work sooner.
— They're likely discussing how I am sinking the ship.
None of the three options is neutral, on purpose — a forced-choice item like this only works if every answer still lets the catastrophizer catastrophize, just along a different axis.
What I’d do differently: write a fourth, deliberately mild option for at least one scenario, to check whether test-takers who don’t catastrophize actually pick it — as drafted, there was no clean way to score someone who isn’t the construct.
3. Merge three formats into one instrument — the actual hard part. Reconciling three separately-drafted approaches into a single scale meant folding overlapping items into one general “I” statement for consistency of tone, and picking one representation whenever two formats aimed at the same underlying worry. We landed on an eleven-item, five-point Likert scale (strongly agree to strongly disagree), scored so an average score under 3 indicated some degree of catastrophizing. No items were reverse-coded at this stage — a gap that would get named for us in the next step.
4. Submit the draft for peer review — and get read closely. Another group of three read our eleven items and sent back a numbered, specific critique. The most substantive point wasn’t about wording at all:
Our main concern is that you may be measuring other constructs on top of catastrophizing, such as social anxiety and pessimism.
That’s a challenge to the instrument’s whole boundary, not a typo to fix. Underneath it were three narrower critiques: items 1 through 3 were flagged as double-barreled — asking two things inside one question. Item 5 — “I worry that because of things that have occurred so far, my life has been ruined” — was flagged as vague and oddly past-tense for a scale meant to anticipate disaster. And somewhere in the same round — the letter we wrote back frames it as something we identified ourselves under a “what we will work on” heading, rather than quoting a specific reviewer line, so I can’t say with certainty it wasn’t partly our own reflection — we also flagged that zero of our eleven items were reverse-coded, which leaves a scale more exposed to acquiescence bias: a test-taker who just agrees with everything scoring as a catastrophizer by default.
5. Decide what to change and what to defend — item by item, in writing. We didn’t fold on everything. On the double-barreled critique, we disagreed and said why: items 1 and 2 were aimed at how a benign social situation produces discomfort, while item 3 targeted the social-anxiety layer of the same construct specifically — different targets, not one question doing two jobs. On a separate note that “I am often” or “I often” opened nearly every item, we kept the repetition on purpose: a consistent stem reduces the chance that a test-taker’s answer shifts because of how a sentence is phrased rather than what it’s asking.
On the past-tense critique, we pushed back harder, and this is where the Step 1 definition did its actual job:
People who catastrophize have prior experiences that affected them enough to cause them to potentially imagine worst case scenarios as well as take that fear into how they view the present; we argue there are experiences from the past that can carry through to the present even without our conscious knowledge. It is thus global and stacks on itself and is a pattern of behavior and these perceptions will influence the future.
Pattern, not a single forward-facing prediction — the word we’d chosen four steps earlier, now doing load-bearing work. A reviewer can talk you out of a word choice. They shouldn’t be able to talk you out of your construct unless they’ve actually found a hole in it, and “we think this only points forward” is a competing definition, not a hole.
I want to flag the less flattering reading of “held the line,” because from the outside it’s indistinguishable from the flattering one: under a semester deadline, defending an item and simply not wanting to rewrite it look identical, and occasionally they feel identical from the inside too. The only tell I trust is whether we wrote the reason down before we’d counted the cost of the alternative. On the double-barreled items we did, and I’ll stand behind those. On one or two of the others, I’m honestly not certain the conviction came first.
On the vagueness critique, we agreed outright and revised: “things that have occurred so far” became “unresolved inter-personal conflicts that have occurred so far.” Same sentence, one word replaced, disagreement on the construct held, defect on the wording fixed.
And on the missing reverse-coded item, we conceded the design gap and picked a candidate to flip. Original item 6 — “I often worry that bad things will happen to people I care about” — became:
I never worry that bad things will happen to people I care about.
One reverse-worded item, out of eleven. A targeted fix for the specific gap that was named, not a full redesign in response to a general complaint.
6. Turn around and review someone else’s measure. The same step that got our scale critiqued also assigned us to critique another group’s — an Emotional Intelligence measure for managers. The review form’s last substantive question asked us to read the items specifically for cultural bias, and our answer is the part of this whole project closest to the inclusion work I care about outside the classroom:
We did have a cross-cultural question about items 19 and 20, as in some geographies it may not be considered a manager's job to manage the emotions of their employees... Item 15 and 16 may be biased to majority culture, as minorities may not always be experts at the social norms which cause interpersonal conflict.
Writing that critique made the double-barreled and vagueness feedback on our own scale land differently in hindsight. It’s easy to read someone else’s item and immediately see whose default experience it was quietly written against. It’s much harder to do that to your own eleven items, written by your own hand, under your own deadline. We were more generous with ourselves than we were with Group 4, and I don’t think that asymmetry is unique to us.
7. Deploy to Qualtrics — and hit a mishap. We ran the finished eleven-item CS alongside a comparison instrument, the Cognitive Distortions Questionnaire (CD-Quest), to test both reliability and convergent validity in one pass. The CD-Quest’s matrix-style response format translated imperfectly onto the Qualtrics platform, and we had to run the whole measure twice with the class to get usable data.
The wrong turn here: we treated survey-platform formatting as a packaging detail to solve after the content was finalized, instead of as part of the instrument itself. A matrix-style item that renders correctly on paper and breaks in Qualtrics is still a broken item, from the test-taker’s side of the screen.
8. Analyze what came back, and report all of it. Using the odd/even split-half method, the unequal-length Spearman-Brown coefficient came out to r = .96. Cronbach’s alpha for the full eleven-item scale was α = .91 — a strongly internally consistent instrument by any conventional threshold. The number sitting underneath both coefficients: N = 7. Seven complete responses, after discarding what the platform mishap had corrupted.
Convergent validity told a quieter story. Eleven total responses came in for the comparison; six were usable after discarding incomplete ones. Recoded and correlated: r(4) = .28, p = .20, one-tailed — weak, non-significant convergence against the comparison instrument overall. Against the single CD-Quest item that most directly targets fortune-telling and catastrophizing specifically, the correlation was moderate: r(4) = .43, p = .20. Neither clears a conventional significance bar. We reported both at the same volume as the flattering alpha, in the same section, not in a footnote.
A .91 next to an N of 7 is the same shape of problem as a 99.9% pass rate next to a test suite that only covers the happy path. The coefficient isn’t lying. It’s answering a much narrower question than the headline number implies, and the narrowness has to travel with it every time it gets repeated.
And “narrower question” is generous. At N = 7 the confidence interval around an alpha that high is wide enough that the two decimal places are mostly decoration — hand me α = .91 on seven people and I wouldn’t call it narrow, I’d call it a number that shouldn’t appear at that precision outside a limitations sentence. The honest one-line version: the scale held together internally on a sample too small to trust the decimals.
9. Write the individual final paper — and a second letter. The group project ended at Step 7; the final write-up was individual, each of us taking the shared instrument and shared data and building our own analysis and argument around it. Mine opened with its own letter, responding point by point to feedback on the draft paper:
One item of reverse scoring has been added to the Catastrophizing Scale as per suggestion. The introduction is now well curated and structured based on your feedback about consistency. Although the IMRaD worksheet was group work, a level of personalization in refinement and analysis was provided to account for differences in N for reliability and validity testing. We tried getting the most from the initial mishap that occurred.
(Lightly copyedited for grammar from the submitted draft — the substance and every number are unchanged.)
The first line is a confirmation, not a new decision: the reverse-coded item from Step 5 made it all the way from the group draft into the individual paper’s methods section, not just into a revision nobody followed up on. The last line I didn’t soften. “The initial mishap that occurred” is the phrase I’d use in an internal retro at work, not a euphemism I saved for a grade — and it’s also, structurally, the same document I write for a living. A letter to reviewers and a test report say the same four things: what I checked, what passed, what didn’t, and what I did about each one. I’ve been writing that document professionally for years, under the title QA lead. This course just handed me a construct to point it at instead of a build.
One bias I should name before I let that metaphor stand: this is a construction log with a single narrator, and the narrator is the one whose profession the whole framing is borrowed from. The “we” throughout was real — the three item formats we merged into one scale were my two teammates’ as much as mine, and the parts that worked were not me arriving to impose QA discipline on a group that lacked it. A one-builder account of a three-person build is exactly the kind of single-source document I’d read skeptically at work, and this is one.
The honest audit
What transfers
A construct definition is a spec. An item is a test case written against that spec. A reviewer’s numbered critique is a code review comment, and the discipline isn’t agreeing with all of it or none of it — it’s writing down, before you respond, which comments changed your mind and which didn’t, and why. Reviewing someone else’s instrument the same week is the part I didn’t expect to matter most: it’s much easier to spot a construct-irrelevant bias in a stranger’s eleven items than in your own, and the gap between those two readings is worth noticing on purpose, not just once, by accident, in a classroom exercise.
None of this makes the project a failure. A strongly reliable, weakly valid, badly underpowered first-pass instrument is what a scale built from nothing in one semester should look like before anyone scales it up. The failure mode would have been reporting only the number that flattered it — or reading our own items more kindly than we read someone else’s, and never noticing we’d done it.
The real debt was never going to come from hitting a platform mishap under schedule pressure and reporting it honestly — that’s just what building something new looks like. It comes if this instrument gets inherited downstream by someone who sees α = .91, copies it into a literature review or a pitch deck, and never checks what seven people can and can’t tell you. The same pattern shows up anywhere a script gets pulled out of the sandbox it was validated in and treated as load-bearing.
Thanks for reading. If you’ve ever built something from a fuzzy idea to a testable spec — a KPI, a rubric, a support-ticket severity scale — I’d like to know two things: which number currently sitting in your own dashboard, deck, or report did you inherit from someone else’s result without ever checking its N, and when’s the last time you read your own work as critically as you’d read a stranger’s?
Comments