Population it was built for
Never stated
Accountable for the interpretation
Nobody
Language it was read in
Unaccounted for
There is a version of me who spends his working life asking where the evidence is, and another version who takes a four-letter personality quiz at eleven at night and thinks oh, that’s so accurate.
They are the same person. Only one of them is on duty.
Instrument assessed at work
"Show me the validity evidence for this population."
Same instrument, aimed at me
"Ooh — which one am I?"
Section 9 of the American Psychological Association’s ethics code covers assessment. Eleven standards, two and a half pages in total, written for people who administer tests that decide things — custody, diagnosis, employment. It is not written about internet quizzes, and I want to be clear that posting one is not an ethics violation. Nobody sees a colour-coded personality graphic on a feed and believes a clinician has been consulted.
What the eleven standards are good for is showing you the shape of a claim that has been properly earned — and therefore how little of that shape a quiz result has.
The shape of a claim that was actually earned
The eleven standards, one sentence each expand
9.01 An opinion needs evidence sufficient to support it.
9.02 Use an instrument only for the population and purpose its evidence covers.
9.03 The person knows the purpose, the limits, and who will see the result.
9.04 Raw responses move with consent, not by default.
9.05 Built with defensible psychometric procedure, not intuition.
9.06 Account for situational, linguistic and cultural factors — and state the limits.
9.07 Don't put assessment in the hands of people untrained to read it.
9.08 An old score is not a current fact about a person.
9.09 The human stays accountable for the interpretation, not the engine.
9.10 Someone qualified explains the outcome to the person it describes.
9.11 Items and protocols stay protected from open circulation.
My working paraphrase. The published text is the authority.
Read as professional housekeeping, they’re dull. Read as acceptance criteria, they are a specification for not doing harm with a number — and the interesting thing is how many of them a free quiz skips without anyone noticing, least of all me.
9.01
An opinion needs evidence sufficient to support it.
A generated paragraph, with no evidence attached.
9.02
Use it only for the population its evidence covers.
No population is named, so none is excluded.
9.03
State the purpose, the limits, and who sees the result.
A share button.
9.05
Built with defensible psychometric procedure.
Procedure unpublished. Possibly none.
9.06
Account for linguistic and cultural factors, and state the limits.
Written in one language. Taken in another. Silent about it.
9.10
Someone qualified explains the outcome to the person.
It appears on a screen at eleven at night.
The bit that isn’t harmless
I take most of my tests in a language I learned second, and 9.06 is the standard that keeps my attention. When an item’s wording is ambiguous, part of my score measures how I parsed the sentence — and that portion is returned to me as a fact about my personality.
I once spent long enough on the word assertive that I answered it three different ways in my head before choosing. There is no version of that where the resulting label is measuring my assertiveness cleanly. Some of it is measuring my English.
Reported to me as
A trait
Partly
The thing being measured
And partly
How I read the sentence
ITEM
One word: assertive.
CONSIDERED
Three different readings, before choosing.
RECORDED
One answer.
RETURNED
One trait, stated as a fact about me.
For years I read those labels as verdicts. Not because I believed the quiz was rigorous — I never thought that — but because a description that fits feels like being seen, and being seen is very hard to argue with at eleven at night.
The eleven standards did not make me feel guilty. They gave me a vocabulary for what was missing, which is a much more useful thing to have than guilt. A label with no stated population, no accountable interpreter and no account of the language it was written in is not a finding. It is a compliment with a chart attached.
What I do now
Four questions, and none of them stop me taking the quiz.
01
Who was this built for?
An instrument that can’t name a population wasn’t built for anyone, including me.
02
Is the language doing the measuring?
If an item is ambiguous to a non-native reader, part of the score is reading comprehension.
03
Who is accountable for the interpretation?
A generated paragraph has no author to disagree with.
04
What would this label cost me if I believed it?
That’s the only one with real stakes, and it’s the one I used to skip.
The professional standards exist because a number, handed to someone about themselves, does something. It sticks. Ask anyone who has ever said “I’m not really a leader type” and meant it, and then try to find where the belief came from.
The finding, as it is actually held
“I’m not really a leader type”
Date issued
Unknown
Interpreter
None on record
Status
Still in use
I still take the quizzes. I just no longer treat the output as information about me — and I’ve stopped assuming that the version of me who audits things for a living is the one who shows up when the subject under test is myself.
Comments