Scrubbing the Answer Key: Contamination Control and Pre-Registration Discipline in Person-Perception Evaluation
Abstract
A participant donates two years of conversations with an AI assistant. Somewhere in month seven, they paste in their personality test results and ask what the scores mean.
Every model that later reads that corpus is no longer inferring the participant’s traits. It is reading the answer key — and its output will look exactly like inference, because fluent generation does not disclose its sources.
This article specifies the contamination controls that a person-perception benchmark built on real conversational history requires, and the pre-registration architecture that keeps its findings confirmatory. Four contamination classes are defined: direct assessment disclosure, trait self-labeling, study-awareness content, and third-party characterization of the participant. For each, a detection procedure and a removal rule are given, along with the codebook discipline — blind raters, dual coding, audited scrub rates, published rules — that makes removal auditable rather than discretionary. Two costs are stated plainly: scrubbing removes genuinely diagnostic material along with contaminating material, and the quantity removed is itself a reportable finding rather than a footnote.
The second half specifies what must be frozen before any model output is observed, and why this benchmark is unusually vulnerable if it is not. When the outcome measure is a correlation, the analyst holds enough degrees of freedom — trait aggregation, exclusion thresholds, prompt selection, parsing rules — to produce a wide range of headline numbers from a single dataset without any single choice appearing unreasonable. Pre-registration is what converts that latitude into a fixed commitment, and a benchmark asking an industry to accept unwelcome results has no standing to skip it.
Keywords: contamination, pre-registration, researcher degrees of freedom, benchmark methodology, person perception, reproducibility
1. The Answer Key Problem
Every evaluation of judgment faces a version of this hazard: the judge must not have access to the criterion. In educational testing it is item exposure. In machine learning it is train–test contamination. In person-perception evaluation built on conversational corpora, it takes a form with no established name and an unusually quiet failure mode.
The stimulus is a person’s own conversation history. People discuss themselves in those conversations. Some of what they say about themselves is exactly what the benchmark is asking models to infer.
A participant who has ever pasted an assessment report into a chat, described themselves as “a classic introvert,” or asked an assistant to help interpret a personality profile has placed the criterion inside the stimulus. A model reading that corpus can retrieve rather than infer, and its output is indistinguishable from inference — a specific instance of the general property that generated text does not carry provenance markers with it (Ji et al., 2023).
The failure is quiet in a way that matters. Contamination in this setting produces higher accuracy, not obvious error. It looks like success. It will be distributed unevenly across participants, since only some people discuss their test results, which means it inflates some cells and not others and cannot be detected by inspecting aggregate results. And it interacts perversely with the dose design: because contaminating material can appear anywhere in a timeline, a truncation level that happens to include a disclosure will show an accuracy jump that mimics a genuine dose effect.
2. Four Classes of Contamination
Contamination is not one thing, and the removal rules differ by class.
Class 1 — Direct assessment disclosure. Verbatim or near-verbatim reproduction of instrument results: score reports, percentiles, type codes, profile summaries, or requests to interpret them. This is the most severe class because it supplies the criterion in the criterion’s own units. Rule: remove the entire exchange containing the disclosure, including the model’s response to it, since the response restates the content.
Class 2 — Trait self-labeling. The participant characterizes themselves in trait language without reference to an instrument: “I’m terrible with details,” “I’m the most impulsive person I know,” “I hate confrontation.” This class is harder, because self-labeling is ordinary conversation and also genuine behavioral evidence about the person.
The distinction that resolves it is between labels and behavior. “I hate confrontation” is a self-report of a trait — a claim in the same currency as the criterion. “I rewrote that email four times before sending it and then deleted it” is behavior from which a judge may infer. The first is removed; the second is retained. This is the boundary the codebook must define most precisely, because it will be applied thousands of times and because it determines what the benchmark is actually measuring: inference from behavior, not reading of self-description.
Class 3 — Study-awareness content. Discussion of the research, the assessment battery the participant completed for the study, or the fact that their corpus is being donated. Beyond direct contamination, this class introduces demand characteristics into any post-enrollment material. Rule: remove; additionally, flag corpora with substantial post-enrollment volume for sensitivity analysis, since awareness may alter conversational behavior in ways scrubbing cannot undo.
Class 4 — Third-party characterization. Others’ descriptions of the participant, quoted or reported by the participant: a performance review, a partner’s observation, a friend’s assessment. This is the subtlest class and the only one where the correct rule is genuinely arguable. Third-party characterization is observer-report — the same evidence class as the informant ratings that carry documented incremental validity in person perception (Connelly & Ones, 2010). It is not the criterion, but it is criterion-adjacent, and a model reading it is being handed another judge’s conclusions rather than forming its own.
The recommended rule is removal with reporting, on the ground that the benchmark’s construct is the system’s own perception rather than its ability to aggregate others’ verdicts. The recommendation is flagged as contestable, and the sensitivity analysis — results with and without Class 4 retained — should be pre-specified so the choice is auditable rather than decisive.
3. Codebook Discipline
Removal rules are only as trustworthy as the process applying them.
Blind application. Scrubbing raters must not have access to participants’ instrument scores. A rater who knows a participant scored high on a dimension cannot un-know it when deciding whether a passage is a self-label, and the direction of the resulting bias is unpredictable.
Dual coding with adjudication. A random subsample — twenty percent is a defensible floor — is coded independently by a second rater, with agreement reported and disagreements adjudicated by a third under written rules. Inter-rater agreement for a scrubbing codebook is a reportable psychometric property of the protocol, and a benchmark that does not report it is asking to be trusted rather than checked.
Published codebook. The rules, with worked examples of retained and removed passages, are published with the protocol. This is what makes independent replication possible and what allows a critic to argue with the boundary rather than merely suspect it.
Audited scrub rates. The proportion of each corpus removed, by class, is reported per participant and in aggregate. This is not administrative detail. Scrub rate is a variable that may correlate with participant characteristics — self-reflective people discuss their traits more — and unreported differential scrubbing is a confound hiding in plain sight.
Automated pre-screening, human adjudication. Keyword and pattern matching can surface candidate passages efficiently; the retain/remove decision should not be automated, because the Class 2 boundary requires judgment that pattern matching cannot supply.
4. What Scrubbing Costs
Two costs are real and must be stated wherever results are reported.
Diagnostic material is lost. The passages most likely to be scrubbed are those where participants discuss themselves — which is also where trait-relevant signal is densest. Accuracy depends on the quality of trait-relevant information and not merely its quantity (Letzring, Wells, & Funder, 2006), and scrubbing removes exactly the highest-quality material. Post-scrub accuracy is therefore a conservative estimate of what a system could achieve on unscrubbed deployment data. The direction of the bias is known, which is the best that can be said for it, and it should be stated in every abstract rather than discovered in a limitations section.
Token counts shift. Because scrubbing removes material, post-scrub corpora are smaller than donated corpora, and by differing amounts. Dose levels must therefore be defined on post-scrub token counts, and the pre-registration must say so, or the dose ladder silently varies in meaning across participants.
A third consideration is not a cost but a design requirement. Because scrubbing may remove material unevenly across a timeline, the removal must be applied before truncation, not after, so that every dose level is drawn from a uniformly cleaned corpus. Applying truncation first would produce dose levels with heterogeneous contamination exposure — the precise artifact that mimics a dose effect.
5. Degrees of Freedom in a Correlation-Based Benchmark
Contamination control protects the stimulus. Pre-registration protects the analysis, and this benchmark needs it more than most.
The mechanism is well established: undisclosed flexibility in data collection and analysis allows almost any dataset to yield apparently significant findings, without any individual choice appearing unreasonable to the analyst making it (Simmons, Nelson, & Simonsohn, 2011). The remedy — specifying decisions before the data can influence them — is equally established (Nosek, Ebersole, DeHaven, & Mellor, 2018).
What makes person-perception evaluation unusually exposed is the sheer number of defensible choices sitting between raw model output and a headline number:
Trait aggregation. Report per-trait, per-domain, or as an overall accuracy figure? Each is defensible; each yields a different headline.
Exclusion thresholds. How much criterion divergence voids a person-trait cell? A threshold chosen after seeing which cells are voided is a threshold chosen for its results.
Prompt selection. Small wording differences change model outputs. Selecting among piloted prompts after seeing performance is prompt-fitting to the outcome.
Parsing rules. Converting free-text output to numeric scores involves edge cases — refusals, ranges, hedged answers. Rules written while examining outputs are rules written toward a result.
Dose ladder definition. Breakpoints adjusted after inspecting curves can convert noise into an inflection.
Model version and decoding parameters. Temperature, sampling, and version choices all move results, and post hoc selection among them is a defensible-sounding path to a preferred finding.
None of these requires bad faith. That is the point of the degrees-of-freedom literature: they operate through ordinary judgment exercised in the presence of results.
6. The Freeze List
The following are fixed before any model output is observed, published as a timestamped pre-registration, and any subsequent deviation is reported as a deviation with its rationale:
Criterion. Instruments, administration order, scoring, the consensus rule, and the divergence threshold that voids a cell.
Stimulus. Scrubbing codebook, class definitions, blind-rating and dual-coding procedure, scrub-rate reporting format, and the Class 4 sensitivity analysis.
Exposure. Dose ladder in post-scrub tokens, chronological truncation rule, and the handling of corpora that fail minimum size after scrubbing.
Elicitation. Verbatim prompts, confidence elicitation format, response schema, decoding parameters, and repetitions per cell.
Parsing. The complete output-to-score procedure, including edge-case rules, with no human discretion at any point.
Analysis. Primary and secondary outcomes, ceiling estimation procedure, per-trait and per-instrument reporting format, mixed-model specification, and the multiple-comparison approach.
Predictions. The failure-mode hypotheses — halo compression, floor fabrication, confidence–evidence decoupling, and sycophantic preference-linked variance — stated as confirmatory tests, so that finding them is confirmation and not discovery.
Exclusions. Participant, corpus, and cell exclusion rules, with anticipated attrition stated in advance.
One clarification about the last item. Pre-registration constrains confirmatory claims; it does not prohibit exploration. Exploratory analyses are welcome and expected — they simply must be labeled as such, and a benchmark that reports every finding as confirmatory is misrepresenting its own evidentiary status regardless of how the analyses were run.
7. Why the Standard Must Be Higher Here
Three features of this benchmark raise the bar above ordinary good practice.
It asks an industry to accept unwelcome findings. A benchmark whose likely results include “systems fabricate people at the floor” and “personalization improves the average rather than the individual” will face motivated scrutiny of its methods. Every unregistered analytic choice is a place for that scrutiny to land, and correctly so.
Its convener has a commercial interest. The convener of a benchmark that could favor psychometrically grounded approaches has an interest in particular outcomes. The only credible answer is structural: freeze the analysis before seeing results, publish the freeze, report per instrument so any single measure can be discarded by a skeptical reader, and hold scoring methodology outside the interested party’s control. Assurances do not substitute for any of this.
It aspires to be a standard. Benchmarks acquire authority that their construct validity may not support, and the field has documented both the phenomenon and its consequences for research priorities (Raji et al., 2021; Liang et al., 2022). A benchmark intended to endure across model generations must be reproducible by parties who did not build it, which requires that everything above be public rather than described.
8. Conclusion
The two disciplines in this article are unglamorous and they are the reason anything else in this program can be believed.
Contamination control is not hygiene. It defines what the benchmark measures — inference from behavior rather than retrieval of self-description — and without it, the most reflective participants would silently produce the highest scores for the models that read them.
Pre-registration is not bureaucracy. In a benchmark whose outcomes are correlations, the analyst holds enough defensible choices to produce a wide range of headline numbers from one dataset. Freezing those choices is what makes the resulting number a finding rather than a selection.
Both are costly. Both remove flexibility that would sometimes have been used well. That is what they are for, and a program asking an industry to change its practices on the strength of its measurements cannot ask for latitude it would not grant the systems it evaluates.
References
Connelly, B. S., & Ones, D. S. (2010). An other perspective on personality: Meta-analytic integration of observers’ accuracy and predictive validity. Psychological Bulletin, 136(6), 1092–1122.
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248, 1–38.
Letzring, T. D., Wells, S. M., & Funder, D. C. (2006). Information quantity and quality affect the realistic accuracy of personality judgment. Journal of Personality and Social Psychology, 91(1), 111–123.
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al. (2022). Holistic evaluation of language models. arXiv:2211.09110.
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606.
Raji, I. D., Denton, E., Bender, E. M., Hanna, A., & Paullada, A. (2021). AI and the everything in the whole wide world benchmark. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks.
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366.
Citation compliance note: all 7 references above were verified against live sources on August 16, 2026, in the session that produced this draft series. No reference is cited from memory.
RELATED
- VIDEO
Scrubbing the Answer Key
A spoken walkthrough of contamination control and pre-registration discipline: four classes of leakage, auditable removal rules, and how findings stay confirmatory.
- VIDEO
The Fluency Demand Taxonomy: Scoring Work Tasks FD0–FD4
A spoken walkthrough of the second axis of occupational exposure: five ordinal levels scoring how much a task’s outcome depends on communication fitted to a specific person.
- VIDEO
Calibration as Co-Equal Criterion
A spoken walkthrough of why confidence–accuracy coupling belongs beside accuracy as a headline criterion, never composited into a single number.