Benchmarks Discipline Industries: Governance Design for a Person-Perception Standard
Abstract
Whoever builds the first credible measurement of a capability shapes how that capability develops for a decade. This article asks what makes such a measurement credible when the party convening it stands to benefit from particular results — which is the position this program is in, and which is more common in benchmark creation than the field’s conventions acknowledge.
The answer proposed here is structural rather than reputational. Four commitments do the work: the convener does not own the scoring methodology; the analysis is frozen and published before any system is evaluated; results are reported per component and per instrument so that any single element can be discarded by a skeptical reader; and no system affiliated with the convener is evaluated by the standard during its establishing phase. Each commitment removes a specific mechanism by which interested parties influence results, and each is checkable from outside.
The article also draws on what benchmark scholarship has established about how measurements fail: they acquire authority their construct validity does not support, they redirect research toward what is measured, and they degrade as contamination accumulates. A person-perception standard carries an additional hazard the technical literature does not face — its subject matter is people, so a defective standard licenses claims about human beings rather than about model capability. That asymmetry argues for governance stricter than the norm, not looser.
Finally, the article draws on the one field that has governed high-stakes judgment of persons for a century: employment testing, whose validity standards remain contested and self-correcting to this day. The relevant lesson is not that such standards can be perfected but that they can be maintained under adversarial conditions, which is the realistic aspiration here.
Keywords: benchmarks, governance, conflict of interest, validity standards, measurement policy, artificial intelligence
1. What Benchmarks Do
Benchmarks are not neutral instruments that observe a field. They organize it.
A public standard makes capability claims falsifiable, makes progress comparable across vendors, and — most consequentially — directs development effort toward whatever it measures. Broad evaluation efforts have demonstrated both the value of systematic measurement and the sensitivity of conclusions to which capabilities are included and how they are reported (Liang et al., 2022). Individual benchmarks have reorganized entire research programs around specific test formats (Hendrycks et al., 2021).
That power is precisely why benchmark construction requires scrutiny. Critical scholarship has documented how evaluations acquire authority their construct validity cannot support — how a measure built for a bounded purpose comes to stand for a general capability, and how the resulting mismatch shapes what the field builds (Raji et al., 2021). The failure is not that benchmarks are wrong on their own terms. It is that they are read as measuring more than they measure, and the field then optimizes for the reading.
For a standard measuring how well AI systems judge people, this failure mode carries stakes the technical benchmarks do not. A mismeasured coding benchmark produces overstated claims about code. A mismeasured person-perception standard produces licensed claims about human beings — that a system knows someone well enough to advise them, evaluate them, or act on their behalf. The object of measurement is a person, and the consequence of error lands on them.
2. The Problem This Program Has
The convener of this benchmark has a commercial interest in psychometrically grounded personalization. A standard establishing that machine person-perception is unvalidated, poorly calibrated, and prone to fabrication would tend to favor approaches built on validated instruments — including the convener’s.
Stating this plainly is necessary but insufficient. Disclosure does not neutralize interest; it only makes it visible. What follows are the structural commitments that make the interest inoperable, each targeting a specific mechanism of influence.
It is worth noting that this situation is not unusual and that pretending otherwise would misdescribe how measurement standards typically originate. Standards are frequently convened by parties with stakes — industry consortia, vendors, and interested institutions — and the credible ones survive not because their conveners were disinterested but because their methodologies were held by parties who were.
3. Four Commitments
Commitment 1: The convener does not own the scoring methodology.
Scientific leadership sits with academic collaborators at multiple institutions. The scoring framework, criterion definitions, and analysis plan are theirs to set and to change. The convener’s role is to originate the framework, fund the establishing phase, and then not control what it concludes.
The mechanism this blocks is the most direct one: choosing the metric that produces the desired result. If scoring is held externally, the convener’s preferences operate only through argument, in public, where they can be rejected.
Commitment 2: Analysis is frozen and published before evaluation.
Everything that stands between raw system output and a headline number — criterion instruments, exclusion thresholds, prompts, decoding parameters, parsing rules, aggregation choices, primary outcomes, and the failure-mode predictions — is specified in a timestamped public pre-registration before any system is evaluated. Deviations are reported as deviations, with rationale.
The mechanism this blocks is subtler and more common: not fraud, but the ordinary exercise of judgment in the presence of results. The literature on researcher degrees of freedom establishes that undisclosed flexibility can produce a wide range of findings from a single dataset without any individual choice appearing unreasonable (Simmons, Nelson, & Simonsohn, 2011), and pre-registration is the established remedy (Nosek, Ebersole, DeHaven, & Mellor, 2018). A benchmark whose outcomes are correlations has an unusually large number of such choices available.
Commitment 3: Results are reported per component and per instrument.
No composite. Accuracy, calibration, fabrication, and sycophancy-linked distortion are reported as separate dimensions, and results are reported for each criterion instrument separately in addition to any consensus analysis.
Two mechanisms are blocked. Compositing hides which failure is present and performs an invisible weighting on behalf of every downstream user — the information loss that broad evaluation work has documented when heterogeneous capabilities are collapsed into single figures (Liang et al., 2022). Per-instrument reporting lets any skeptical reader discard any instrument — including any with commercial affiliation — and check whether conclusions survive on the remainder. This converts a potential liability into a robustness test that the reader performs rather than trusts.
Commitment 4: No system affiliated with the convener is evaluated during the establishing phase.
The standard’s first published results contain no entry the convener benefits from directly. Once methodology is externally held and the protocol is public, later evaluation of any system — including affiliated ones — is possible, but by then the scoring is not the convener’s to shape.
The mechanism this blocks is the crudest and most damaging: a standard whose author’s product happens to score well on it.
4. Failure Modes of Benchmarks, and What This One Requires
Four failure modes recur across evaluation efforts, and each has a specific countermeasure here.
Construct drift. A benchmark built for a bounded question becomes read as measuring a general capability (Raji et al., 2021). Countermeasure: every published result states its scope explicitly — measurement validity for trait judgments under specified conditions, not general understanding of persons, and not fitness for any deployment.
Optimization to the metric. Once a measure matters, systems are tuned to it, and the tuning may not generalize. Countermeasure: rotating the held-out portion of the criterion battery across releases, version-stamping every result, and treating an implausibly high score — one exceeding the criterion instruments’ own agreement — as a signal to investigate rather than a record to publish.
Contamination. Test material leaks into training or into the stimulus. Here contamination has a domain-specific form: participants who have discussed their own assessment results place the criterion inside the stimulus. Countermeasure: the published scrubbing protocol and audited scrub rates specified elsewhere in this program.
Staleness. A standard measures the systems of its founding moment. Countermeasure: version-stamped releases and a protocol designed for repeatable administration against future generations — the property that distinguishes a standing benchmark from a single study.
To these, one hazard specific to this benchmark: misuse as a certification. A system scoring well demonstrates measurement validity for trait judgments. It has not been shown safe, appropriate, or consented-to for any particular use. Every published result should carry that limit in its own words, because the commercial temptation to read a validity score as a deployment license will be immediate and considerable.
5. The Century-Old Precedent
There is one field that has governed high-stakes judgment of persons under adversarial conditions for a hundred years, and its record is instructive precisely because it is imperfect.
Employment testing developed validity standards, professional bodies, and a large technical literature to govern selection decisions with substantial consequences for individuals. It is also a field in active methodological dispute: recent work revisiting the field’s meta-analytic estimates argues that decades of accepted validity figures were systematically overstated through corrections for range restriction, prompting substantial reassessment of what selection methods actually predict (Sackett, Zhang, Berry, & Lievens, 2022).
Two lessons follow, and the second matters more.
The first is that standards for judging persons are achievable. A century of accumulated methodology exists, and it disciplines practice.
The second is that they require permanent maintenance. The field’s own core numbers were wrong for a long time and were corrected from within, by researchers examining their predecessors’ assumptions rather than defending them. A person-perception standard should expect the same trajectory — early estimates revised, methodological assumptions challenged, headline figures moved — and should be designed to make that correction easy rather than embarrassing. Publishing full protocols, per-instrument results, and machine-readable summary data is what makes reanalysis by critics possible, and reanalysis by critics is the mechanism by which such corrections actually occur.
6. What Would Falsify the Governance Claim
Governance commitments are testable, and stating the tests is part of making them binding.
The claim that this standard is independent of its convener’s interest fails if: scoring methodology remains under the convener’s control at first publication; the pre-registration is absent, late, or materially deviated from without disclosure; results are published as composites rather than per component and per instrument; an affiliated system appears in the establishing results; or funding sources go undisclosed in any publication.
Each is externally checkable. None requires trusting anyone’s intentions, which is the point.
7. Conclusion
The measurement described in this program will exist. Person-perception claims are proliferating, regulators will require some basis for evaluating them, and someone will build the first credible standard.
If that standard is a vendor’s self-assessment, the field will spend a decade arguing with it, and the arguing will be justified. The alternative is a standard whose methodology is held by parties with nothing to gain from the results, frozen before evaluation, reported in components no one can composite, and published in enough detail that its critics can attempt to overturn it.
That arrangement is not offered as evidence of good intentions. It is offered because good intentions are not evidence of anything, and because a standard measuring how well machines judge people should be held to the standard of proof it intends to impose.
References
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., & Steinhardt, J. (2021). Measuring massive multitask language understanding. In Proceedings of the International Conference on Learning Representations (ICLR).
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al. (2022). Holistic evaluation of language models. arXiv:2211.09110.
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606.
Raji, I. D., Denton, E., Bender, E. M., Hanna, A., & Paullada, A. (2021). AI and the everything in the whole wide world benchmark. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks.
Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366.
Citation compliance note: all 6 references above were verified against live sources on August 16, 2026, in the session that produced this draft series. No reference is cited from memory.
RELATED
- VIDEO
Scrubbing the Answer Key
A spoken walkthrough of contamination control and pre-registration discipline: four classes of leakage, auditable removal rules, and how findings stay confirmatory.
- VIDEO
The Fluency Demand Taxonomy: Scoring Work Tasks FD0–FD4
A spoken walkthrough of the second axis of occupational exposure: five ordinal levels scoring how much a task’s outcome depends on communication fitted to a specific person.
- VIDEO
Calibration as Co-Equal Criterion
A spoken walkthrough of why confidence–accuracy coupling belongs beside accuracy as a headline criterion, never composited into a single number.