The AI as Candidate Instrument: A Full Validity Battery for Machine Judgments of Persons
Abstract
Suppose you wanted to know whether an AI system actually understands the people it talks to. Not whether users like it. Whether it is right about them.
You would need a validity battery — the same apparatus psychology built for evaluating any new measure of persons, adapted for a judge that reads conversation instead of administering items. This article specifies one. It has five components, each answering a different question that the others cannot: does the system's rating of a trait agree with validated measurement of that trait (convergent validity); does it distinguish traits from one another or collapse them into a single impression (discriminant validity); does its accuracy grow lawfully with exposure to the person (dose–response sensitivity); does its confidence track its accuracy (calibration); and does it reproduce the shape of an individual's profile rather than merely ranking people on isolated traits (profile accuracy).
Two design commitments run through all five. Accuracy is scored against the empirical ceiling set by how well the criterion instruments agree with each other, not against a perfect criterion that does not exist. And accuracy and calibration are reported separately, never composited, because a moderately accurate system that knows its limits and a more accurate system that does not are different products for a deployment decision, and a single number hides which one you have.
The battery is specified in enough operational detail to be pre-registered, implemented, and — the point of the exercise — run against the author's own systems by someone else.
Keywords: construct validity, multitrait-multimethod, calibration, person perception, benchmark methodology, large language models
1. What This Article Is For
A companion argument holds that AI systems rendering trait judgments about their users are functionally psychometric instruments and should be validated as such. That argument is only as good as the protocol that follows from it. This article supplies the protocol.
The framing is deliberately unglamorous: treat the system as a candidate instrument and run it through the validation any new measure of persons must survive. Construct validation has a specification — the nomological network of relationships a measure must exhibit to warrant its interpretation (Cronbach & Meehl, 1955) — and a century of practice has turned that specification into procedures. Nothing in what follows is methodologically novel. Its novelty lies entirely in the target.
Two properties of the machine case make the battery both easier and stranger to run than its human counterpart. Easier, because the judge is replicable: a machine judge can be run repeatedly on identical evidence, which converts the reliability of the judge from an inferential problem into a directly measurable quantity. Stranger, because the judge's exposure to the target is continuously manipulable — an interaction history can be truncated chronologically at any point — which gives the battery a dimension no human validation has ever had.
The five components are presented in the order in which failure is most consequential. A system that fails convergent validity is not measuring. A system that passes convergence but fails discrimination is measuring one thing and reporting five. A system that passes both but shows no dose sensitivity is not learning the person. A system that passes all three but is uncalibrated cannot be safely deployed. A system that passes all four but cannot reproduce a profile is ranking people, not knowing them.
2. The Criterion Problem, Settled First
Every validity claim is a claim relative to a criterion, and criterion choices made carelessly determine results more thoroughly than any analysis that follows. Four commitments settle the criterion before any model output is observed.
Multiple instruments, independently developed. Ground truth comes from a battery of validated self-report measures drawn from distinct instrument traditions — for instance, a hierarchical Big Five inventory in the public domain (Soto & John, 2017) alongside an independently developed six-factor instrument (Lee & Ashton, 2004). Using more than one is not redundancy. It is what makes the ceiling estimable (Section 3) and what allows any single instrument's idiosyncrasies to be detected rather than inherited.
Convergence defines the criterion; divergence voids the cell. Where the instruments agree about an individual's standing on a trait, that consensus is the criterion. Where they disagree beyond a pre-specified threshold, the criterion for that person-trait cell is undefined, and the cell is excluded from scoring rather than charged against the system. A candidate instrument should never be penalized for the criterion's internal disagreement.
Per-instrument reporting, always. All results are reported for each instrument separately, in full, in addition to the consensus analysis. Any reader may then discard any instrument and check whether conclusions survive on the remainder. This is the structural answer to every concern about instrument selection, including concerns about instruments with commercial affiliations: a finding that replicates across independently developed instruments is a robustness demonstration, and a finding that does not is a finding about the instrument.
Non-ipsative instruments only. Forced-choice, intra-individual formats do not yield the normative trait scores that observer ratings can be correlated against, and their inclusion would silently change the meaning of every coefficient in the battery.
One further criterion commitment is trait-specific and rarely observed in machine work. Self-report is not uniformly authoritative: the self holds privileged access to internally defined traits, while observers hold advantages for highly observable and highly evaluative ones (Vazire, 2010), and observer reports carry genuine incremental validity of their own (Connelly & Ones, 2010). A battery that scores machine judgments against self-report alone is therefore applying a criterion of varying quality across the trait space. Where feasible, informant ratings should be collected alongside self-report, and the criterion's differing authority across traits should be stated as an interpretive constraint rather than assumed away.
3. Component One: Convergent Validity, Scored Against the Ceiling
The question. Does the system's rating of a trait agree with validated measurement of that trait?
The analysis. For each trait, model, and exposure level, compute the correlation across persons between the system's ratings and the criterion consensus scores. This is the battery's backbone statistic and its most straightforward.
The complication that makes it honest. A raw correlation invites an unanswerable question: agreement with what would count as perfect? Two established instruments measuring the same construct do not agree perfectly with each other (McCrae & Costa, 1987; Pace & Brannick, 2010). Nothing can agree with a criterion more strongly than the criterion agrees with itself.
The battery therefore reports convergent accuracy as a proportion of the empirically observed inter-instrument agreement within the study's own data. If the criterion instruments correlate at .70 in this sample and a system correlates at .55 with their consensus, the system has attained roughly 80 percent of instrument-grade accuracy for that trait. The ceiling is estimated from the same participants under the same conditions, not imported from the literature — which matters, because ceilings vary by trait, sample, and administration.
This framing does two jobs at once. It forecloses the inflated claim, since a system cannot report accuracy against a fictional perfect standard. And it forecloses the unfair penalty, since a system is not charged for variance no instrument can explain. It also makes the headline number interpretable to non-specialists without distorting it, which is a requirement for any measure intended to inform deployment decisions.
What must be reported alongside. Confidence intervals, and the sample size that produced them. Correlation estimates are unstable in small samples and do not settle within useful bounds until sample sizes reach the range that stability analyses identify (Schönbrodt & Perugini, 2013); a validity battery that reports point estimates from small samples is producing noise with decimal places.
4. Component Two: Discriminant Validity, or Is It One Judgment or Five?
The question. Does the system distinguish traits from one another, or does it emit a single global impression wearing five labels?
The analysis. The multitrait–multimethod matrix, applied with methods defined as {each criterion instrument, each model} (Campbell & Fiske, 1959). The classic requirements transfer without modification: monotrait–heteromethod values (system's extraversion vs. instrument extraversion) should exceed heterotrait–heteromethod values (system's extraversion vs. instrument conscientiousness), and the pattern of trait intercorrelations in the system's ratings should resemble the pattern in the criterion rather than being uniformly higher.
The failure mode this exists to catch. Halo — the collapse of distinct judgments onto a single evaluative dimension — was identified in human raters a century ago (Thorndike, 1920) and has an obvious machine analog with an additional mechanism behind it. A system trained toward inoffensiveness may rate nearly everyone as agreeable, conscientious, stable, and open, near population means. Such a system can post a respectable mean-level correspondence with the criterion while carrying almost no between-person information. Convergent validity alone will not catch this; the discriminant analysis will.
A machine-specific extension. Because a machine judge can be run repeatedly on identical evidence, the matrix can be augmented with a within-model reliability estimate: the agreement among multiple stochastic runs of the same model on the same corpus. This is the machine analog of a perceiver-reliability estimate in the componential tradition of interpersonal perception (Kenny, 2004), and it bounds everything else in the battery — a judge that does not agree with itself cannot agree with a criterion. Within-model reliability is a reportable psychometric property of a model, and at present no vendor reports it.
A second extension worth pre-registering. Compare each model's between-person variance in trait ratings against the criterion's between-person variance. Variance compression is halo's quantitative signature and is detectable even when correlations look acceptable.
5. Component Three: Dose–Response Sensitivity, or Is It Learning the Person?
The question. Does accuracy improve lawfully as the system is given more of the person, and where does it plateau?
The design. Chronological truncation. A participant's interaction history is cut at standardized token thresholds — the first ten thousand tokens, the first fifty thousand, and upward on a log scale — and every truncation of every participant's corpus is judged by every model. Truncation must be chronological rather than sampled across the record, because the point is to reproduce the way an impression forms along a relationship's timeline.
Why this is the strongest design in the battery. Each participant serves as their own control across exposure levels. In the human acquaintanceship literature, exposure is a between-relationships variable confounded with selection: people who have known each other longer differ systematically from people who have not (Biesanz, West, & Millevoi, 2007). Truncation removes that confound entirely. Within-person exposure effects are estimated with the target, the criterion, and the judge held constant — a design affordance the human literature has never had.
The hypothesis that matters is compositional, not aggregate. The naive prediction is that accuracy rises with tokens. The literature supports a sharper one. What improves with acquaintance in human judges is specific: differential accuracy and self–other agreement rise while stereotype-based accuracy falls, meaning the judgment shifts from what people in general are like toward what this person is like (Biesanz et al., 2007). The battery therefore decomposes accuracy at each dose into stereotype and differential components, and the pre-registered prediction is a compositional shift.
This distinction is not academic. A system whose total accuracy rises across doses without the compositional shift is getting better at the average person — which is precisely the personalization failure the industry is at risk of shipping while reporting improving correlations.
The floor condition is a fabrication detector. At near-zero exposure, genuine signal exists but is modest; trait inferences at zero acquaintance carry real but limited validity in human judges (Borkenau & Liebler, 1992). A valid machine judge at the floor should therefore be modestly accurate and visibly uncertain. A system that produces rich, confident, elaborate profiles from negligible evidence is doing something other than inference — the person-perception form of the fluent-output-without-evidentiary-support tendency documented across generative systems (Ji et al., 2023). Floor behavior should be reported prominently rather than buried as a baseline condition.
A confound that must be controlled. Accuracy depends on the quality of trait-relevant information, not merely its quantity (Letzring, Wells, & Funder, 2006). Two participants with identical token counts may supply radically different evidence. Corpora must therefore be characterized descriptively — topic composition, self-disclosure density — or the dose curve will silently confound exposure with content.
6. Component Four: Calibration, as a Co-Equal Headline
The question. When the system expresses confidence in a judgment, does that confidence predict accuracy?
The analysis. Elicit a confidence rating for every trait judgment; score the confidence–accuracy relationship using the standard machinery — proper scoring rules with their decomposition into calibration and resolution components (Brier, 1950), calibration curves, and overconfidence indices interpreted against the known structure of overconfidence phenomena (Moore & Healy, 2008). The extension of these methods to contemporary neural systems is well developed, and the standing finding is not reassuring: modern networks tend toward systematic overconfidence (Guo, Pleiss, Sun, & Weinberger, 2017).
Why it is co-equal rather than supplementary. Consider two systems. The first attains 60 percent of the ceiling and reliably signals when it is unsure. The second attains 65 percent and expresses uniform high confidence regardless of evidence. On a single composite score, the second wins. In deployment, the first is safer in every application where a wrong read carries cost, because its errors are flagged and the other system's are not.
The non-negotiable reporting rule: accuracy and calibration are never composited. A single number formed from both hides exactly the trade-off a deployment decision requires. The battery reports two headline dimensions and lets consumers of the results weigh them for their own use case. This is a stronger version of the general principle that benchmarks lose information when heterogeneous capabilities are collapsed into a scalar (Liang et al., 2022; Raji et al., 2021).
What calibration failure looks like specifically. Uniform confidence across doses is the primary signature — a system as certain at the floor as at a million tokens is not tracking its own evidence. Uniform confidence across traits is the secondary signature, since traits genuinely differ in how legible they are from verbal behavior, and a judge insensitive to that difference is not modeling its own uncertainty.
7. Component Five: Profile Accuracy, or Does It Know the Person?
The question. Does the system reproduce the shape of an individual's trait profile, not merely rank that individual correctly among others on isolated traits?
Why the first four components leave this open. Convergent validity is computed across persons within traits. A system could rank people correctly on every trait separately while never assembling a coherent picture of any individual — which is the wrong unit of analysis for deployment, where the system's output is a model of one person.
The analysis. Within-person profile correlations: the correlation, across trait dimensions, between the system's profile of an individual and that individual's criterion profile. The essential correction is normative. Because human trait profiles share substantial common shape, an uncorrected profile correlation is inflated by the population average, exactly as stereotype accuracy inflates single-trait agreement. Profile accuracy must therefore be reported both raw and distinctively — with the normative profile partialled — and the distinctive value is the one that answers the question. The parallel to the stereotype/differential decomposition in Component Three is exact, and both derive from the same insight in the accuracy literature: agreement about people in general is not knowledge of a person.
8. Pre-Registration and What It Must Contain
A validity battery is only as credible as the analytic freedom it forecloses. The methodological reform literature has established both the mechanism of the problem — undisclosed flexibility can produce apparently significant findings from almost any dataset (Simmons, Nelson, & Simonsohn, 2011) — and the remedy (Nosek, Ebersole, DeHaven, & Mellor, 2018). For this battery, pre-registration must fix, before any model output is observed:
the criterion instruments and the divergence threshold that voids a cell; the exact prompts, held identical across models, and all decoding parameters; the dose ladder and truncation rules; the number of stochastic repetitions per cell; the automated parsing procedure from model output to numeric score, with no human discretion at any point; exclusion rules for corpora and instrument protocols; the ceiling estimation procedure; and the three failure-mode predictions — halo compression, floor fabrication, and confidence–accuracy decoupling — stated as confirmatory tests rather than exploratory observations.
One contamination control is specific to this battery and belongs in the pre-registration rather than in the discussion section. When the stimulus is a person's own conversational history, that history may contain the person discussing their own assessment results or trait self-descriptions. Such content converts inference into retrieval: the system reads the answer key. Because fluent output does not announce its provenance (Ji et al., 2023), this contamination is invisible in results. It must be removed by published codebook, applied by raters blind to instrument scores, audited on a random subsample, with scrub rates reported. A companion article in this series develops the codebook in full.
9. What the Battery Cannot Do
Four limits are structural and should be stated wherever results are reported.
It measures perception, not benefit. Nothing in the battery establishes that accurate person-perception improves outcomes. That is a separate empirical question with its own framework, and treating a validity score as a benefit score would be an inferential leap the battery does not license.
It inherits the criterion's limits. Self-report instruments carry their own biases and vary in authority across the trait space (Vazire, 2010). The ceiling-referenced metric handles this honestly rather than pretending otherwise, but honest handling is not elimination.
It is version-stamped by necessity. Models change on provider timescales. Every result is a measurement of a specific version under specific conditions, and the protocol's value lies in being repeatable against future generations — the property that distinguishes a standing benchmark from a single study.
It does not license deployment. A system that scores well on this battery has demonstrated measurement validity for trait judgments. It has not been shown safe, appropriate, or consented-to for any particular use. Validity is a precondition for responsible deployment of person-perception claims, not a substitute for the governance those claims require.
10. Conclusion
The apparatus in this article is, in a strict sense, unoriginal. Convergent and discriminant validation is sixty-seven years old. The proper scoring rule is seventy-six. Construct validity is seventy-one. The insight that agreement about people in general is not knowledge of a person predates every system this battery would evaluate.
That is the argument. A field currently deploying millions of unvalidated instruments does not need new measurement theory. It needs to run the theory it inherited against the systems it built — and to publish what comes back, per instrument, per trait, per dose, with calibration beside accuracy and neither hidden inside a composite.
The battery is specified here in enough detail to be implemented by any laboratory, against any system, including the author's. That is not a gesture. It is the only condition under which the results of such a battery would mean anything at all.
References
Biesanz, J. C., West, S. G., & Millevoi, A. (2007). What do you learn about someone over time? The relationship between length of acquaintance and consensus and self–other agreement in judgments of personality. Journal of Personality and Social Psychology, 92(1), 119–135.
Borkenau, P., & Liebler, A. (1992). Trait inferences: Sources of validity at zero acquaintance. Journal of Personality and Social Psychology, 62, 645–657.
Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3.
Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105.
Connelly, B. S., & Ones, D. S. (2010). An other perspective on personality: Meta-analytic integration of observers' accuracy and predictive validity. Psychological Bulletin, 136(6), 1092–1122.
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302.
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (PMLR 70, pp. 1321–1330).
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248, 1–38.
Kenny, D. A. (2004). PERSON: A general model of interpersonal perception. Personality and Social Psychology Review, 8(3), 265–280.
Lee, K., & Ashton, M. C. (2004). Psychometric properties of the HEXACO Personality Inventory. Multivariate Behavioral Research, 39(2), 329–358.
Letzring, T. D., Wells, S. M., & Funder, D. C. (2006). Information quantity and quality affect the realistic accuracy of personality judgment. Journal of Personality and Social Psychology, 91(1), 111–123.
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al. (2022). Holistic evaluation of language models. arXiv:2211.09110.
McCrae, R. R., & Costa, P. T., Jr. (1987). Validation of the five-factor model of personality across instruments and observers. Journal of Personality and Social Psychology, 52(1), 81–90.
Moore, D. A., & Healy, P. J. (2008). The trouble with overconfidence. Psychological Review, 115(2), 502–517.
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606.
Pace, V. L., & Brannick, M. T. (2010). How similar are personality scales of the "same" construct? A meta-analytic investigation. Personality and Individual Differences, 49, 669–676.
Raji, I. D., Denton, E., Bender, E. M., Hanna, A., & Paullada, A. (2021). AI and the everything in the whole wide world benchmark. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks.
Schönbrodt, F. D., & Perugini, M. (2013). At what sample size do correlations stabilize? Journal of Research in Personality, 47(5), 609–612.
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366.
Soto, C. J., & John, O. P. (2017). The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of Personality and Social Psychology, 113(1), 117–143.
Thorndike, E. L. (1920). A constant error in psychological ratings. Journal of Applied Psychology, 4(1), 25–29.
Vazire, S. (2010). Who knows what about a person? The self–other knowledge asymmetry (SOKA) model. Journal of Personality and Social Psychology, 98(2), 281–300.
Citation compliance note: all 22 references above were verified against live sources on August 16, 2026, in the session that produced this draft series. No reference is cited from memory.
RELATED
- VIDEO
Scrubbing the Answer Key
A spoken walkthrough of contamination control and pre-registration discipline: four classes of leakage, auditable removal rules, and how findings stay confirmatory.
- VIDEO
The Fluency Demand Taxonomy: Scoring Work Tasks FD0–FD4
A spoken walkthrough of the second axis of occupational exposure: five ordinal levels scoring how much a task’s outcome depends on communication fitted to a specific person.
- VIDEO
Calibration as Co-Equal Criterion
A spoken walkthrough of why confidence–accuracy coupling belongs beside accuracy as a headline criterion, never composited into a single number.