Benchmarks and Standards
What an independent, pre-registered benchmark for machine person-perception would have to measure, and who is allowed to hold the answer key.
- VIDEO
Scrubbing the Answer Key
A spoken walkthrough of contamination control and pre-registration discipline: four classes of leakage, auditable removal rules, and how findings stay confirmatory.
- ARTICLE
A Century of Knowing Persons: What Accuracy Research Already Solved — and AI Evaluation Ignores
Machine learning is rediscovering questions personality psychology solved long ago. This article maps four established research traditions onto the machine-judge case and derives testable hypotheses about where AI should be accurate, biased, or simply wrong.
- VIDEO
The Unmeasured Instrument
A spoken walkthrough of the argument: AI person-perception is a psychometric problem, and no benchmark yet validates whether machine judgments of people are accurate.
- ARTICLE
The D0 Problem: Fabricated Familiarity at Zero Acquaintance
Why the floor of exposure belongs in every benchmark: specifying the D0 condition and the fabricated-familiarity rate as a detector for confident portraits built from nothing.
- ARTICLE
Scrubbing the Answer Key: Contamination Control and Pre-Registration Discipline in Person-Perception Evaluation
What it takes to keep a person-perception benchmark honest: four classes of contamination, auditable removal rules, and the pre-registration architecture that keeps findings confirmatory.
- ARTICLE
Where Fluency Pays: Communicative Load as the Moderator of Personalization Benefit
A pre-registered prediction and the decisive experiment: accuracy about a person converts into benefit only in proportion to a task’s communicative load.
- ARTICLE
Benchmarks Discipline Industries: Governance Design for a Person-Perception Standard
Four structural commitments that make a benchmark credible when its convener has a commercial stake in the results.
- ARTICLE
The Unmeasured Instrument: AI Person-Perception as Psychometrics' Missing Validation Problem
AI systems now judge personality at scale, yet no benchmark validates whether those judgments are accurate. This piece argues that person-perception is a psychometric problem — and sketches what an independent validation standard would require.