A Century of Knowing Persons: What Accuracy Research Already Solved — and AI Evaluation Ignores
A Century of Knowing Persons: What Accuracy Research Already Solved — and AI Evaluation Ignores
Abstract
Machine learning is rediscovering, in slow motion and without attribution, a set of problems personality psychology solved decades ago.
Can a system judge a person accurately from limited evidence? How much evidence is enough? Which traits are easy to read and which are nearly invisible? Whose report counts as truth? Does knowing someone longer mean knowing them better? Every one of these has a mature theoretical answer, an empirical literature, and in several cases a formal model with estimated parameters.
This article maps that inheritance onto the machine case through four bodies of work: the Realistic Accuracy Model's decomposition of how judgment becomes accurate; Kenny's componential account of what an interpersonal judgment is made of; the self–other knowledge asymmetry, which specifies whose report is authoritative for which traits; and the acquaintanceship and thin-slice literatures, which establish the shape of the information–accuracy relationship. Each mapping yields testable hypotheses about machine judges — several counterintuitive, two predicting that machines should be systematically better than humans in specifiable ways and systematically worse in others. One implies that conversational AI occupies a position in the informant taxonomy that no human occupies.
The claim is not that machines are people. It is that a century of theory about judging persons was built at a level of abstraction that never specified what kind of thing the judge had to be — and that ignoring it means paying twice for knowledge already purchased.
Keywords: person perception, accuracy, Realistic Accuracy Model, acquaintanceship, trait visibility, large language models
1. The Inheritance Nobody Claimed
A researcher today asks whether a language model can accurately judge a person's personality from their text. She runs the study, computes a correlation, plots it, and publishes.
She has just asked a question with a hundred-year pedigree, and answered it with no theory at all.
This is not a criticism of the empirical work, which is often excellent. Machine judges have been scored against self-report at scale, with results that would have astonished the field's founders: models rating personality from digital behavioral residue can approach or exceed the accuracy of human informants who know the target personally (Kosinski, Stillwell, & Graepel, 2013; Youyou, Kosinski, & Stillwell, 2015), language-based assessment converges with self-report across enormous samples (Schwartz et al., 2013), and contemporary large language models infer psychological dispositions from user text with no training for the task (Peters & Matz, 2024). These are real findings about a real capability.
But findings without a theory accumulate; they do not compound. They cannot say why a model reads one trait better than another. Or where the accuracy curve should bend as evidence piles up. Or whose report counts as the criterion when self and informants disagree. Or what would show that apparent accuracy is an artifact of something else.
Personality psychology has answers to all four. They were developed studying human judges, and written at a level of abstraction that never specified what kind of thing the judge had to be.
That abstraction is the inheritance. This article claims it on behalf of machine person-perception, taking four literatures in sequence and deriving, from each, hypotheses that a validation program can test.
2. The Process Model: What Accuracy Requires
The Realistic Accuracy Model (RAM) begins from a position that machine learning shares without usually stating it: personality traits are real attributes of persons, and judgments about them can therefore be more or less correct (Funder, 1995). Accuracy is then decomposed into a sequence of conditions, all of which must be met: relevant behavioral information must be relevant to the trait, available to the judge, detected by the judge, and correctly utilized (Funder, 1995, 2012). The model's power lies in the sequence's fragility — failure at any stage caps accuracy regardless of performance at the others — and in the fact that the four classic moderators of accuracy (good judge, good target, good trait, good information) each map onto specific stages.
The mapping to machine judges is direct and immediately generative.
Relevance and availability become properties of the corpus. A system's evidence about a person consists of whatever the interaction happened to contain; a user who discusses only technical matters supplies abundant tokens with low relevance to, say, emotional stability. This yields the first hypothesis: machine accuracy should be predicted better by the trait-relevant content of an interaction history than by its raw volume. Two users with identical token counts and different disclosure profiles should be judged with materially different accuracy — a prediction that also implies a confound to control in any dose–response design.
Detection is where the two kinds of judge most plausibly diverge, and where RAM predicts machine superiority. Human detection is bounded by attention, by memory, and by the plain fact that nobody can re-read a relationship. Machine detection is bounded by none of these. The entire record is available at once, and subtle statistical regularities across hundreds of exchanges — base rates, lexical patterns, consistency across contexts — are exactly what these systems extract well. The relevant human literature is instructive here: judgments from very brief behavioral samples achieve non-trivial accuracy (Ambady & Rosenthal, 1992), which establishes that diagnostic cues are dense in behavior; a judge that can attend to all of them at once should, on the detection stage alone, do better than one that cannot. Hypothesis: machine judges will outperform human judges disproportionately on traits whose cues are distributed and statistical rather than salient and episodic.
Utilization is where RAM predicts machine inferiority, or at least machine idiosyncrasy. Correct utilization requires the judge to weight detected cues according to their actual validity. Human judges err here through stereotype and implicit personality theory; machine judges have their own version, inherited from training distributions and from optimization pressures that have nothing to do with accuracy — most notably the documented tendency of assistant-tuned systems to shape outputs toward user agreement rather than truth (Perez et al., 2022; Sharma et al., 2023). A system may detect a cue perfectly and utilize it toward flattery. Hypothesis: machine accuracy will show a characteristic utilization signature — good detection paired with systematically compressed or socially desirable weighting — detectable as intact rank-order sensitivity on some traits alongside inflated means on evaluative ones.
RAM thus supplies not just a vocabulary but a diagnostic: when a machine judge fails, the model tells you where to look, and the four stages generate distinguishable error signatures.
3. The Componential Model: What a Judgment Is Made Of
If RAM describes the process, Kenny's PERSON model describes the composition. It decomposes any interpersonal judgment into components — including the target's actual personality, the stereotype the perceiver applies, the perceiver's idiosyncratic opinion, shared norms, and error — and, from parameter estimates, derives implications about how judgment behaves in the aggregate (Kenny, 2004). Several of its conclusions are directly relevant here, and one is deeply counterintuitive: the model accounts for the finding that consensus among judges does not increase with greater acquaintance nearly as much as intuition expects, and that acquaintance is generally less important in person perception than commonly assumed (Kenny, 2004).
Three implications for machine evaluation follow.
First, stereotype and norm components will masquerade as accuracy unless the design controls for them. A machine judge that assigns every user the population-modal profile will achieve a non-trivial correlation with truth, because population-modal profiles are correct on average. This is the machine form of what the componential literature calls stereotype accuracy, and it is why any credible evaluation must separate the ability to rank persons against one another from the ability to reproduce the average person. The distinction is not merely technical: an assistant that treats every user as the population mean is precisely the "personalization" failure the industry is at risk of shipping while reporting positive correlations.
Second, judgments are decomposable into shared and unique variance, and both are interpretable for machines. Multiple stochastic runs of the same model on the same corpus estimate something analogous to a perceiver's reliability; different models judging the same person estimate something analogous to consensus. Both are reportable psychometric properties of a machine judge, and neither is currently reported by anyone.
Third, Kenny's counterintuitive result about acquaintance sets up the sharpest empirical question in this territory. If human consensus rises far less with acquaintance than folk psychology predicts — because the added information is partly redundant with what stereotype and early impressions already supply — then the machine analog is an open and highly consequential question: does machine accuracy continue to climb with exposure past the point where human accuracy plateaus? A system with perfect recall of a million tokens is not subject to the memory and attention constraints that plausibly flatten the human curve. If machine curves keep rising where human curves flatten, that is the strongest single argument that machine person-perception is a distinct capability worth measuring on its own terms rather than a mechanical approximation of the human one.
4. The Asymmetry Model: Whose Report Is Truth?
Any evaluation of a judge requires a criterion, and the criterion question in person perception has a well-developed answer that machine researchers routinely flatten. The Self–Other Knowledge Asymmetry model establishes that neither the self nor observers are globally superior: the self has privileged access to internal states and is the better judge of traits defined by thoughts and feelings (such as neuroticism-related traits), while observers have better access to overt behavior and can be superior judges of traits that are highly observable or highly evaluative (Vazire, 2010). Meta-analytic integration of observer ratings supports the broader point that other-reports carry genuine, incremental validity — and that intimacy and interaction frequency shape how much (Connelly & Ones, 2010).
The implications for machine evaluation are pointed.
The criterion should be trait-specific, and machine researchers should stop treating "self-report" as a monolithic gold standard. A machine judge scored only against self-report is being scored against a criterion of varying quality across the trait space — held to a demanding standard on traits the self knows well, and to a partially invalid one on traits the self knows poorly.
The machine occupies a position on the self–other axis that no human occupies. This is the most interesting theoretical point in this article. A conversational AI is not an observer in the classic sense — it never watches behavior. It is obviously not the self. It sits in a third position: recipient of a person's self-narration over time, including things the person may never have said aloud to anyone. Its evidence is verbal, longitudinal, and in some respects more intimate than that of most human informants, while lacking behavioral observation entirely. SOKA predicts a specific and testable profile from this position: machine judges should be relatively strong on internally defined traits ordinarily accessible only to the self — where the person's own disclosures are the evidence — and relatively weak on the highly observable, behavior-dependent traits at which human observers excel. If confirmed, this would establish machine person-perception not as a substitute for either perspective but as a third informant class with its own validity profile — a finding of genuine theoretical interest to personality science, independent of any application.
Evaluative traits carry a known distortion, and machines have their own version of it. Where SOKA identifies motivated distortion in self-reports of evaluative traits, machine judges face an analogous pressure from a different source: training that rewards agreeable outputs. The prediction is convergent — inflated ratings on socially desirable dimensions — but the mechanism is distinct, and separating the two requires exactly the multi-instrument, per-trait reporting that the criterion literature already recommends.
5. The Information Model: The Shape of the Curve
The final inheritance concerns the relationship between how much a judge has seen and how accurately they judge — the question that any dose–response evaluation of machine person-perception must inherit rather than reinvent.
Three findings anchor the human literature. Accuracy is a function of information quantity and quality, with both contributing and quality often dominating (Letzring, Wells, & Funder, 2006). Judgment from very thin behavioral slices is already non-trivially accurate, and — crucially — longer observation did not reliably yield better prediction in the meta-analytic record (Ambady & Rosenthal, 1992). And even at zero acquaintance, trait inferences carry genuine validity from appearance and minimal behavioral cues (Borkenau & Liebler, 1992). Against these, the acquaintance-length literature finds that what improves over time is specific: self–other agreement and differential accuracy increase with acquaintance while stereotype-based accuracy decreases, meaning that longer relationships shift judgment from what people in general are like toward what this person is like (Biesanz, West, & Millevoi, 2007).
That last decomposition is the single most useful import for machine evaluation, and it reframes the dose question entirely. The naive hypothesis is that accuracy rises with tokens. The theoretically informed one is sharper and easier to kill: as exposure grows, machine judgment should change composition — stereotype accuracy falling, differential accuracy rising — even if total accuracy barely moves.
A system whose total accuracy rises without that shift is getting better at the average person. Not at the individual. Nothing in current machine-personality research measures this, and it is arguably the most diagnostic single test of whether a system is personalizing at all.
Two further predictions follow directly. From the thin-slice literature: a floor condition should show above-chance but modest machine accuracy, since real signal exists at minimal exposure (Borkenau & Liebler, 1992) — which makes the floor a fabrication detector, because a system reporting rich, confident profiles there is producing something other than inference (a tendency the generative-systems literature documents in general form; Ji et al., 2023). And from the quantity–quality distinction: matched token counts should yield different accuracies as a function of corpus content, which means any credible dose design must characterize what is in the corpus, not merely how much of it there is.
6. What the Machine Case Adds Back
An inheritance argument risks implying that machine person-perception is merely an application of existing theory. It is not, and three features of the machine case press on the human literature in return.
Exposure becomes continuously manipulable within subjects. In human research, acquaintance length is a between-relationships variable, confounded with selection: people who know each other longer differ systematically from people who do not. A machine judge's exposure can be truncated chronologically at arbitrary points within the same person's record, holding the target, the criterion, and the judge constant. This is a design affordance the human literature has never had, and it permits estimation of the information–accuracy function free of the between-relationship confounds that complicate its human counterpart (Biesanz et al., 2007; Letzring et al., 2006).
The judge becomes replicable. A human judge cannot be run twice on the same evidence without contamination. A machine judge can be run indefinitely, which turns judge reliability from an inferential problem into a measurable quantity, and which allows the componential decomposition of judgment (Kenny, 2004) to be estimated with a precision the human paradigm cannot reach.
A new informant class enters the taxonomy. As Section 4 argued, a system whose evidence is longitudinal self-narration occupies a position on the self–other axis that has no human occupant. Diaries and therapists are the nearest analogs, and neither is quite right — the diary does not judge, and the therapist observes. If the SOKA-derived predictions hold, personality psychology gains a new perspective to place in its models, with its own characteristic validity profile. That is a contribution running from the machine case back into the parent discipline, not merely a borrowing from it.
7. Conclusion: Stop Paying Twice
The empirical question of whether machines can judge persons accurately is being answered, at considerable expense, without the theory that would make the answers cumulative. This article has argued that the theory exists, that it was written at the right level of abstraction, and that its transfer generates a research program rather than a footnote: process-stage error signatures from RAM; stereotype-versus-differential decomposition and reliability estimation from the componential tradition; trait-specific criteria and a novel informant position from SOKA; and, from the acquaintanceship and thin-slice literatures, the correct form of the dose question — not does accuracy rise with exposure but does its composition shift from the average person to this person.
The hypotheses derived here are stated to be tested and several are stated to fail: machine judges may show no compositional shift with exposure, may prove weak precisely on the internally defined traits SOKA predicts they should own, may plateau exactly where human judges do. Any of those outcomes would be informative, and none is currently known, because the studies have not been run in a framework capable of interpreting them.
Personality psychology spent a century learning how to ask whether one mind has correctly read another. The field now faces judges that read minds at industrial scale — and the discipline best equipped to evaluate them has mostly watched.
The inheritance is unclaimed. It should not stay that way.
References
Ambady, N., & Rosenthal, R. (1992). Thin slices of expressive behavior as predictors of interpersonal consequences: A meta-analysis. Psychological Bulletin, 111(2), 256–274.
Biesanz, J. C., West, S. G., & Millevoi, A. (2007). What do you learn about someone over time? The relationship between length of acquaintance and consensus and self–other agreement in judgments of personality. Journal of Personality and Social Psychology, 92(1), 119–135.
Borkenau, P., & Liebler, A. (1992). Trait inferences: Sources of validity at zero acquaintance. Journal of Personality and Social Psychology, 62, 645–657.
Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105.
Connelly, B. S., & Ones, D. S. (2010). An other perspective on personality: Meta-analytic integration of observers' accuracy and predictive validity. Psychological Bulletin, 136(6), 1092–1122.
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302.
Funder, D. C. (1995). On the accuracy of personality judgment: A realistic approach. Psychological Review, 102(4), 652–670.
Funder, D. C. (2012). Accurate personality judgment. Current Directions in Psychological Science, 21(3), 177–182.
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248, 1–38.
Kenny, D. A. (2004). PERSON: A general model of interpersonal perception. Personality and Social Psychology Review, 8(3), 265–280.
Kosinski, M., Stillwell, D., & Graepel, T. (2013). Private traits and attributes are predictable from digital records of human behavior. Proceedings of the National Academy of Sciences, 110, 5802–5805.
Letzring, T. D., Wells, S. M., & Funder, D. C. (2006). Information quantity and quality affect the realistic accuracy of personality judgment. Journal of Personality and Social Psychology, 91(1), 111–123.
McCrae, R. R., & Costa, P. T., Jr. (1987). Validation of the five-factor model of personality across instruments and observers. Journal of Personality and Social Psychology, 52(1), 81–90.
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al. (2022). Discovering language model behaviors with model-written evaluations. arXiv:2212.09251.
Peters, H., & Matz, S. C. (2024). Large language models can infer psychological dispositions of social media users. PNAS Nexus, 3(6), pgae231.
Schwartz, H. A., Eichstaedt, J. C., Kern, M. L., Dziurzynski, L., Ramones, S. M., Agrawal, M., Shah, A., Kosinski, M., Stillwell, D., Seligman, M. E. P., & Ungar, L. H. (2013). Personality, gender, and age in the language of social media: The open-vocabulary approach. PLoS ONE, 8(9), e73791.
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., et al. (2023). Towards understanding sycophancy in language models. arXiv:2310.13548.
Thorndike, E. L. (1920). A constant error in psychological ratings. Journal of Applied Psychology, 4(1), 25–29.
Vazire, S. (2010). Who knows what about a person? The self–other knowledge asymmetry (SOKA) model. Journal of Personality and Social Psychology, 98(2), 281–300.
Youyou, W., Kosinski, M., & Stillwell, D. (2015). Computer-based personality judgments are more accurate than those made by humans. Proceedings of the National Academy of Sciences, 112(4), 1036–1040.
Citation compliance note: all 20 references above were verified against live sources on August 16, 2026, in the session that produced this draft series. No reference is cited from memory.
RELATED
- VIDEO
Scrubbing the Answer Key
A spoken walkthrough of contamination control and pre-registration discipline: four classes of leakage, auditable removal rules, and how findings stay confirmatory.
- VIDEO
The Fluency Demand Taxonomy: Scoring Work Tasks FD0–FD4
A spoken walkthrough of the second axis of occupational exposure: five ordinal levels scoring how much a task’s outcome depends on communication fitted to a specific person.
- VIDEO
Calibration as Co-Equal Criterion
A spoken walkthrough of why confidence–accuracy coupling belongs beside accuracy as a headline criterion, never composited into a single number.