← Back to blog
ARTICLE

Beyond the Stranger: Machine Person-Perception Versus Human Judges Across Acquaintance

James Niblick, PsyD9 min read

Abstract

Everyone asks the question. Does the machine know me better than my colleagues do? Better than my friends? Better than my spouse?

The benchmark program this article concludes has deliberately refused to answer it, for a methodological reason: comparing a machine judge to human judges requires holding the evidence constant, and evidence is exactly what differs between them. A spouse has decades of shared life; a model has a conversation history. Comparing their accuracy without addressing that asymmetry produces a number that means whatever the reader wants it to mean.

This article specifies how to ask the question properly. The design pairs the machine dose ladder — accuracy as a function of tokens of interaction history — against human judges at graded acquaintance levels, scored against the same multi-instrument criterion, in the same participants, on the same trait dimensions. The resulting comparison is not a single verdict but two curves, and the interesting quantities are their crossing points: how much conversational exposure buys a machine the accuracy of a colleague, of a close friend, of a spouse — and whether the machine curve continues climbing where the human curve flattens.

Four predictions are stated, one of which is counterintuitive and consequential: because a conversational system’s evidence is a person’s self-narration over time, it may exceed human observers precisely on the internally defined traits observers read worst, while remaining inferior on the behavioral traits observers read best. If confirmed, machine person-perception is not a better or worse version of human judgment but a third informant class with its own validity profile — a finding for personality science independent of any application.

Keywords: person perception, acquaintanceship, informant accuracy, self–other agreement, machine judgment, research agenda


1. The Question the Benchmark Brackets

A measurement program for machine person-perception invites an obvious comparison, and this one has declined it on purpose.

The reason is not modesty. It is that the comparison is easy to run badly and the bad version is what would circulate. Machine and human judges differ in the evidence available to them, in the modality of that evidence, in the conditions under which it was gathered, and in what the judge was doing while gathering it. A headline of the form “AI beats friends at reading personality” collapses all of that into a number whose meaning depends entirely on design choices that never appear in the headline.

The literature already contains the seed of this problem. Machine judgments from digital behavioral residue have been found to exceed the accuracy of human informants who know the target personally (Youyou, Kosinski, & Stillwell, 2015) — a genuinely important result that is also, in its popular retelling, precisely the collapsed claim. The comparison deserves better than that, and the dataset assembled for the benchmark program makes better possible.

2. Why the Comparison Is Hard

Four asymmetries make naive comparison uninterpretable, and any serious design must address each.

Evidence volume and kind. A spouse has observed behavior across years, contexts, and stress conditions. A conversational model has text from one channel. These are not more and less of the same thing.

Acquaintance is confounded in humans. Human judges cannot be assigned to acquaintance levels. People who have known a target for twenty years differ systematically from people who have known them for two, and so do the relationships. Every human acquaintance estimate carries selection confounds that the machine condition — where exposure is manipulable within a single relationship — does not (Biesanz, West, & Millevoi, 2007).

The criterion is not neutral across judges. Self-report is the usual criterion, and it privileges some perspectives over others: the self has better access to internally defined traits while observers have advantages on observable and evaluative ones (Vazire, 2010). A comparison scored solely against self-report is not a fair contest; it is a contest on the self’s home ground.

Human judges are not a homogeneous class. Other-ratings carry substantial validity that varies with the rater’s relationship to the target and with how much they have observed (Connelly & Ones, 2010). “Human judges” is a category containing enormous variance, and pooling it produces a comparison against a fictional average person.

3. The Design

Participants. The establishing study’s participants, who already hold multi-instrument criterion profiles and have donated conversational corpora — which is why this study requires no new recruitment and can be specified in advance as a designated follow-on.

Human judge tiers. Each participant nominates judges at graded acquaintance: a zero-acquaintance judge (no prior contact, brief exposure), a low-acquaintance judge (a work colleague), a moderate-acquaintance judge (a friend), and a high-acquaintance judge (a partner, family member, or comparably close relation). Judges rate the participant on the same trait dimensions the criterion instruments measure, with confidence ratings, under identical instructions to those given the models.

Machine tiers. The dose ladder already specified — chronologically truncated corpora across a logarithmic token progression, evaluated by multiple models with stochastic repetitions.

Common criterion. Both judge classes are scored against the same multi-instrument consensus, with accuracy expressed against the empirically observed inter-instrument ceiling, per trait. This is what makes the comparison a comparison: identical criterion, identical trait definitions, identical scoring.

Matched elicitation. Human and machine judges receive the same rating instrument, the same anchors, and the same confidence elicitation. Differences in output format are a design flaw, not a finding.

The evidence-matched sub-condition. A subset of human judges rate the participant from the same truncated corpus the models receive, with no personal acquaintance. This condition is the only true like-for-like comparison in the design: identical evidence, different judge. It isolates judging capability from evidentiary advantage, and it should be treated as the study’s cleanest single contrast even though it is the least ecologically natural.

4. Four Predictions

P1 — Crossing points, not a verdict. The primary output is two curves and their intersections: the token dose at which machine accuracy equals colleague-level, friend-level, and close-relation-level human accuracy, estimated per trait. There will be no single answer to “does AI know you better,” and the design is built to make that visible rather than to force a scalar.

P2 — Divergent plateaus. Human accuracy is expected to flatten with acquaintance, consistent with findings that consensus rises less with acquaintance than intuition predicts (Kenny, 2004) and that longer observation does not reliably improve prediction (Ambady & Rosenthal, 1992). Machine accuracy may continue rising, because the constraints plausibly producing the human plateau — attention, memory, the impossibility of re-reading a relationship — do not apply. If machine curves keep climbing where human curves flatten, that is the strongest available evidence that machine person-perception is a distinct capability rather than a mechanical approximation of the human one.

P3 — Opposite trait profiles. This is the agenda’s most consequential prediction. A conversational system’s evidence is a person’s self-narration over time — including material never said aloud to any human. Human observers’ evidence is behavior. SOKA implies the two illuminate different regions of the trait space (Vazire, 2010). The prediction is therefore an interaction rather than a main effect: machines should be relatively strong on internally defined traits that observers read worst, and relatively weak on the observable behavioral traits observers read best. Confirmation would establish machine judges as a third informant class rather than a substitute for either existing perspective.

P4 — Complementarity beats either alone. If P3 holds, combining machine and human ratings should predict criterion scores better than either class alone, and the incremental validity should be largest for traits where the two classes have complementary access. This is the practically important prediction: it implies the right deployment question is not whether to trust the machine or the person, but how to combine them.

A stated non-prediction. Nothing here predicts that machines will be better or worse overall, and the design is built to prevent that question from having a clean answer. If a summary is demanded, the honest one is: better at some things, worse at others, and different in kind from both existing perspectives.

5. What Each Field Gains

Personality science gains a new perspective for its models. The self–other framework has two positions; this design tests whether a third exists, with evidence that is verbal, longitudinal, private, and self-narrated. That is not a machine-learning result borrowed by psychology — it is a psychological finding about what kinds of access to a person are possible, produced using a machine as an instrument.

Machine learning gains an accuracy reference class more meaningful than self-report agreement in the abstract. “This system reaches colleague-level accuracy at roughly this exposure, on these traits, and does not reach close-relation-level on these others” is interpretable in a way a correlation is not.

Deployment practice gains the complementarity result. If P4 holds, systems that present their reads as replacements for human judgment are misusing them, and the appropriate design is one that surfaces machine and human perspectives as distinct inputs.

6. Limits

Judge nomination is not random. Participants nominate their own judges, introducing selection at every acquaintance tier. This is unavoidable — one cannot assign spouses — and must be characterized rather than solved: judge relationship, contact frequency, and duration are measured and reported as covariates.

The evidence-matched condition is artificial. Asking a human to read a stranger from a token-truncated transcript is not how humans judge people, and its ecological validity is low even as its internal validity is the design’s highest. Both conditions are reported; neither is the answer alone.

Human judges are unblinded to their relationship. A spouse knows they are a spouse, and expectations about their own accuracy may shape their ratings and confidence. Machine judges have no such expectation, which is an asymmetry the design can measure through confidence data but cannot remove.

The criterion remains imperfect. Self-report instruments carry known biases and differing authority across traits (Vazire, 2010; Connelly & Ones, 2010). Ceiling-referencing handles this honestly rather than eliminating it, and per-trait reporting keeps the criterion’s uneven authority visible.

Version dependence. Machine results are properties of specific model versions. Human results are not, which means the comparison ages asymmetrically and requires re-running rather than updating.

7. Conclusion

The question the public asks about this technology is the one the benchmark program has refused to answer, and the refusal was a sequencing decision rather than an evasion. Measurement first, comparison second: a contest between two judges is meaningless until the scoring is settled, the criterion is defined, and the ceiling is known.

With those in place, the comparison becomes worth running — not to determine a winner, but to locate two curves relative to each other and to find where they cross. The most likely finding is the least satisfying to headline and the most useful to know: that a machine reading a person’s self-narration and a human watching that person live are doing different things, arriving at different parts of the truth, and better together than either alone.

The dataset for this study will already exist. The design is specified above. It should be run by someone with no stake in the answer.


References

Ambady, N., & Rosenthal, R. (1992). Thin slices of expressive behavior as predictors of interpersonal consequences: A meta-analysis. Psychological Bulletin, 111(2), 256–274.

Biesanz, J. C., West, S. G., & Millevoi, A. (2007). What do you learn about someone over time? The relationship between length of acquaintance and consensus and self–other agreement in judgments of personality. Journal of Personality and Social Psychology, 92(1), 119–135.

Connelly, B. S., & Ones, D. S. (2010). An other perspective on personality: Meta-analytic integration of observers’ accuracy and predictive validity. Psychological Bulletin, 136(6), 1092–1122.

Kenny, D. A. (2004). PERSON: A general model of interpersonal perception. Personality and Social Psychology Review, 8(3), 265–280.

Vazire, S. (2010). Who knows what about a person? The self–other knowledge asymmetry (SOKA) model. Journal of Personality and Social Psychology, 98(2), 281–300.

Youyou, W., Kosinski, M., & Stillwell, D. (2015). Computer-based personality judgments are more accurate than those made by humans. Proceedings of the National Academy of Sciences, 112(4), 1036–1040.


Citation compliance note: all 6 references above were verified against live sources on August 16, 2026, in the session that produced this draft series. No reference is cited from memory.