← Back to blog
ARTICLE

The Halo in the Machine: Discriminant-Validity Failure and Impression Compression in LLM Person-Judgments

James Niblick, PsyD10 min read

Abstract

Ask a language model to assess ten different people from their conversation histories, and read the ten assessments side by side.

They will be different. They will also, we predict, be more similar to each other than the ten people are — clustered near the population mean, tilted toward the flattering end of every evaluative dimension, and correlated across traits in a way the criterion instruments are not. Everyone comes out reasonably conscientious, fairly open, mostly stable, and basically agreeable.

This is halo: the collapse of distinct judgments onto a single global impression. Thorndike identified it in human raters in 1920. This article argues it has a machine descendant with an additional mechanism behind it — systems trained toward inoffensiveness have gradient pressure toward the flattering center that human raters do not — and that it is the most commercially dangerous failure mode in machine person-perception, because it is invisible to the statistics vendors report.

A system exhibiting halo compression can post respectable correlations with criterion instruments while carrying almost no information distinguishing one user from another. It looks like it is working. The article specifies three measurements that catch it — multitrait–multimethod structure, between-person variance ratios, and evaluative tilt — states pre-registered hypotheses, and argues that variance compression rather than correlation should be the headline diagnostic for whether a system is personalizing at all.

Keywords: halo effect, discriminant validity, multitrait-multimethod, variance compression, person perception, large language models


1. The Thermometer That Always Reads 98.6

Consider a clinical thermometer that returns 98.6 degrees for every patient, with small random variation. Across a population of mostly healthy people, its readings would correlate positively with true temperature — because most people are near normal, and so are its readings. Its mean would be approximately correct. On several plausible summary statistics, it would look like a working instrument.

It is useless. It carries no information about any individual patient, which is the only thing a thermometer is for.

This article’s claim is that machine person-judgment is structurally vulnerable to exactly this failure, that the vulnerability has a specific mechanism, and that the statistics currently reported would not detect it.

The failure has a century-old name. Thorndike observed that raters assessing people on multiple dimensions produced ratings contaminated by a general impression, such that a person judged favorably on one attribute was judged favorably on others regardless of the evidence — a constant error in psychological ratings (Thorndike, 1920). The machinery for detecting it is nearly as old: the multitrait–multimethod matrix was designed precisely to distinguish genuine trait variance from variance attributable to method and general impression (Campbell & Fiske, 1959).

What is new is the mechanism, and it is worth being precise about why the machine case may be worse rather than merely analogous.

2. Why Machines Should Compress More Than Humans Do

Three pressures push machine person-judgments toward the flattering center, and none of them operates on a human rater in the same way.

Training toward inoffensiveness. Assistant-tuned systems are optimized against human preference data, and that data rewards responses people find agreeable. Model-written evaluations documented systematic behavioral tendencies emerging from this training regime (Perez et al., 2022), and subsequent work traced sycophantic behavior in part to preference data in which convincingly agreeable responses are preferred over correct ones at non-negligible rates (Sharma et al., 2023). A characterization of a person that is unflattering is, by construction, a candidate for being rated poorly by human evaluators. The gradient points toward charitable readings.

Generation from central tendency. Where evidence is thin, a generative system produces well-formed output regardless of evidentiary support (Ji et al., 2023). The most probable well-formed characterization of an unknown person is the modal one — which is exactly the population average. Thin evidence therefore does not produce uncertainty in the output; it produces the population mean, delivered fluently.

No cost for being wrong in the safe direction. A human rater who describes everyone as agreeable faces social consequences when the description fails. A deployed system faces none, and the errors that would be noticed — describing a user unflatteringly and being wrong — are precisely the errors that training suppresses. The asymmetry is systematic rather than random.

The prediction that follows is not that machines will be inaccurate. It is that they will be accurate in the wrong way: correct about people in general, uninformative about any particular person, and unlikely to say anything a person would rather not hear.

3. Why Correlations Will Not Catch It

This is the article’s central practical claim, and it explains why the failure could persist unnoticed through a great deal of apparently successful evaluation.

Consider a system whose ratings have correct means and compressed variance — the right center, too narrow a spread. Its correlation with criterion scores across persons can be substantially positive, because correlation is scale-invariant: it measures whether the system ranks people similarly, not whether it distinguishes them by the right amount. A system that ranks correctly but compresses everyone toward the middle will report a respectable coefficient while conveying a fraction of the information the coefficient implies.

Now add the second signature. If the system’s ratings across traits are more intercorrelated than the criterion’s — everyone who reads as conscientious also reads as agreeable and stable — then the ratings are carrying one dimension wearing five labels. Per-trait correlations may each look acceptable while the profile as a whole contains a single judgment.

Neither problem appears in the statistic most likely to be reported. Both appear immediately in the analyses specified below, which is the argument for making those analyses mandatory rather than supplementary.

4. Three Measurements

Multitrait–multimethod structure. Construct the matrix with methods defined as {each criterion instrument, each model} (Campbell & Fiske, 1959). The diagnostic comparisons are the classic ones: monotrait–heteromethod values should exceed heterotrait–heteromethod values, and the pattern of trait intercorrelations within the system’s own ratings should resemble the criterion’s pattern rather than exceeding it uniformly. Halo appears as elevated heterotrait values within the model’s method block — the signature of a general impression contaminating every judgment.

Between-person variance ratio. For each trait, compute the ratio of the model’s between-person variance in ratings to the criterion’s between-person variance, both standardized. A ratio near 1 indicates the system distinguishes people by roughly the right amount. A ratio well below 1 is compression: the system is rating everyone more similarly than they are.

This ratio deserves to be a headline number. It is interpretable without training — this system rates people as being about half as different from one another as they actually are — and it captures the failure that correlations miss.

Evaluative tilt. For traits with a socially desirable pole, compute the mean displacement of the system’s ratings from the criterion’s mean, signed toward desirability. Positive displacement across evaluative traits is the flattering-center signature, distinguishable from simple compression because it has a direction. Compression without tilt suggests uncertainty collapsing toward the mean; compression with tilt suggests the training-pressure mechanism in Section 2.

Two auxiliary analyses complete the picture. Within-model reliability — agreement across repeated stochastic runs on identical evidence — bounds everything else, since a judge that disagrees with itself cannot distinguish people reliably; this is the machine analog of perceiver reliability in the componential tradition (Kenny, 2004). And profile distinctiveness — within-person profile correlations with the normative profile partialled out — tests whether individual profiles carry person-specific shape or merely reproduce the population’s average profile.

5. Pre-Registered Hypotheses

Stated before any model output has been observed, per the pre-registration standards this program adopts (Nosek, Ebersole, DeHaven, & Mellor, 2018; Simmons, Nelson, & Simonsohn, 2011).

H1 — Variance compression. Between-person variance ratios will be below 1 for most traits and most models: systems will rate people as more similar to one another than the criterion instruments do.

H2 — Elevated heterotrait structure. Trait intercorrelations within model ratings will exceed those within criterion measurement, indicating a general-impression component beyond what the criterion’s own trait covariance explains.

H3 — Evaluative tilt. Displacement toward the socially desirable pole will be positive on evaluative traits and larger than displacement on traits with weaker desirability loadings.

H4 — Compression decreases with exposure, tilt does not. As corpus exposure increases, variance compression should attenuate — more evidence permitting more differentiation — while evaluative tilt should persist, because tilt originates in training pressure rather than in evidence scarcity. This differential prediction is the sharpest available test that the two phenomena have distinct sources, and it is the hypothesis whose disconfirmation would most inform the diagnosis.

H5 — Correlation–information dissociation. Rank-order correlations with the criterion will be substantially positive even where variance ratios show marked compression, demonstrating empirically that correlation is insufficient as a personalization diagnostic. This hypothesis is included precisely because it targets the reporting practice this article seeks to change.

H6 (exploratory). Compression will vary by trait visibility, with larger compression on traits less legible from verbal behavior — consistent with the perspective asymmetries in person knowledge (Vazire, 2010) and with the finding that observer accuracy varies by trait and by information conditions (Connelly & Ones, 2010).

6. Results

[Reserved. The version of record will report: full MTMM matrices per model; between-person variance ratios by trait, model, and exposure level; evaluative tilt estimates; within-model reliability; profile distinctiveness with and without normative partialling; and the H4 differential-attenuation test.]

7. Implications

For evaluation. Variance ratio belongs beside correlation in every report of machine person-perception. A benchmark reporting only correlations cannot distinguish a system that knows people from one that has learned the population and ranks approximately. If H5 holds, that distinction is not hypothetical: the two will produce similar headline coefficients.

For product claims. A system with a variance ratio of, say, 0.4 is delivering personalization at less than half amplitude — its “personalized” outputs differ across users far less than the users differ. This is checkable by any vendor on their own system, today, without a benchmark, and the fact that it is not currently reported is informative about what is currently measured.

For the failure’s plausibility. The mechanism proposed in Section 2 predicts that the most agreeable systems will compress the most. If confirmed, this creates a direct tension between two things AI laboratories currently optimize simultaneously: inoffensiveness in output and accuracy in person-modeling. Halo compression would be the measurable form of that tension, and it would imply that reducing it requires accepting characterizations that some users will not enjoy reading.

For personality science. The machine case supplies an unusual test of halo theory. Human halo is typically attributed to cognitive economy — a single impression is cheaper than five independent judgments. Machine judges face no such constraint; they can attend to all evidence simultaneously. If they nonetheless exhibit halo, cognitive economy cannot be the whole explanation, and the alternative account — that halo arises from optimization pressure toward agreeable outputs, or from generation defaulting to central tendency under thin evidence — becomes worth taking seriously as a general mechanism rather than a machine-specific artifact.

8. Conclusion

The failure this article describes does not look like failure. A system exhibiting full halo compression produces assessments that are readable, plausible, mostly correct in aggregate, and flattering. Users will not complain. Correlational summaries will look acceptable. Nothing in the standard reporting apparatus will flag it.

What such a system will not do is distinguish one person from another — which is the entire content of the claim that it understands its users.

Thorndike’s finding was that raters produce a constant error when a general impression contaminates specific judgments. A hundred years later, the systems most likely to reproduce that error are the ones trained hardest to be agreeable, and the statistic most likely to conceal it is the one everyone reports.

Report the variance ratio.


References

Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105.

Connelly, B. S., & Ones, D. S. (2010). An other perspective on personality: Meta-analytic integration of observers’ accuracy and predictive validity. Psychological Bulletin, 136(6), 1092–1122.

Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248, 1–38.

Kenny, D. A. (2004). PERSON: A general model of interpersonal perception. Personality and Social Psychology Review, 8(3), 265–280.

Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606.

Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al. (2022). Discovering language model behaviors with model-written evaluations. arXiv:2212.09251.

Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., et al. (2023). Towards understanding sycophancy in language models. arXiv:2310.13548.

Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366.

Thorndike, E. L. (1920). A constant error in psychological ratings. Journal of Applied Psychology, 4(1), 25–29.

Vazire, S. (2010). Who knows what about a person? The self–other knowledge asymmetry (SOKA) model. Journal of Personality and Social Psychology, 98(2), 281–300.


Citation compliance note: all 10 references above were verified against live sources on August 16, 2026, in the session that produced this draft series. No reference is cited from memory.