Calibration as Co-Equal Criterion: Confidence–Accuracy Coupling in Machine Judgments of Humans
Abstract
Two systems read the same user. The first is right about 60 percent of what a validated instrument would capture, and it reliably signals when it is unsure. The second is right about 65 percent and expresses the same high confidence no matter what it has seen.
On a single accuracy score, the second system wins. In any deployment where a wrong read carries cost, the first is safer — because its errors announce themselves and the second's do not.
This article argues that confidence--accuracy coupling is not a supplementary diagnostic in machine person-perception but a co-equal headline criterion, and that the two must never be composited into one number. The argument has three parts. First, the theoretical case: knowing the limits of one's read on a person is constitutive of competence in judging persons, not an add-on to it. Second, the empirical case: modern neural systems are documented to be systematically overconfident, and the failure has a specific and dangerous form when the judgment is about a person rather than an image class. Third, the measurement case: the machinery for scoring confidence against accuracy is seventy-six years old and transfers directly, with three adaptations the person-perception setting requires — dose-conditional calibration, trait-conditional calibration, and a fabrication index defined at the floor of exposure.
The recommendation is concrete. Report accuracy and calibration as two headline dimensions, never one. A system that is confidently wrong about people is not a slightly worse version of a system that is uncertainly wrong. It is a different kind of object, and the reporting format should make that visible before anyone deploys it.
Keywords: calibration, overconfidence, confidence--accuracy, person perception, benchmark reporting, large language models
1. The Deployment Question a Single Number Cannot Answer
Imagine buying an assistant for a clinical intake team, a wealth advisory practice, or a school. You are shown one number: the system reads people at 65 percent of instrument-grade accuracy.
You still cannot make the decision, and the reason is not that the number is too low. It is that the number omits the variable that governs how the system's errors will propagate. If the system flags its uncertain reads, your staff can treat those as prompts to ask rather than conclusions to act on, and the system's 35 percent shortfall degrades gracefully. If it does not — if it delivers every judgment with the same assurance — then its errors arrive dressed identically to its successes, and the shortfall propagates silently into decisions about people.
Two systems with identical accuracy are therefore not equivalent products. And a system with slightly lower accuracy and honest uncertainty may be the better one to deploy. No accuracy score, however carefully constructed, contains this information.
The remedy is not a better composite. It is refusing to composite.
2. Knowing What You Don't Know Is Part of Knowing
The strong version of the argument is that calibration is not adjacent to person-perception competence but internal to it.
Consider what expert human judges of persons actually do. A skilled clinician forms provisional impressions and holds them provisionally, revising as evidence accumulates. What distinguishes expertise is not the absence of error but the tracking of error: knowing which impressions rest on thin evidence, which traits are hard to read in this particular person, and when to withhold judgment altogether.
The person-perception literature supplies the theoretical basis for treating this as constitutive rather than optional. Accurate judgment requires that relevant information be available, detected, and correctly utilized (Funder, 1995, 2012); each stage can fail, and a judge with no representation of which stage is failing has no basis for modulating confidence. The information available about a person varies enormously in quantity and quality (Letzring, Wells, & Funder, 2006), and traits differ in how legible they are from any given vantage point — the self and observers hold asymmetric access to different regions of the trait space (Vazire, 2010). A judge insensitive to these differences, expressing uniform confidence across traits and across evidence conditions, is not merely miscalibrated. It is failing to represent the structure of the judgment problem.
This is why calibration cannot be demoted to a robustness check. A system that produces trait ratings without any internal model of their evidentiary basis is doing something less than judging. It is generating.
3. What Is Already Known About Machine Confidence
The empirical prior is not encouraging, and it is specific rather than vague.
Modern neural networks are systematically overconfident: their reported confidence exceeds their accuracy, and the discrepancy has been documented across architectures and tasks (Guo, Pleiss, Sun, & Weinberger, 2017). Separately, generative systems produce fluent, well-formed output regardless of whether the underlying evidence supports it — the tendency documented across natural language generation as hallucination (Ji et al., 2023). And assistant-tuned systems exhibit sycophancy: outputs drift toward what the user appears to want, traceable in part to training feedback that rewards agreeable responses (Perez et al., 2022; Sharma et al., 2023).
Each of these becomes more consequential when the object of judgment is a person.
An overconfident image classifier produces a wrong label with a high score. An overconfident person-judge produces a character portrait — coherent, specific, plausible, and wrong — which the recipient has strong psychological reasons to accept. People extend social responses to machines automatically (Nass & Moon, 2000) and attribute humanlike understanding under predictable conditions (Epley, Waytz, & Cacioppo, 2007). Fluent confidence about a person is therefore received in a way fluent confidence about a photograph is not: it feels like being seen.
Sycophancy adds a further mechanism specific to this domain. A system whose ratings drift toward what the user wants to hear will produce confident judgments that are agreeable rather than accurate, and its confidence will be highest exactly where agreement is easiest — that is, where the user has signaled a self-view most clearly. This is the worst possible confidence structure: certainty concentrated where the evidence is most contaminated by the target's own preferences.
The literature on human overconfidence supplies the interpretive frame for what to expect and how to name it. Overconfidence is not one phenomenon but several distinguishable ones — overestimation of one's performance, overplacement relative to others, and overprecision in the certainty attached to beliefs (Moore & Healy, 2008). Machine judgments of persons plausibly exhibit the third form most acutely, and the distinctions matter because they imply different diagnostics.
4. The Measurement Machinery Transfers
Scoring confidence against accuracy is a solved problem with a long pedigree. Proper scoring rules were formalized for weather forecasting three-quarters of a century ago, and the decomposition they permit — separating calibration from resolution — remains the standard apparatus (Brier, 1950). Modern extensions to neural systems supply calibration curves, expected calibration error, and temperature-based post-hoc correction (Guo et al., 2017).
Four measurements constitute the core battery for person-perception:
Calibration curves. Bin judgments by reported confidence; plot mean confidence against observed accuracy within each bin. Perfect calibration lies on the diagonal; systematic elevation above it is overconfidence. This is the primary visual and the primary summary.
Proper scoring with decomposition. A proper score, decomposed into calibration and resolution components (Brier, 1950), separates two distinct virtues: whether stated confidences match observed frequencies, and whether the system discriminates between cases at all. A system can be well calibrated and useless (uniformly 50 percent confident and uniformly 50 percent accurate) or poorly calibrated and informative. Both components are reported.
Confidence--accuracy correlation within person. Across the trait dimensions judged for a single individual, does higher confidence accompany greater accuracy? This is the within-person analog of the classic confidence--accuracy question and is closest to the deployment-relevant property: when this system tells me it is sure about this aspect of this person, should I believe it more?
Overconfidence index. Mean confidence minus mean accuracy, reported per trait and per exposure level, with the Moore--Healy distinctions used to characterize its form rather than merely its magnitude (Moore & Healy, 2008).
5. Three Adaptations the Person-Perception Case Requires
Standard calibration analysis assumes a fixed evidence condition. Person-perception does not, and three adaptations follow.
Dose-conditional calibration. Because exposure to the person is manipulable — a conversational history can be truncated chronologically at any point — calibration can and must be measured at each exposure level. This yields a diagnostic unavailable in ordinary calibration work: does the system's confidence respond to how much it has seen? A well-behaved judge should be visibly less certain at ten thousand tokens than at a million. Flat confidence across a dose ladder is a specific, nameable failure — confidence decoupled from evidence — and it is arguably more damning than poor calibration at any single dose, because it demonstrates that the confidence signal is not tracking evidence at all.
Trait-conditional calibration. Traits are not equally legible from verbal behavior, and the perspective asymmetries in person knowledge are well established (Vazire, 2010; Connelly & Ones, 2010). A judge that is appropriately confident on traits it reads well and appropriately hesitant on traits it reads poorly has an internal model of its own competence structure. A judge with uniform confidence across traits does not. Trait-conditional calibration is therefore reported as a matrix, never as a single figure.
The fabrication index. At the floor of exposure — negligible evidence, full elicitation — genuine signal exists but is modest, as the zero-acquaintance literature establishes for human judges (Borkenau & Liebler, 1992). A valid judge at the floor should be modestly accurate and visibly uncertain. The fabrication index is defined as the confidence expressed at the floor condition, scaled against the accuracy actually attained there. A system that produces elaborate, assured portraits from almost nothing scores high on this index, and the number has a plain-language interpretation: how readily this system invents a person.
Of everything in this article, the fabrication index is the measurement most likely to matter outside the research literature, because it answers the question a regulator, a procurement officer, or a journalist would actually ask.
6. The Reporting Rule
Accuracy and calibration are reported as two headline dimensions and are never combined into a single score.
The rule is stated as a prohibition because the pressure to violate it will be constant. Leaderboards want scalars. Marketing wants a number. Benchmark scholarship has documented what is lost when heterogeneous capabilities are compressed into single figures, and the resulting distortion of the field's priorities (Liang et al., 2022; Raji et al., 2021).
Here the loss is not merely informational but decision-relevant. The trade-off between accuracy and calibration is precisely what a deployment decision must weigh, and it varies by application: a system used to draft low-stakes suggestions can tolerate confident error better than one used to brief a clinician. Compositing performs that weighting invisibly, once, on behalf of every downstream user, and there is no weighting that is correct for all of them.
Two practical corollaries follow. First, any published benchmark table must present the two dimensions side by side, with neither derivable from the other. Second, claims of the form "System X is the most accurate person-reader" are incomplete by construction and should be treated as such by reviewers and reporters alike.
7. Empirical Protocol and Pre-Registered Hypotheses
Elicitation. Every trait judgment is accompanied by a confidence rating on a standardized, anchored scale, using prompt language frozen during piloting and held identical across models. Confidence is elicited per trait, not per profile, because the trait-conditional analysis requires it.
Conditions. Confidence data are collected at every cell of the dose × trait × model design specified in a companion article, with multiple stochastic repetitions per cell.
Pre-registered hypotheses, stated before any model output has been observed (Nosek, Ebersole, DeHaven, & Mellor, 2018; Simmons, Nelson, & Simonsohn, 2011):
C1. Systems will exhibit net overconfidence: mean confidence will exceed ceiling-referenced accuracy, consistent with the documented tendency of modern neural systems (Guo et al., 2017).
C2. Confidence will be substantially less responsive to exposure than accuracy is. Formally, the slope of confidence across the dose ladder will be shallower than the slope of accuracy. This predicts the decoupling signature described in Section 5.
C3. Confidence will be less differentiated across traits than accuracy is, indicating a weak internal model of trait-specific legibility.
C4. Fabrication indices will vary substantially across models, and this variation will not be predicted by accuracy rank. That is, the most accurate system will not necessarily be the least fabricating one — the finding that would most directly justify separate reporting.
C5 (exploratory). Sycophancy-linked structure: confidence will be elevated where the corpus contains explicit self-description by the participant, relative to matched cells without it, indicating that agreement with a stated self-view inflates certainty. This is exploratory because the scrubbing protocol removes assessment-related self-description, leaving only the subtler forms.
Results. [Reserved for the version of record.]
8. Conclusion
The field's implicit position is that accuracy is the criterion and calibration is a refinement. This article has argued the opposite ordering is closer to right: for judgments about persons, a system that does not know the limits of its own read is not a less accurate judge but a less trustworthy object, and no accuracy figure can carry that information.
The measurement machinery has existed since 1950. The adaptations needed for this domain are three, and each is straightforward: measure calibration at every exposure level, measure it per trait, and define a fabrication index at the floor where evidence is negligible and elicitation is complete.
What follows is a reporting discipline rather than a scientific difficulty. Two numbers, side by side, neither derived from the other, for every system evaluated. The alternative — one number, weighting performed invisibly — hands every downstream decision-maker a figure that cannot answer the question they are asking.
A system that is confidently wrong about a person is not a slightly worse version of one that is uncertainly wrong. It should not be scored as though it were.
References
Borkenau, P., & Liebler, A. (1992). Trait inferences: Sources of validity at zero acquaintance. Journal of Personality and Social Psychology, 62, 645--657.
Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1--3.
Connelly, B. S., & Ones, D. S. (2010). An other perspective on personality: Meta-analytic integration of observers' accuracy and predictive validity. Psychological Bulletin, 136(6), 1092--1122.
Epley, N., Waytz, A., & Cacioppo, J. T. (2007). On seeing human: A three-factor theory of anthropomorphism. Psychological Review, 114(4), 864--886.
Funder, D. C. (1995). On the accuracy of personality judgment: A realistic approach. Psychological Review, 102(4), 652--670.
Funder, D. C. (2012). Accurate personality judgment. Current Directions in Psychological Science, 21(3), 177--182.
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (PMLR 70, pp. 1321--1330).
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248, 1--38.
Letzring, T. D., Wells, S. M., & Funder, D. C. (2006). Information quantity and quality affect the realistic accuracy of personality judgment. Journal of Personality and Social Psychology, 91(1), 111--123.
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al. (2022). Holistic evaluation of language models. arXiv:2211.09110.
Moore, D. A., & Healy, P. J. (2008). The trouble with overconfidence. Psychological Review, 115(2), 502--517.
Nass, C., & Moon, Y. (2000). Machines and mindlessness: Social responses to computers. Journal of Social Issues, 56(1), 81--103.
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600--2606.
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al. (2022). Discovering language model behaviors with model-written evaluations. arXiv:2212.09251.
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., et al. (2023). Towards understanding sycophancy in language models. arXiv:2310.13548.
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359--1366.
Vazire, S. (2010). Who knows what about a person? The self--other knowledge asymmetry (SOKA) model. Journal of Personality and Social Psychology, 98(2), 281--300.
Citation compliance note: all 17 references above were verified against live sources on August 16, 2026, in the session that produced this draft series. No reference is cited from memory.
RELATED
- VIDEO
Scrubbing the Answer Key
A spoken walkthrough of contamination control and pre-registration discipline: four classes of leakage, auditable removal rules, and how findings stay confirmatory.
- VIDEO
The Fluency Demand Taxonomy: Scoring Work Tasks FD0–FD4
A spoken walkthrough of the second axis of occupational exposure: five ordinal levels scoring how much a task’s outcome depends on communication fitted to a specific person.
- VIDEO
Calibration as Co-Equal Criterion
A spoken walkthrough of why confidence–accuracy coupling belongs beside accuracy as a headline criterion, never composited into a single number.