The Unmeasured Instrument: AI Person-Perception as Psychometrics' Missing Validation Problem
Abstract
Every major AI platform now claims, in one phrasing or another, to understand its users. No one can check.
The industry benchmarks mathematics, coding, factual recall, and long-context retrieval. It does not benchmark the capability that every assistant, tutor, advisor, and companion product silently depends on: how accurately a system models the individual human in front of it. This article argues that the gap is not a hard measurement problem but an unclaimed one. A memory-enabled AI rendering judgments about its user is not like a psychometric instrument. It is functioning as one — a non-ipsative, observer-report measure whose item pool is the interaction history it has observed. Once that is seen, a century of validation methodology transfers intact: construct validity, convergent and discriminant analysis, dose-dependent accuracy, criterion ceilings, and calibration. The article specifies what such a validation requires, names three failure modes it must be built to catch — fabricated familiarity, halo compression, and sycophantic drift — answers the obvious objections, and argues that the moment to establish an independent, pre-registered standard is now, before a vendor's self-assessment becomes the reference point by default.
Fluency in mathematics is benchmarked. Fluency in persons is asserted. That asymmetry has outlived its excuse.
Keywords: person perception, personality assessment, construct validity, large language models, benchmarks, calibration
1. A Claim Without a Measure
A model's performance on graduate-level mathematics is a public number. So is its accuracy on medical licensing questions, its success rate on competitive programming problems, and its ability to retrieve a fact buried in a million tokens of context. When a laboratory claims improvement, the claim is checkable, because the field agreed — imperfectly but publicly — on what would count as evidence (Hendrycks et al., 2021; Liang et al., 2022).
Now ask a simpler question about the same systems. This assistant has talked with a particular person for eight months. How well does it know them?
There is no number. There is no test. There is not even an agreed definition of what the number would measure — despite the fact that this capability, and not competitive programming, is what every assistant, tutor, advisor, and companion product is actually selling. Call it person-perception, borrowing from the psychological literature that has studied it in human judges for a century.
The consequences of getting it wrong are not subtle. A system that misreads a user's decisiveness will over-explain or under-explain every answer it gives. One that misreads risk tolerance will calibrate advice to a person who is not there. One that cannot distinguish this user's trait profile from the next user's is not personalizing at all; it is running a generic script with someone's name attached.
And the claim is not marketing garnish. It is the load-bearing premise beneath personalization features, memory architectures, and a growing share of enterprise AI. Benchmark scholars have warned that influential evaluations can acquire authority their construct validity cannot support (Raji et al., 2021). Here the failure runs the other way, and it is worse: a consequential capability claim circulating with no evaluation at all.
The measure is missing not because it is hard to build, but because two fields have not noticed that one of them already built it. Psychology spent a hundred years on the discipline of instrument validation — what it means for a new way of measuring persons to be valid, and how to find out (Cronbach & Meehl, 1955; Campbell & Fiske, 1959). Machine learning spent the last several years shipping systems that render trait judgments about persons at population scale, without asking whether those judgments would survive the validation any conventional instrument must pass.
The argument here is that the transfer is direct. An AI system judging a person is a candidate psychometric instrument, and should be validated as one.
2. The Functional Argument: AI as Candidate Instrument
Calling an AI an instrument sounds like a metaphor. It is not. It is a functional description, and it can be written in the standard vocabulary of assessment.
Watch what a memory-enabled system does across eight months with a user. It observes an extended sample of that person's verbal behavior — word choices, topics, decisions narrated aloud, reactions to suggestions, revisions, hesitations, requests. From that sample it forms a representation of the person, explicit or not, and that representation shapes what it explains, how deeply, in what tone, and what it recommends. Ask it to state the representation out loud — "this user prefers directness," "this user is risk-averse" — and what comes out is an observer-report trait rating derived from a behavioral sample.
That is a non-ipsative, observer-report measure whose item pool is the interaction history. Every term earns its place. It is observer-report rather than self-report: the system occupies the position of the informant in the classic self–other paradigm, with the informational advantages and disadvantages that position carries (Vazire, 2010; Connelly & Ones, 2010). It is non-ipsative: nothing forces its judgments into intra-individual rankings, so its ratings can in principle be correlated with normative trait scores. And its item pool is naturalistic rather than administered: the "items" are whatever the relationship happened to contain, which places the system squarely within the tradition of judging persons from unstructured behavioral information rather than questionnaire responses (Ambady & Rosenthal, 1992; Borkenau & Liebler, 1992).
This is not speculation about a future capability. Machine models rating personality from digital behavioral residue have produced valid trait estimates and, under some conditions, beaten the accuracy of human informants who know the target personally (Kosinski, Stillwell, & Graepel, 2013; Youyou, Kosinski, & Stillwell, 2015), with language-based assessment converging with self-report across very large samples (Schwartz et al., 2013). Contemporary large language models infer users' psychological dispositions from their text with no training for the task at all (Peters & Matz, 2024). The capability exists. It is deployed. Its validation does not exist.
One asymmetry in the current literature is worth naming. A growing research program administers personality inventories to language models, asking what synthetic traits the models themselves display (Serapio-García et al., 2023). That work treats the model as a test-taker. Almost nothing treats the model as what it functionally is in every deployed personalization context: a test-giver, rendering consequential judgments about human beings.
The field has been characterizing the psychology of the judge while declining to score the judge's judgments.
3. The Machinery Already Exists
If an AI is a candidate instrument, then "is it any good?" is not a new question. It is the oldest question in psychometrics, and the answers transfer with almost embarrassing directness.
Construct validity asks whether the instrument measures what it purports to measure (Cronbach & Meehl, 1955). For an AI judging extraversion from conversation, the question is whether its "extraversion" ratings track the construct as established instruments measure it — not whether they track user satisfaction, engagement, or agreement, which are the criteria deployed systems are actually optimized against. The distinction is not pedantic; Section 4 argues it is where the most consequential failure mode lives.
Convergent and discriminant validation supplies the analytic architecture (Campbell & Fiske, 1959). The multitrait–multimethod logic maps cleanly onto the machine case: AI ratings of trait X should correlate with validated instrument measurements of trait X (convergence) and should correlate less with instrument measurements of trait Y (discrimination). A system whose trait ratings collapse into a single evaluative dimension — rating agreeable people as also conscientious, stable, and open — is exhibiting the machine descendant of the constant error Thorndike identified in human raters over a century ago (Thorndike, 1920). The MTMM matrix was built to detect exactly this, and it does not care whether the method is a questionnaire, an informant, or a transformer.
The accuracy literature supplies the process theory and the moderators. Funder's Realistic Accuracy Model decomposes accurate judgment into the availability, detection, and utilization of relevant cues (Funder, 1995, 2012) — a decomposition that applies without modification to a system whose available cues are the tokens of an interaction history. The acquaintanceship literature establishes that human judgmental accuracy improves with information quantity and quality, with diminishing returns and trait-dependent trajectories (Letzring, Wells, & Funder, 2006; Biesanz, West, & Millevoi, 2007; Kenny, 2004). This yields the machine analog directly: exposure can be quantified in tokens of interaction history, truncated chronologically to produce a dose–response curve — accuracy as a function of how much of the person the system has seen. Whether machine curves inflect and plateau where human curves do is an open empirical question of considerable theoretical interest; that it is the right question is settled by the human literature.
Reliability theory and the criterion ceiling supply honesty about the maximum. Two established instruments measuring the same trait do not agree perfectly with each other (McCrae & Costa, 1987; Pace & Brannick, 2010) — and nothing can agree with a criterion more strongly than the criterion agrees with itself. Machine accuracy must therefore be scored against the observed inter-instrument ceiling rather than against a perfect criterion that does not exist. This cuts both ways, which is the point: it forecloses the inflated claim ("90% accurate!") and the unfair penalty (charging a system for noise no instrument can explain). The stability requirements of correlational estimation additionally discipline study design: correlations computed across persons do not stabilize in small samples (Schönbrodt & Perugini, 2013), which constrains what any credible validation can claim and at what n.
Calibration theory supplies the co-equal criterion. A judge of persons who is moderately accurate and knows the limits of its own read differs categorically from a judge who is somewhat more accurate and uniformly, unjustifiably confident. The formal machinery for scoring the confidence–accuracy relationship is mature (Brier, 1950; Moore & Healy, 2008), and its extension to modern machine-learning systems is well developed — with the sobering standing result that contemporary neural systems tend toward systematic overconfidence (Guo et al., 2017). For judgments about persons, calibration is not a technical footnote. It is half of what fluency in persons means.
None of this machinery needs to be invented. It needs to be pointed at the right target.
4. What Validation Must Be Built to Detect
A standard earns its authority partly by what it commits in advance to catching. Three failure modes qualify, because each turns apparent competence into measurement failure.
Fabricated familiarity. Generative systems produce fluent, confident output whether or not the evidence supports it, a tendency documented across natural language generation (Ji et al., 2023). Pointed at a person, it produces something specific and unsettling: a rich, confident character portrait of someone the system has barely met. The zero-acquaintance literature establishes that some genuine signal is available even in minimal exposure (Borkenau & Liebler, 1992) — which is exactly why fabrication is detectable: a valid judge's accuracy and confidence at near-zero exposure should be modest and should scale with evidence, whereas a fabricating judge produces rich profiles whose confidence bears no relationship to dose. A floor condition — negligible exposure, full elicitation — functions as a fabrication detector, and a system's behavior under it should be a headline result of any evaluation.
Halo compression. As noted above, the MTMM machinery exists to detect judgments that collapse into global impressions (Campbell & Fiske, 1959; Thorndike, 1920). For machine judges there is a plausible mechanism beyond cognitive economy: systems trained toward inoffensiveness may rate all persons as near-population-mean on socially desirable traits, producing adequate mean-level agreement with near-zero between-person discrimination. On aggregate statistics such a system looks accurate. It carries no individual information whatsoever. It is a thermometer that always reads 98.6.
Sycophantic drift. The best-documented behavioral distortion in assistant-tuned language models is sycophancy: the tendency, traceable in part to human feedback on which the systems are trained, to produce outputs that match the user's beliefs and preferences rather than the truth (Perez et al., 2022; Sharma et al., 2023). Restated psychometrically, sycophancy is construct invalidity. An instrument whose readings drift toward what the subject wants to hear is not measuring the subject. A flattering scale is not a scale.
That restatement does two things. It converts a product-quality complaint into a falsifiable measurement property. And it explains why the missing standard has gone unmissed: people experience agreement as being understood (Nass & Moon, 2000), so a system optimized to agree feels perceptive exactly when it is measuring nothing. The market cannot tell the feeling from the fact. Only measurement can.
A validation standard must also specify its criterion with the same discipline. Ground truth should be defined by the convergence of multiple validated instruments from independent traditions — for example, a public-domain hierarchical Big Five measure (Soto & John, 2017) alongside an independently developed six-factor inventory (Lee & Ashton, 2004) — with results reported per instrument so that no single measure carries the conclusions, and with cells where the instruments themselves diverge excluded rather than counted against the candidate. Where the stimulus is a person's own conversational history, contamination control becomes a first-order design problem: segments in which the person discusses their own assessment results would allow a system to read the answer key rather than infer the person, and their removal must follow a published codebook under blind audit. The broader methodological-reform literature supplies the enforcement mechanism — pre-registration of the analysis plan, prompts, exclusion rules, and failure-mode predictions before any model output is observed (Nosek et al., 2018; Simmons, Nelson, & Simonsohn, 2011).
5. Why the Standard Must Exist, and Why Now
Three arguments elevate this from a methodological curiosity to an obligation of the field.
The trust argument. People extend social responses to machines automatically and reliably (Nass & Moon, 2000), anthropomorphize under predictable conditions (Epley, Waytz, & Cacioppo, 2007), and have been projecting understanding onto conversational programs since the first one — a projection its own creator found alarming (Weizenbaum, 1966). Today's systems intensify every input to that dynamic: they converse fluently, they remember, and increasingly they speak aloud. Meanwhile the underlying capacity is genuinely advancing; frontier models now match or exceed human performance on substantial portions of theory-of-mind batteries (Strachan et al., 2024).
That combination is precarious. Interfaces that manufacture trust, capabilities that partly justify it, and no instrument marking where the justification stops. Trust calibrated to nothing is not trust. It is exposure.
The stakes argument. Person-perception claims are migrating into consequential settings. Psychological targeting demonstrably moves behavior at scale (Matz, Kosinski, Nave, & Stillwell, 2017), and large language models make personalized persuasion cheap to generate (Matz et al., 2024). Algorithmic judgment is restructuring managerial control inside organizations (Kellogg, Valentine, & Christin, 2020). The economic backdrop sharpens the point: as language models absorb routine cognitive work (Autor, Levy, & Murnane, 2003; Eloundou, Manning, Mishkin, & Rock, 2024; Felten, Raj, & Seamans, 2021), the labor market increasingly prices interpersonal capability (Deming, 2017) — and vendors will increasingly sell machine versions of it. There is instructive precedent for governing high-stakes judgment of persons with validity standards: a century of employment testing, whose meta-analytic self-corrections continue to this day (Sackett, Zhang, Berry, & Lievens, 2022), demonstrates both that such standards are achievable and that they require permanent methodological vigilance. Regulation is arriving regardless — comprehensive data-protection and AI regimes are already in force (Regulation [EU] 2016/679; Regulation [EU] 2024/1689) — and regulators will need exactly what the field currently lacks: a defensible way to measure whether a system's claims about persons are true.
The feasibility argument. What was missing until recently was not method but material. The methodology sections above require a stimulus set with three properties: authentic, longitudinal, and consented. Prior computational work necessarily relied on scraped social exhaust and short texts (Kosinski et al., 2013; Schwartz et al., 2013). But the deployment of conversational AI at scale has produced, for the first time, extended interaction histories between individual persons and AI systems — the exact substrate on which deployed systems actually form impressions — exportable by the person, with the person's consent, and joinable to instrument-based ground truth the person chooses to provide. The ecological validity of this corpus type is not incremental; it is categorical. And its consent architecture is not a compliance detail but a scientific commitment: the only person-perception research this framework validates or normalizes is research in which the profile belongs to the person.
A less elevated consideration argues for urgency too. Standards are shaped by whoever builds the first credible one. If the first widely cited evaluation of AI person-perception is a vendor's own self-assessment, the field spends a decade arguing with it. The alternative — pre-registered, multi-institution, scored against a criterion no commercial party controls — is available now, and the intellectual precedent for offering a strong claim in deliberately falsifiable form is a distinguished one: Mori advanced the uncanny valley explicitly as a conjecture inviting test, and it organized a research field for fifty years (Mori, 1970/2012).
6. Objections
"An AI is not an instrument — no fixed items, no standardization, no stable form." Neither is a human informant, and the field has treated informant judgment as measurement for decades, scoring its reliability and validity without pretending it is a questionnaire (Connelly & Ones, 2010; Kenny, 2004). The candidate-instrument framing requires only that the system produce trait judgments that can be scored against criteria — which deployed systems demonstrably can (Peters & Matz, 2024). Version instability is real and is handled the way psychometrics handles form changes: version-stamped evaluation with repeatable protocols, which is precisely what makes this a standing standard rather than a single study.
"Self-report ground truth is itself flawed, so the criterion is circular." The objection proves the design rather than defeating it. Self-report's imperfections are known, quantified, and partially complemented by the observer perspective (Vazire, 2010); the multi-instrument consensus criterion, per-instrument reporting, and ceiling-referenced scoring exist exactly because the criterion is imperfect. The standard does not claim the instruments are true; it claims they are the best-validated available definition of the construct, and it scores candidates against what those instruments can agree on — no more.
"Communication-heavy judgment tasks are too context-dependent for a general benchmark." Context-dependence is an argument for careful scope, not for no measurement. The claim under test is narrow and operational: given this behavioral record, how accurately does this system estimate this person's standing on well-defined trait dimensions, and how well does it know when it doesn't know? That is a bounded question with a hundred-year methodological literature behind it. What lies beyond it — whether accurate perception translates into benefit, and under what communicative conditions — is a separable research program with its own theoretical foundations in the study of communication itself (Shannon, 1948; Grice, 1975; Clark & Brennan, 1991; Daft & Lengel, 1986; Bell, 1984), and is deliberately bracketed here.
"This will be used to sell products." Yes. That is what benchmarks are for, and it is why governance is a design requirement rather than an afterthought. The commitments are structural: scientific leadership by academic collaborators across multiple institutions; pre-registration before data collection; open protocol, prompts, codebook, and scoring; per-instrument reporting; and a scoring methodology controlled by no commercial entity — a constraint that must include, explicitly and especially, any entity associated with this article's author. A standard that a claimant controls is a brochure. The entire value of the proposed standard lies in the fact that no one who might benefit from a particular result can produce it.
7. Conclusion: The Last Unmeasured Capability
The argument compresses to three lines. Deployed AI systems render trait judgments about the people they talk to. That activity is functionally psychometric measurement. Psychometric measurement without validation is precisely what a century of assessment science exists to prevent.
Without intending to, the industry has released millions of unvalidated instruments into the most sensitive measurement context there is — the modeling of individual human beings — and wrapped them in interfaces that make their judgments feel true whether or not they are.
Psychology is the discipline that knows how to answer the question these systems raise, because it is the discipline that spent a hundred years learning to ask it precisely: is this a valid way of measuring a person? The machinery — construct validity, convergent and discriminant analysis, accuracy process models, ceiling-referenced criteria, calibration — is built, tested, and waiting. The stimulus material now exists in consented form. The failure modes are specifiable in advance. What remains is for the field to claim the problem.
Colleagues with expertise in personality assessment, person perception, quantitative methods, and machine behavior are invited to challenge the framework's premises, sharpen its protocol, and lead the science that follows. The capability in question is the one every human-facing AI deployment silently presumes. It should be the next thing we measure, and the last thing we continue to take on faith.
References
Ambady, N., & Rosenthal, R. (1992). Thin slices of expressive behavior as predictors of interpersonal consequences: A meta-analysis. Psychological Bulletin, 111(2), 256–274.
Autor, D. H., Levy, F., & Murnane, R. J. (2003). The skill content of recent technological change: An empirical exploration. Quarterly Journal of Economics, 118(4), 1279–1333.
Bell, A. (1984). Language style as audience design. Language in Society, 13(2), 145–204.
Biesanz, J. C., West, S. G., & Millevoi, A. (2007). What do you learn about someone over time? The relationship between length of acquaintance and consensus and self–other agreement in judgments of personality. Journal of Personality and Social Psychology, 92(1), 119–135.
Borkenau, P., & Liebler, A. (1992). Trait inferences: Sources of validity at zero acquaintance. Journal of Personality and Social Psychology, 62, 645–657.
Brier, G. W. (1950). Verification of forecasts expressed in terms of probability. Monthly Weather Review, 78(1), 1–3.
Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105.
Clark, H. H., & Brennan, S. E. (1991). Grounding in communication. In L. B. Resnick, J. M. Levine, & S. D. Teasley (Eds.), Perspectives on socially shared cognition (pp. 127–149). American Psychological Association.
Connelly, B. S., & Ones, D. S. (2010). An other perspective on personality: Meta-analytic integration of observers' accuracy and predictive validity. Psychological Bulletin, 136(6), 1092–1122.
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302.
Daft, R. L., & Lengel, R. H. (1986). Organizational information requirements, media richness and structural design. Management Science, 32(5), 554–571.
Deming, D. J. (2017). The growing importance of social skills in the labor market. Quarterly Journal of Economics, 132(4), 1593–1640.
Eloundou, T., Manning, S., Mishkin, P., & Rock, D. (2024). GPTs are GPTs: Labor market impact potential of LLMs. Science, 384(6702), 1306–1308.
Epley, N., Waytz, A., & Cacioppo, J. T. (2007). On seeing human: A three-factor theory of anthropomorphism. Psychological Review, 114(4), 864–886.
Felten, E., Raj, M., & Seamans, R. (2021). Occupational, industry, and geographic exposure to artificial intelligence: A novel dataset and its potential uses. Strategic Management Journal, 42(12), 2195–2217.
Funder, D. C. (1995). On the accuracy of personality judgment: A realistic approach. Psychological Review, 102(4), 652–670.
Funder, D. C. (2012). Accurate personality judgment. Current Directions in Psychological Science, 21(3), 177–182.
Guo, C., Pleiss, G., Sun, Y., & Weinberger, K. Q. (2017). On calibration of modern neural networks. In Proceedings of the 34th International Conference on Machine Learning (PMLR 70, pp. 1321–1330).
Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., & Steinhardt, J. (2021). Measuring massive multitask language understanding. In Proceedings of the International Conference on Learning Representations (ICLR).
Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248, 1–38.
Kellogg, K. C., Valentine, M. A., & Christin, A. (2020). Algorithms at work: The new contested terrain of control. Academy of Management Annals, 14(1), 366–410.
Kenny, D. A. (2004). PERSON: A general model of interpersonal perception. Personality and Social Psychology Review, 8(3), 265–280.
Kosinski, M., Stillwell, D., & Graepel, T. (2013). Private traits and attributes are predictable from digital records of human behavior. Proceedings of the National Academy of Sciences, 110, 5802–5805.
Lee, K., & Ashton, M. C. (2004). Psychometric properties of the HEXACO Personality Inventory. Multivariate Behavioral Research, 39(2), 329–358.
Letzring, T. D., Wells, S. M., & Funder, D. C. (2006). Information quantity and quality affect the realistic accuracy of personality judgment. Journal of Personality and Social Psychology, 91(1), 111–123.
Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., et al. (2022). Holistic evaluation of language models. arXiv:2211.09110.
Matz, S. C., Kosinski, M., Nave, G., & Stillwell, D. J. (2017). Psychological targeting as an effective approach to digital mass persuasion. Proceedings of the National Academy of Sciences, 114, 12714–12719.
Matz, S. C., Teeny, J. D., Vaid, S. S., Peters, H., Harari, G. M., & Cerf, M. (2024). The potential of generative AI for personalized persuasion at scale. Scientific Reports, 14, 4692.
McCrae, R. R., & Costa, P. T., Jr. (1987). Validation of the five-factor model of personality across instruments and observers. Journal of Personality and Social Psychology, 52(1), 81–90.
Moore, D. A., & Healy, P. J. (2008). The trouble with overconfidence. Psychological Review, 115(2), 502–517.
Mori, M. (2012). The uncanny valley (K. F. MacDorman & N. Kageki, Trans.). IEEE Robotics & Automation Magazine, 19(2), 98–100. (Original work published 1970)
Nass, C., & Moon, Y. (2000). Machines and mindlessness: Social responses to computers. Journal of Social Issues, 56(1), 81–103.
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606.
Pace, V. L., & Brannick, M. T. (2010). How similar are personality scales of the "same" construct? A meta-analytic investigation. Personality and Individual Differences, 49, 669–676.
Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., Pettit, C., Olsson, C., Kundu, S., Kadavath, S., et al. (2022). Discovering language model behaviors with model-written evaluations. arXiv:2212.09251.
Peters, H., & Matz, S. C. (2024). Large language models can infer psychological dispositions of social media users. PNAS Nexus, 3(6), pgae231.
Raji, I. D., Denton, E., Bender, E. M., Hanna, A., & Paullada, A. (2021). AI and the everything in the whole wide world benchmark. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks.
Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 (General Data Protection Regulation). Official Journal of the European Union, L 119, 1–88.
Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 (Artificial Intelligence Act). Official Journal of the European Union.
Sackett, P. R., Zhang, C., Berry, C. M., & Lievens, F. (2022). Revisiting meta-analytic estimates of validity in personnel selection: Addressing systematic overcorrection for restriction of range. Journal of Applied Psychology, 107(11), 2040–2068.
Schönbrodt, F. D., & Perugini, M. (2013). At what sample size do correlations stabilize? Journal of Research in Personality, 47(5), 609–612.
Schwartz, H. A., Eichstaedt, J. C., Kern, M. L., Dziurzynski, L., Ramones, S. M., Agrawal, M., Shah, A., Kosinski, M., Stillwell, D., Seligman, M. E. P., & Ungar, L. H. (2013). Personality, gender, and age in the language of social media: The open-vocabulary approach. PLoS ONE, 8(9), e73791.
Serapio-García, G., Safdari, M., Crepy, C., Sun, L., Fitz, S., Romero, P., Abdulhai, M., Faust, A., & Matarić, M. (2023). Personality traits in large language models. arXiv:2307.00184.
Shannon, C. E. (1948). A mathematical theory of communication. Bell System Technical Journal, 27, 379–423, 623–656.
Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., Cheng, N., Durmus, E., Hatfield-Dodds, Z., Johnston, S. R., et al. (2023). Towards understanding sycophancy in language models. arXiv:2310.13548.
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366.
Soto, C. J., & John, O. P. (2017). The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of Personality and Social Psychology, 113(1), 117–143.
Strachan, J. W. A., Albergo, D., Borghini, G., Pansardi, O., Scaliti, E., Gupta, S., Saxena, K., Rufo, A., Panzeri, S., Manzi, G., Graziano, M. S. A., & Becchio, C. (2024). Testing theory of mind in large language models and humans. Nature Human Behaviour, 8(7), 1285–1295.
Thorndike, E. L. (1920). A constant error in psychological ratings. Journal of Applied Psychology, 4(1), 25–29.
Vazire, S. (2010). Who knows what about a person? The self–other knowledge asymmetry (SOKA) model. Journal of Personality and Social Psychology, 98(2), 281–300.
Weizenbaum, J. (1966). ELIZA — a computer program for the study of natural language communication between man and machine. Communications of the ACM, 9(1), 36–45.
Youyou, W., Kosinski, M., & Stillwell, D. (2015). Computer-based personality judgments are more accurate than those made by humans. Proceedings of the National Academy of Sciences, 112(4), 1036–1040.
Citation compliance note: all 47 references above were verified against live sources on August 16, 2026, in the session that produced this draft. No reference is cited from memory.
RELATED
- VIDEO
Scrubbing the Answer Key
A spoken walkthrough of contamination control and pre-registration discipline: four classes of leakage, auditable removal rules, and how findings stay confirmatory.
- VIDEO
The Fluency Demand Taxonomy: Scoring Work Tasks FD0–FD4
A spoken walkthrough of the second axis of occupational exposure: five ordinal levels scoring how much a task’s outcome depends on communication fitted to a specific person.
- VIDEO
Calibration as Co-Equal Criterion
A spoken walkthrough of why confidence–accuracy coupling belongs beside accuracy as a headline criterion, never composited into a single number.