← Back to blog
ARTICLE

From Digital Footprints to Living Dialogue: The Ecological Validity Problem in Computational Personality Inference

James Niblick, PsyD13 min read

From Digital Footprints to Living Dialogue: The Ecological Validity Problem in Computational Personality Inference

Abstract

For a decade, computational personality inference has been built on a substrate of convenience: social media likes, public posts, short elicited essays. That substrate produced foundational findings. It also has a defect that has gone largely unexamined — it is not the material from which deployed AI systems actually form impressions of their users.

Deployed systems see something else entirely: private, longitudinal, bidirectional dialogue accumulated across months or years. This article argues the gap between research stimulus and deployment stimulus is a first-order ecological validity problem rather than a matter of degree. Five properties separate conversational history from digital footprints — it is elicited by the judge rather than merely observed, bidirectional, longitudinal within a single relationship, private rather than audience-directed, and the actual deployment substrate. Each one breaks transfer on its own.

The article then confronts why the field avoided this material: it cannot be scraped. The consent architecture that fact forces is not a regrettable cost but an upgrade — scientific as well as ethical — and it resolves a standing tension in a literature whose most cited findings double as demonstrations of surveillance capability. A research agenda follows: what must be measured, what confounds the new substrate introduces, and what the field can no longer claim without it.

Keywords: ecological validity, computational personality inference, digital footprints, conversational data, informed consent, external validity


1. A Literature Built on What Was Available

The computational personality literature began the way empirical literatures usually do: with the data that could be obtained rather than the data that was wanted.

The foundational demonstrations used what a platform could export in bulk. Facebook Likes predicted a range of sensitive personal attributes, including personality traits, with accuracy that startled both researchers and the public (Kosinski, Stillwell, & Graepel, 2013). Open-vocabulary analysis of social media language recovered personality, age, and gender signals across tens of thousands of users, producing the field's most influential methodological template (Schwartz et al., 2013). And the capstone comparison — machine judgments from Likes outperforming the judgments of the target's own friends and family on self-report criteria — established that machines had entered the person-perception arena as serious judges (Youyou, Kosinski, & Stillwell, 2015). More recently, large language models have been shown to infer psychological dispositions directly from users' social media text without task-specific training (Peters & Matz, 2024).

These are genuine achievements, and nothing here disputes their internal validity. The question is the one ecological validity always asks: do these findings transfer to the setting we actually care about?

That setting is a person in sustained conversation with an AI system that is forming — and acting on — a model of who they are. It differs from the research substrate not in one respect but in five, and each difference is enough to break transfer by itself.

2. Five Properties That Do Not Transfer

Elicitation. A digital footprint is observed; a conversation is co-produced by the judge. When a system converses, it asks questions, offers interpretations, proposes options, and thereby shapes the very evidence on which it will later be scored. This is not a nuisance variable — it is a structural difference in the accuracy problem itself. The Realistic Accuracy Model treats the availability of trait-relevant information as a precondition the judge does not control (Funder, 1995, 2012); a conversational judge partially controls it. This creates a capability that no footprint study can measure — active cue elicitation — and a confound no footprint study must handle: a system's apparent accuracy may reflect skill at provoking diagnostic behavior rather than skill at reading it. Both belong in any serious evaluation, and neither is visible in scraped corpora.

Bidirectionality. Social media text is monologic; conversation is a joint activity in which meaning is established through mutual grounding, at a cost that varies with medium and partner (Clark & Brennan, 1991). What a person reveals in dialogue is a function of what the partner has said, and person-perception from dialogue is therefore perception of behavior that the perceiver's own conduct helped produce. The interpersonal-perception tradition has long modeled judgment as containing perceiver-specific components alongside target-driven ones (Kenny, 2004); with a conversational judge, the perceiver's contribution enters the evidence, not merely the interpretation. Findings from monologic corpora cannot speak to this.

Longitudinality within one relationship. Footprint datasets are typically cross-sectional: a snapshot of accumulated public output. Deployment is a relationship with a history, and the human literature is unambiguous that the information–accuracy relationship has structure over time — differential accuracy and self–other agreement increase with acquaintance while stereotype-based accuracy declines (Biesanz, West, & Millevoi, 2007), and information quantity and quality jointly govern judgmental accuracy (Letzring, Wells, & Funder, 2006). A cross-sectional corpus cannot distinguish a system that is getting better at this person from one that is getting better at people in general. That distinction is the whole question in deployment.

Privacy and audience design. Social media output is audience-directed by construction, and speakers systematically adjust style to their audience (Bell, 1984). A public post is a performance whose intended recipients are known and human. A conversation with an AI system is, subjectively, a private exchange — and often contains material the person has told no one: uncertainties, drafts of decisions, admissions framed as questions. The self–other knowledge asymmetry model establishes that observability is precisely what determines whose report is authoritative for which traits, with the self holding privileged access to internal states (Vazire, 2010). It follows that the two substrates differ not just in quantity but in which region of the trait space they illuminate: public footprints are rich in the observable and evaluative; private dialogue is rich in the internal. Accuracy estimates from one substrate should not be expected to hold on the other, and may well be strongest on different traits.

Deployment identity. The last property is the simplest and the most decisive. Conversational history is what deployed systems condition on. Everything else is a proxy. When a memory-enabled assistant forms an impression of its user, it does so from the accumulated record of exchanges with that user — not from their Likes, and not from an essay they wrote for a study.

A benchmark's stimulus should be the deployment stimulus. Where it is not, the benchmark measures a capability adjacent to the one in question.

3. What the Substrate Gap Costs

Three specific claims the field would like to make cannot currently be supported, and naming them precisely is more useful than a general appeal to ecological validity.

Transfer of accuracy estimates. Reported correlations between machine judgments and self-report on footprint data are the field's headline numbers, and they are routinely read as estimates of what deployed systems know about their users. They are not. They are estimates of what a model can infer from public, audience-directed, cross-sectional traces. Whether accuracy on private longitudinal dialogue is higher (more intimate evidence, more of it) or lower (task-dominated, less trait-relevant content) is an open empirical question with credible arguments on both sides — which is precisely why it must be measured rather than assumed.

Trait-level profiles of machine judgability. Because substrates differ in which regions of the trait space they illuminate, the pattern of which traits machines read well is likely substrate-dependent. A field that has measured this pattern only on public text has measured it once, in one place, and should not generalize the profile.

Claims about the effect of exposure. The most consequential deployment question — does the system's model of me improve as our history lengthens? — cannot be answered on cross-sectional data at all. It requires a corpus with an internal timeline, chronologically truncatable, from a single continuous relationship.

To this list, one methodological hazard should be added that is unique to the new substrate and must be handled before it can be used. A person's conversation with an AI system may contain discussion of their own assessment results, their trait self-descriptions, or the research study itself. Any such content converts inference into retrieval: the system reads the answer key rather than judging the person. Because generative systems are fluent regardless of evidentiary support (Ji et al., 2023), this contamination is not self-announcing in the output. It must be removed by protocol — a published scrubbing codebook, blind raters, audited scrub rates — and the removal must be specified in advance rather than justified afterward, in keeping with the pre-registration standards the broader methodological reform literature has established (Nosek, Ebersole, DeHaven, & Mellor, 2018; Simmons, Nelson, & Simonsohn, 2011).

4. The Reason the Field Avoided This Material

There is a straightforward explanation for a decade of footprints rather than dialogue. Footprints could be obtained at scale without asking anyone. Dialogue cannot.

Private conversation between a person and an AI system exists in one of two places — the person's account, or the provider's servers. The first requires the person's active participation to export. The second is not available to independent researchers, and its use without individual consent would be indefensible on grounds this literature has itself established. Read in the other direction, the field's most cited finding is a demonstration that intimate attributes can be extracted from behavioral residue people did not know they were producing (Kosinski et al., 2013) — a result its own authors framed with explicit concern about privacy.

A literature whose foundational results double as surveillance demonstrations should be alert to the difference between what can be taken and what has been given.

This is where the constraint becomes an asset rather than an obstacle. Consented donation of one's own conversational history — with the person as an active participant who exports the corpus, receives compensation, and retains the right to withdraw it — produces a research substrate whose provenance is unambiguous. And it aligns the science with a regulatory environment that is no longer optional: comprehensive data-protection law already governs the processing of personal data at this level of sensitivity, and dedicated AI regulation now sits alongside it (Regulation [EU] 2016/679; Regulation [EU] 2024/1689). Research designs that depend on unconsented collection are not merely ethically awkward; they are increasingly non-viable as a foundation for a durable standard.

There is also a scientific dividend, and it is larger than the compliance argument. Consented participation makes possible something scraped data can never supply: ground truth from the same person. A footprint study must take self-report where it can find it, if at all; a consented study can administer a full battery of validated instruments to the very people whose corpora it holds, with multiple independent instruments defining the criterion by convergence rather than by assumption (Soto & John, 2017; Lee & Ashton, 2004), and with honest acknowledgment that instruments agree with one another imperfectly (McCrae & Costa, 1987; Pace & Brannick, 2010). The participation consent requires is the same participation that produces a criterion worth scoring against. Ethics and measurement quality point the same way here. That is not always true, and it is worth saying plainly when it is.

5. What a Conversational-Corpus Research Program Must Do

Adopting the deployment substrate imposes obligations that footprint research did not face. Four are structural.

Characterize corpus content, not merely corpus size. Because accuracy depends on the quality of trait-relevant information and not only its quantity (Letzring et al., 2006), two participants with identical token counts may supply radically different evidence. Topic-level descriptive coding of corpora is therefore not an optional refinement but a requirement for interpreting any dose relationship, and unmeasured content heterogeneity is the most likely source of spurious findings in this design.

Design for the timeline. The substrate's central advantage is its internal chronology, and exploiting it means truncating corpora chronologically rather than sampling across them — the machine analog of relationship length. This permits within-person estimation of the information–accuracy function, free of the between-relationship selection confounds that complicate the human acquaintance literature (Biesanz et al., 2007), and it makes the compositional question testable: whether increasing exposure shifts judgment from the population-typical toward the person-specific.

Report the sample's boundary conditions honestly. A corpus large enough for high-exposure analysis exists only for heavy users of AI systems, who are not a random sample of humanity in age, occupation, disposition, or self-disclosure. This constrains generalization and should be stated as a limitation rather than buried; the within-subject design mitigates it for exposure effects specifically, since dose is manipulated within rather than between persons, but it does not resolve it for population-level accuracy estimates.

Treat the criterion with the care the criterion literature demands. Ground truth should be multi-instrument, reported per instrument, and referenced to the empirical agreement among instruments rather than to a fictional perfect standard — the discipline construct validation has required since its formalization (Cronbach & Meehl, 1955; Campbell & Fiske, 1959). Where instruments disagree about an individual on a trait, the criterion for that cell is undefined, and scoring should exclude it rather than charge the disagreement to the machine.

6. Conclusion: The Substrate Has Arrived

For most of the history of computational personality inference, the ideal stimulus did not exist in accessible form. Researchers used public traces because private, longitudinal, bidirectional dialogue between a person and a machine was either rare, short, or locked away. That is no longer true. Hundreds of millions of people now hold extended conversational histories with AI systems — the exact material on which those systems form their impressions — and those histories are exportable by the people who made them.

The field's response to this development will determine what its findings mean for the next decade. Continuing to infer deployment-relevant conclusions from scraped public text is a choice to keep measuring the proxy after the target has become available. Moving to the deployment substrate costs more, requires consent, demands new controls for contamination and content heterogeneity, and restricts the accessible sample. It also produces, for the first time, external validity for the claim the field's results are read as supporting: that AI systems can form accurate models of the individual people they talk to.

That claim is now testable on the material it is actually about.

The remaining question is whether it will be tested by people who ask permission first.


References

Bell, A. (1984). Language style as audience design. Language in Society, 13(2), 145–204.

Biesanz, J. C., West, S. G., & Millevoi, A. (2007). What do you learn about someone over time? The relationship between length of acquaintance and consensus and self–other agreement in judgments of personality. Journal of Personality and Social Psychology, 92(1), 119–135.

Campbell, D. T., & Fiske, D. W. (1959). Convergent and discriminant validation by the multitrait-multimethod matrix. Psychological Bulletin, 56(2), 81–105.

Clark, H. H., & Brennan, S. E. (1991). Grounding in communication. In L. B. Resnick, J. M. Levine, & S. D. Teasley (Eds.), Perspectives on socially shared cognition (pp. 127–149). American Psychological Association.

Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52(4), 281–302.

Funder, D. C. (1995). On the accuracy of personality judgment: A realistic approach. Psychological Review, 102(4), 652–670.

Funder, D. C. (2012). Accurate personality judgment. Current Directions in Psychological Science, 21(3), 177–182.

Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248, 1–38.

Kenny, D. A. (2004). PERSON: A general model of interpersonal perception. Personality and Social Psychology Review, 8(3), 265–280.

Kosinski, M., Stillwell, D., & Graepel, T. (2013). Private traits and attributes are predictable from digital records of human behavior. Proceedings of the National Academy of Sciences, 110, 5802–5805.

Lee, K., & Ashton, M. C. (2004). Psychometric properties of the HEXACO Personality Inventory. Multivariate Behavioral Research, 39(2), 329–358.

Letzring, T. D., Wells, S. M., & Funder, D. C. (2006). Information quantity and quality affect the realistic accuracy of personality judgment. Journal of Personality and Social Psychology, 91(1), 111–123.

McCrae, R. R., & Costa, P. T., Jr. (1987). Validation of the five-factor model of personality across instruments and observers. Journal of Personality and Social Psychology, 52(1), 81–90.

Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606.

Pace, V. L., & Brannick, M. T. (2010). How similar are personality scales of the "same" construct? A meta-analytic investigation. Personality and Individual Differences, 49, 669–676.

Peters, H., & Matz, S. C. (2024). Large language models can infer psychological dispositions of social media users. PNAS Nexus, 3(6), pgae231.

Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 (General Data Protection Regulation). Official Journal of the European Union, L 119, 1–88.

Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 (Artificial Intelligence Act). Official Journal of the European Union.

Schwartz, H. A., Eichstaedt, J. C., Kern, M. L., Dziurzynski, L., Ramones, S. M., Agrawal, M., Shah, A., Kosinski, M., Stillwell, D., Seligman, M. E. P., & Ungar, L. H. (2013). Personality, gender, and age in the language of social media: The open-vocabulary approach. PLoS ONE, 8(9), e73791.

Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366.

Soto, C. J., & John, O. P. (2017). The next Big Five Inventory (BFI-2): Developing and assessing a hierarchical model with 15 facets to enhance bandwidth, fidelity, and predictive power. Journal of Personality and Social Psychology, 113(1), 117–143.

Vazire, S. (2010). Who knows what about a person? The self–other knowledge asymmetry (SOKA) model. Journal of Personality and Social Psychology, 98(2), 281–300.

Youyou, W., Kosinski, M., & Stillwell, D. (2015). Computer-based personality judgments are more accurate than those made by humans. Proceedings of the National Academy of Sciences, 112(4), 1036–1040.


Citation compliance note: all 23 references above were verified against live sources on August 16, 2026, in the session that produced this draft series. No reference is cited from memory.