← Back to blog
ARTICLE

Dose–Response Person-Perception: Acquaintance Effects in Machines, With the Token as the Unit of Exposure

James Niblick, PsyD12 min read

Abstract

How long does a machine have to know you before it knows you?

The question has an exact human analog with a century of research behind it, and a machine version that has never been asked properly. This article specifies the machine version. Exposure is quantified in tokens rather than hours, because tokens are what models consume and because time confounds interaction volume with typing speed. A participant's real conversational history with an AI system is truncated chronologically at standardized thresholds—the first ten thousand tokens, the first fifty thousand, and upward on a log scale—and every truncation is judged by every model under study. Each participant is therefore their own control across exposure levels, which eliminates the confound that has complicated the human acquaintance literature since its inception: people who have known each other longer are systematically different from people who have not.

The central hypothesis is not that accuracy rises with exposure. It is that accuracy changes composition—stereotype-based accuracy declining, differential accuracy rising—as the judge shifts from knowing what people in general are like to knowing what this person is like. A system whose aggregate accuracy improves without that compositional shift is getting better at the average user, which is the precise failure mode a personalization industry would ship without noticing. Four additional hypotheses concern the floor condition as a fabrication detector, trait-specific dose curves, the quantity–quality confound, and the open question of whether machine curves plateau where human curves do.

Keywords: acquaintanceship, person perception, dose–response, exposure, large language models, within-subject design


1. The Question Nobody Has Asked Correctly

An assistant that has exchanged three messages with you knows almost nothing about you. An assistant that has worked alongside you for two years knows—something. Between those points lies a function, and nobody has drawn it.

The function matters commercially and scientifically. Commercially, because every memory feature shipped in the last three years is an implicit bet on its shape: that accumulated history yields better-fitted responses, monotonically, without limit. Scientifically, because the function's shape is the single most informative thing one could learn about machine person-perception. A curve that rises steeply and plateaus early implies a capability that saturates on thin evidence. A curve that keeps climbing implies something that genuinely accumulates. A curve that rises in aggregate while its composition stays fixed implies a system getting better at people in general and no better at you.

The human literature has drawn versions of this curve for decades, and its results are more surprising than intuition predicts. Judgments from very brief behavioral samples achieve non-trivial accuracy, and the meta-analytic record found that longer observation did not reliably improve prediction (Ambady & Rosenthal, 1992). Trait inferences at zero acquaintance carry genuine validity from minimal cues (Borkenau & Liebler, 1992). Accuracy depends on the quality of trait-relevant information and not merely its quantity (Letzring, Wells, & Funder, 2006). And the componential literature has found that consensus among judges rises considerably less with acquaintance than folk psychology expects (Kenny, 2004).

Most importantly, the acquaintance literature has established that what improves over time is specific: self–other agreement and differential accuracy increase with acquaintance while stereotype-based accuracy decreases (Biesanz, West, & Millevoi, 2007). Longer relationships do not simply produce more accuracy. They produce a different kind of accuracy.

That finding, transplanted, generates the hypothesis this article is built around.

2. Why Machines Make This Measurable

Human acquaintance research faces a structural obstacle that machine research does not.

In human studies, acquaintance length is a between-relationships variable. Comparing judges who have known targets for one month against judges who have known targets for five years compares different people, in different relationships, selected non-randomly: relationships that survive five years differ systematically from those that do not, and people who form them differ too. Every estimate of the acquaintance effect carries that selection confound (Biesanz et al., 2007).

A machine judge's exposure is manipulable within the same relationship. One participant's conversational history can be truncated at ten thousand tokens, fifty thousand, two hundred fifty thousand, and a million—and each truncation judged separately, with the target held constant, the criterion held constant, and the judge held constant. Differences in accuracy across exposure levels cannot be attributed to differences between persons, because there are no differences between persons.

This is a design affordance the human literature has never had, and it makes the machine case the cleaner test of a question the human case has been asking for fifty years. Two further affordances follow. The judge is replicable: the same model can be run repeatedly on identical evidence, turning judge reliability into a measured quantity rather than an inference (Kenny, 2004). And exposure is continuous: the dose ladder can be as fine-grained as compute allows, permitting estimation of curve shape rather than comparison of two or three arbitrary bins.

3. Design

Exposure unit. Tokens, not hours or sessions. Tokens are what models actually consume, and temporal measures confound interaction volume with typing speed, session length, and conversational style.

Truncation rule. Strictly chronological. The D1 corpus is the first ten thousand tokens of the participant's history, not a random sample of ten thousand tokens drawn from across it. Chronological truncation reproduces the way an impression forms along a relationship's timeline and preserves the ecological meaning of the dose curve. Random sampling would measure something else entirely—accuracy from a representative slice, which no deployed system ever sees.

Dose ladder. A logarithmic progression: a floor condition below ten thousand tokens, then ten thousand, fifty thousand, two hundred fifty thousand, and one million, with breakpoints subject to revision during piloting based on the corpus sizes the recruitment pool can actually supply. Logarithmic spacing is chosen because the human literature's returns to information are diminishing, and linear spacing would waste measurements in the flat region.

Cells. Every dose level of every participant's corpus is judged by every model, with multiple stochastic repetitions per cell. A modest participant sample multiplies into a large number of assessment cells; the marginal cost of cells is computational rather than human, which is what makes the within-subject design efficient as well as clean.

Scoring. Each cell yields a trait rating vector and a confidence rating per trait, scored against multi-instrument criterion consensus, expressed as a proportion of the empirically observed inter-instrument ceiling for that trait, per the scoring framework specified in a companion article. Accuracy is additionally decomposed into stereotype and differential components (Section 4).

Contamination control. Corpora are scrubbed of any content in which participants discuss their own assessment results or trait self-descriptions, since such content converts inference into retrieval and fluent output does not announce its provenance (Ji et al., 2023). Scrubbing follows a published codebook applied by raters blind to instrument scores, with audited scrub rates.

Analysis. Within-subject accuracy curves across dose, with inflection and plateau estimation; mixed models with participants as random effects; per-trait and per-model curves reported separately, never pooled into a single aggregate curve.

4. Hypotheses

All hypotheses are stated before any model output has been observed, per the pre-registration standards this program adopts (Nosek, Ebersole, DeHaven, & Mellor, 2018; Simmons, Nelson, & Simonsohn, 2011).

H1—Compositional shift (primary). As exposure increases, the composition of accuracy shifts: stereotype-based accuracy declines and differential accuracy rises, following the pattern established in human acquaintance research (Biesanz et al., 2007). This is the primary hypothesis, and it is deliberately more demanding than an aggregate-accuracy prediction.

H1a—The null that matters. Aggregate accuracy may rise across doses without a compositional shift. This pattern would indicate a system converging on the population-typical profile—getting better at the average person, not at the individual. H1a is stated as a distinct, nameable outcome rather than as a failure of H1, because it is the outcome with the greatest practical significance for deployed personalization, and because a design that could not distinguish it from H1 would be uninformative regardless of what it found.

H2—Floor behavior as fabrication detection. At the floor condition, accuracy should be modest but above chance, since genuine signal exists at minimal exposure (Borkenau & Liebler, 1992), and confidence should be correspondingly low. Two failure signatures are pre-specified: elaborate, confident profiles at the floor (fabricated familiarity), and confidence that does not vary across the dose ladder (calibration decoupling). Both are reportable results in their own right, not merely diagnostics.

H3—Trait-specific curves. Dose curves will differ by trait, with traits that are more legible in verbal self-narration reaching their plateaus earlier and higher than traits that depend on observable behavior. This follows from the trait-visibility structure of person perception and from the asymmetry of what different perspectives can know (Vazire, 2010; Connelly & Ones, 2010). Pooling traits into a single curve is therefore prohibited in the analysis plan.

H4—Quality moderates quantity. Accuracy at matched token counts will vary systematically with corpus content—specifically, with the density of trait-relevant self-disclosure—consistent with the finding that both quantity and quality of information govern accuracy (Letzring et al., 2006). Corpora are therefore descriptively coded, and content is entered as a covariate. Without this control, the dose curve silently confounds exposure with disclosure style.

H5—Plateau location (exploratory). Whether machine curves plateau where human curves do is left open. The human literature suggests early saturation (Ambady & Rosenthal, 1992; Kenny, 2004), plausibly because human judges are constrained by attention and memory in ways machines are not. A machine judge with perfect recall of the full record may continue improving past the human plateau. This is stated as exploratory because either outcome is theoretically interesting and neither is predicted with confidence; if machine curves keep rising where human curves flatten, that is the strongest available evidence that machine person-perception is a distinct capability rather than a mechanical approximation of the human one.

5. Results

[Reserved. Results from the establishing study will be reported here: per-trait, per-model dose curves with confidence intervals; stereotype/differential decomposition at each dose; floor-condition accuracy and confidence distributions; content-covariate models; and inflection and plateau estimates where estimable. No results are reported in this version.]

6. Anticipated Interpretive Constraints

Four constraints will apply to whatever the curves show, and stating them in advance prevents their being negotiated after the fact.

Sample conditioning at the upper doses. Corpora large enough to populate the highest exposure levels exist only for heavy users of AI systems, who are not representative in age, occupation, disposition, or self-disclosure. The within-subject design protects exposure estimates from this selection, because dose is manipulated within persons; it does not protect population-level accuracy estimates, which must be reported as conditioned on the sample.

Curve shape depends on the criterion. Because accuracy is scored against instrument consensus and expressed against the inter-instrument ceiling, and because ceilings differ by trait (McCrae & Costa, 1987; Pace & Brannick, 2010), curves are not comparable across traits in raw units. Only ceiling-referenced curves are comparable, and even then the comparison is of proportions attained, not of absolute knowledge.

Estimates require adequate n. Every point on every curve is a correlation across persons, and correlations do not stabilize within useful bounds at small samples (Schönbrodt & Perugini, 2013). Curves from underpowered samples will be reported with intervals wide enough to make their instability visible rather than smoothed into a suggestive shape.

Version-stamping. Dose curves are properties of specific model versions under specific conditions. The protocol is designed for repeatable administration against future model generations, which is the property that makes this a standing measurement rather than a snapshot.

7. Why This Design Is Worth Running

Three payoffs justify the cost, and they accrue to different audiences.

For deployment, the curve answers a question every product team currently answers by assumption: how much history does a system need before its model of the user is worth acting on? If the plateau is early, memory architectures are over-engineered for person-perception purposes and the marginal token is being stored for nothing. If it is late, the systems currently shipping are operating well below their own ceiling. Either answer changes engineering priorities, and neither is currently known.

For personality science, the design supplies something the human literature has never had: an acquaintance function estimated free of between-relationship selection, with a replicable judge, at arbitrary resolution. Questions about the information–accuracy relationship that have been argued from confounded designs for fifty years become directly estimable—in a judge that is not human, which is a limitation, but under conditions of control that no human study can approach.

For measurement policy, the floor condition produces the single most consequential number in this research program: how confidently systems characterize people they have barely observed. That figure bears directly on whether unmeasured person-perception claims should be permitted to circulate, and it is obtainable only in a design where the floor is deliberately included rather than treated as an uninteresting baseline.

8. Conclusion

The question is old: how well can one party know another, and how much acquaintance does knowing require? Psychology has been asking it since before the instruments existed to answer it properly, and has been answering it with designs that confound acquaintance with everything acquaintance travels with.

The machine case removes the confound. It also raises the stakes, because the judges in question are being deployed at population scale with no measurement of what their exposure buys them.

The hypotheses above are stated in advance, the design is specified to be replicated, and the outcome that would most embarrass the industry—aggregate accuracy improving while the composition never shifts, systems learning the average user and calling it personalization—is named as a distinct, testable result rather than filed as a failure to confirm. That is the arrangement under which the curve should be drawn.


References

Ambady, N., & Rosenthal, R. (1992). Thin slices of expressive behavior as predictors of interpersonal consequences: A meta-analysis. Psychological Bulletin, 111(2), 256–274.

Biesanz, J. C., West, S. G., & Millevoi, A. (2007). What do you learn about someone over time? The relationship between length of acquaintance and consensus and self–other agreement in judgments of personality. Journal of Personality and Social Psychology, 92(1), 119–135.

Borkenau, P., & Liebler, A. (1992). Trait inferences: Sources of validity at zero acquaintance. Journal of Personality and Social Psychology, 62, 645–657.

Connelly, B. S., & Ones, D. S. (2010). An other perspective on personality: Meta-analytic integration of observers' accuracy and predictive validity. Psychological Bulletin, 136(6), 1092–1122.

Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248, 1–38.

Kenny, D. A. (2004). PERSON: A general model of interpersonal perception. Personality and Social Psychology Review, 8(3), 265–280.

Letzring, T. D., Wells, S. M., & Funder, D. C. (2006). Information quantity and quality affect the realistic accuracy of personality judgment. Journal of Personality and Social Psychology, 91(1), 111–123.

McCrae, R. R., & Costa, P. T., Jr. (1987). Validation of the five-factor model of personality across instruments and observers. Journal of Personality and Social Psychology, 52(1), 81–90.

Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606.

Pace, V. L., & Brannick, M. T. (2010). How similar are personality scales of the "same" construct? A meta-analytic investigation. Personality and Individual Differences, 49, 669–676.

Schönbrodt, F. D., & Perugini, M. (2013). At what sample size do correlations stabilize? Journal of Research in Personality, 47(5), 609–612.

Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366.

Vazire, S. (2010). Who knows what about a person? The self–other knowledge asymmetry (SOKA) model. Journal of Personality and Social Psychology, 98(2), 281–300.


Citation compliance note: all 13 references above were verified against live sources on August 16, 2026, in the session that produced this draft series. No reference is cited from memory.