Where Fluency Pays: Communicative Load as the Moderator of Personalization Benefit
Abstract
Personalization is sold as a general good. Know the user better, serve them better — everywhere, on everything, without limit.
This article predicts that the claim is false across most of the tasks people actually bring to AI systems, and specifies the experiment that would establish it. The prediction follows from a simple moderation hypothesis: accuracy about a person converts into benefit only in proportion to how much of a task’s outcome rides on communication being fitted to that person. Where communicative load is near zero, an accurate model of the user should produce no measurable improvement over a person-blind baseline, no matter how accurate it is. Where communicative load is high, the same accuracy should produce substantial improvement — and an inaccurate model should produce active harm, worst precisely where accuracy would have helped most.
The existing literature is consistent with the prediction but cannot test it, because the studies that found strong personalization effects were run almost exclusively on high-load tasks, and the null results at low load are largely unpublished. The decisive design is a factorial experiment crossing three profile conditions — true, absent, and deliberately inverted — against a task battery independently scored for communicative load, with outcomes graded blind and felt-understanding collected separately. The predicted signature is a crossover interaction that no additive account of personalization can produce.
The article also states what a confirmation would cost the industry that sells personalization, and why the experiment is nonetheless worth running by parties who would prefer a different answer.
Keywords: personalization, communicative load, moderation, human–AI interaction, psychological targeting, experimental design
1. The Claim Nobody Has Tested
Every major AI product now personalizes, and the implicit theory behind that investment is monotone: a better model of the user yields better outcomes, on any task, without an upper bound and without exception.
Notice what would be required for this to be true. It would have to be the case that knowing a user’s disposition improves the conversion of a file, the retrieval of a phone number, the summation of a column, the resolution of an arithmetic question. On reflection nobody believes this. But the belief that personalization has limits has never been converted into a statement about where the limits fall, which means product decisions, capability claims, and research designs all proceed as though the limits do not exist.
The prediction developed here is specific: benefit from person-accuracy is proportional to communicative load. Formally it follows from a multiplicative account developed in a companion article; stated without notation, it says that accuracy has nothing to multiply against on tasks whose outcomes do not ride on communication being fitted to a particular person.
Three consequences distinguish this from the truism that context matters.
Low-load tasks should show nulls regardless of accuracy. Not small effects — nulls. A perfectly accurate person-model on a file conversion should produce no measurable benefit over no model at all.
Inaccurate models should produce harm, not merely absent benefit. A system whose model of the person is wrong in a specific direction — treating the blunt person as needing cushioning, the autonomous person as wanting direction — should perform worse than a system with no model at all, and worst on the tasks where a correct model would have mattered most.
The relationship should appear as an interaction, not a main effect. Studies estimating an average effect of personalization across mixed task sets will produce exactly the heterogeneous, hard-to-replicate literature the field currently has.
2. What the Existing Evidence Can and Cannot Show
Where the evidence sits. The strongest demonstrations of person-matched communication come from psychological targeting: messages matched to recipients’ psychological characteristics outperform mismatched or generic messages at scale (Matz, Kosinski, Nave, & Stillwell, 2017), and large language models generate such tailored content cheaply, producing measurable effects even from coarse targeting cues (Matz, Teeny, Vaid, Peters, Harari, & Cerf, 2024). Adjacent work in the study of language establishes that speakers routinely adapt style to audience as a basic communicative competence (Bell, 1984).
Why this is consistent with the prediction rather than a test of it. Persuasion is a high-load task by construction: its outcome is almost entirely constituted by whether the message reached a particular mind. The prediction says effects should be large there. Finding them confirms nothing about the low-load half of the claim, because the low-load half has not been studied.
Why the low-load evidence is missing. The null result at low load is the least publishable finding in this area. A study reporting that personalization did not improve performance on routine tasks faces an unenthusiastic reception from journals and an actively hostile one from the commercial parties best positioned to run it. The absence of published nulls is therefore weak evidence of anything, and the general mechanism by which such absences distort literatures is well documented (Simmons, Nelson, & Simonsohn, 2011).
What is entirely absent. No published work examines inverted person-models — systems given a systematically wrong model of the user. The condition is easy to construct and nobody constructs it, because vendors have no reason to and academics have not had the ground-truth profiles required to invert.
This is why the prediction requires a purpose-built experiment rather than a meta-analysis. The literature cannot answer the question because it has sampled one corner of the design space.
3. The Decisive Design
Structure. A factorial experiment: three profile conditions crossed against a task battery spanning the load range, within participants.
Profile conditions.
True profile. The system operates with an accurate model of the participant, established from validated instruments administered before the study.
No profile (baseline). The system operates person-blind. This is the comparison against which benefit is defined.
Inverted profile. The system operates with a systematically wrong model, constructed by reflecting the participant’s true profile through the scale midpoints. This condition is what makes the design decisive, since it tests the harm prediction that no observational study can reach.
Task battery. Tasks scored independently for communicative load, using the ordinal taxonomy developed in a companion article, spanning from tasks whose outcomes are recipient-independent (format transformation, defined retrieval, computation) through category-fitted tasks (technical explanation to a stated audience) to individually fitted tasks (advisory work, coaching, complex decision support). Load scoring is performed by raters blind to the study’s hypotheses, with agreement reported.
Outcome measurement. Task outcomes graded by raters blind to condition, against criteria specified in advance for each task. Felt understanding — the participant’s sense of being understood — is collected separately and is not an outcome. It is a distinct dependent variable used to test the dissociation described in Section 4.
Controls. Task order randomized; profile condition order counterbalanced; identical system, prompts, and decoding parameters across conditions with only the profile varying; participants blind to condition.
Ethics of the inverted condition. Participants consent to receiving responses generated from an intentionally incorrect model of them, are debriefed, and receive the accurate condition’s outputs afterward. The inverted condition is restricted to tasks where a poor outcome carries no consequence beyond the study, which excludes the highest-load real-world tasks and is a real limitation on the design’s reach rather than a formality.
4. Pre-Registered Hypotheses
Stated before data collection, per the pre-registration standards this program adopts (Nosek, Ebersole, DeHaven, & Mellor, 2018).
M1 — The interaction. The effect of profile condition on task outcome is moderated by communicative load. This is the primary hypothesis; the interaction term, not the main effect of personalization, carries the prediction.
M2 — The null zone. At the lowest load levels, outcomes in the true-profile condition do not differ meaningfully from baseline. This is a predicted null, tested with equivalence procedures rather than by failing to reject, since “no difference” is a claim requiring its own evidence.
M3 — The harm zone. At high load, the inverted-profile condition produces outcomes below baseline, and the magnitude of the deficit increases with load. This is the design’s most distinctive prediction and the one with the clearest practical implication: bad person-modeling is not merely wasted effort.
M4 — The crossover signature. Condition differences converge at low load and diverge at high load, with the inverted condition falling below baseline where the spread is widest. No additive account of personalization benefit produces this pattern.
M5 — Felt understanding dissociates from benefit. Participant-reported feelings of being understood will be less sensitive to profile condition than task outcomes are, and at low load may show condition differences where outcomes show none. If confirmed, the industry’s principal proxy for personalization success is measuring something other than benefit — which would be the finding with the largest immediate consequences for how these products are evaluated.
M6 (exploratory). Accuracy dose: where profiles of varying accuracy can be constructed, benefit at high load should scale with profile accuracy rather than being binary. Exploratory because constructing graded-accuracy profiles is methodologically delicate.
5. Results
[Reserved for the version of record: outcome models by condition and load level; equivalence tests at low load; harm estimates at high load with confidence intervals; the felt-understanding dissociation; and per-task descriptive results.]
6. What Confirmation Would Cost, and Who Should Run It
A confirmed result is uncomfortable for the industry that would cite it, and stating that plainly is part of the argument for the experiment’s credibility.
For vendors, M2 means a large share of deployed personalization produces no measurable benefit. Personalization applied to low-load tasks would be, on this evidence, decoration — and possibly worse than decoration, given the privacy cost of collecting person-data that improves nothing.
For products, M3 means person-modeling carries downside risk that no current product measures. A system with an inaccurate model of a user is not a neutral system; it is one whose errors concentrate in the interactions that matter most.
For evaluation practice, M5 would invalidate satisfaction and felt-understanding metrics as evidence of personalization benefit. Those are the metrics the industry currently reports.
The author’s own commercial interest runs against all three findings, which is precisely why the design is specified in enough detail for independent replication and why the profile conditions are constructed from validated instruments rather than proprietary ones. A prediction that would embarrass its proponent if confirmed, tested by parties with no stake, is the only form in which this claim should enter the literature.
There is also a positive case that survives confirmation. If M1 and M3 hold, the value of accurate person-modeling is not diminished but relocated: it becomes concentrated, measurable, and defensible on the tasks where it demonstrably pays, and honest about the tasks where it does not. A field able to say where its capability matters is in a stronger position than one asserting it matters everywhere.
7. Limits
Task batteries are not workplaces. Laboratory tasks scored for load approximate but do not reproduce the stakes, duration, and relational history of real high-load work. The highest-load tasks in life — crisis conversations, serious diagnoses, conflict under real consequences — cannot ethically be studied in the inverted condition, which means the harm prediction is testable only in its milder range and must be extrapolated with explicit caution.
Benefit is task outcome. The design measures whether outcomes improve, not whether people are better off in ways outcomes miss. Relational goods, long-run trust, and the value of being known are outside the measurement and should not be claimed or denied on its basis.
Profile accuracy is bounded by the criterion. Profiles come from validated instruments, whose own agreement is imperfect, and the “true” condition is therefore true-to-instrument rather than true-to-person — the same ceiling constraint that governs accuracy measurement throughout this program.
Single-session estimation. Effects at high load may require relationship duration that a session cannot supply, which would bias the design toward finding smaller high-load effects than deployment would produce, making M1’s confirmation conservative and its disconfirmation ambiguous.
8. Conclusion
The industry sells personalization as universal. The claim has never been tested against tasks where it should fail, and the condition that would reveal its failure mode — a wrong model, held confidently — has never been run.
The prediction here is that accuracy about a person pays in proportion to how much the task rides on reaching that person, and that a wrong model on a high-load task is worse than no model at all. The design that tests it is straightforward, the hypotheses are stated in advance, the predicted null is treated as a claim requiring evidence rather than a failure to find one, and the outcome that would most damage the author’s own commercial position is the one the design is built to detect.
That is the arrangement under which a personalization claim should be evaluated. It has not been the arrangement so far.
References
Bell, A. (1984). Language style as audience design. Language in Society, 13(2), 145–204.
Matz, S. C., Kosinski, M., Nave, G., & Stillwell, D. J. (2017). Psychological targeting as an effective approach to digital mass persuasion. Proceedings of the National Academy of Sciences, 114, 12714–12719.
Matz, S. C., Teeny, J. D., Vaid, S. S., Peters, H., Harari, G. M., & Cerf, M. (2024). The potential of generative AI for personalized persuasion at scale. Scientific Reports, 14, 4692.
Nosek, B. A., Ebersole, C. R., DeHaven, A. C., & Mellor, D. T. (2018). The preregistration revolution. Proceedings of the National Academy of Sciences, 115(11), 2600–2606.
Simmons, J. P., Nelson, L. D., & Simonsohn, U. (2011). False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant. Psychological Science, 22(11), 1359–1366.
Citation compliance note: all 5 references above were verified against live sources on August 16, 2026, in the session that produced this draft series. No reference is cited from memory.
RELATED
- VIDEO
Scrubbing the Answer Key
A spoken walkthrough of contamination control and pre-registration discipline: four classes of leakage, auditable removal rules, and how findings stay confirmatory.
- VIDEO
The Fluency Demand Taxonomy: Scoring Work Tasks FD0–FD4
A spoken walkthrough of the second axis of occupational exposure: five ordinal levels scoring how much a task’s outcome depends on communication fitted to a specific person.
- VIDEO
Calibration as Co-Equal Criterion
A spoken walkthrough of why confidence–accuracy coupling belongs beside accuracy as a headline criterion, never composited into a single number.