← Back to blog
ARTICLE

Consented Ground Truth: A Data Architecture Where the Profile Belongs to the Person

James Niblick, PsyD10 min read

Abstract

There is a line running through this entire field, and almost nothing in the literature draws it.

On one side: a person decides to be assessed, receives the result, and controls where it goes. On the other: an institution infers a psychological profile of someone who never asked for one, and acts on it. Both produce a trait profile. They are not variants of the same practice, and the difference is not a matter of data-handling hygiene.

This article specifies the architecture that operationalizes the first and forecloses the second, and argues that the choice is scientific as well as ethical. The argument’s uncomfortable premise is that the field’s most cited findings are demonstrations of the second practice: intimate attributes are recoverable from behavioral residue people did not know they were producing. A literature whose foundational results double as surveillance proofs cannot treat consent as a compliance formality appended to its methods sections.

Six architectural commitments follow, each stated as a design constraint rather than a principle: the profile belongs to the person; consent is layered, purpose-bound, and revocable; scopes govern uses rather than access; subjects can see how their profiles have been consulted; hard limits exist that no party can lift; and profiles are portable across institutions. The article then makes the scientific case — that consented architectures yield better data, not merely permissible data — and closes with what remains unresolved, including the two problems the architecture does not solve: consent given under employment pressure, and inference about people who never participated at all.

Keywords: informed consent, data architecture, psychological profiling, research ethics, data protection, person perception


1. The Line

Two scenarios, deliberately close together.

In the first, a person completes a validated personality assessment, receives their results, and authorizes an AI system to use the profile so that its communication fits them better. They can see the profile, change the authorization, and withdraw it.

In the second, an organization runs its members’ conversation logs through a model that infers psychological profiles. The subjects are not asked, are not told, and never see the output. The profiles inform decisions about them.

Both produce trait profiles from data. The technical pipelines might be nearly identical. But the first is assessment and the second is surveillance, and no amount of encryption, access control, or retention policy converts the second into the first.

The distinction is not that one involves consent paperwork. It is directionality: whether the assessment originates with the person being assessed or is performed upon them. A research and product architecture can encode that distinction structurally, or it can leave it to the discretion of whoever holds the data. This article argues for the former and specifies what it requires.

2. Why This Field Owes the Question More Than Most

Computational personality inference has an uncomfortable relationship with its own foundational results.

The finding that launched the modern literature is that a range of sensitive personal attributes — including traits people would not volunteer — are predictable from ordinary digital behavior, a result its authors published with explicit concern about the privacy consequences (Kosinski, Stillwell, & Graepel, 2013). Language-based assessment then demonstrated the same capability at scale from public text (Schwartz et al., 2013). Machine judgments from behavioral residue were shown to exceed the accuracy of the target’s own friends and family (Youyou, Kosinski, & Stillwell, 2015). Most recently, large language models have been shown to infer psychological dispositions from user text with no task-specific training at all (Peters & Matz, 2024).

Read one way, these are advances in measurement. Read the other way — the way their own authors have often read them — they are demonstrations that unconsented psychological profiling is technically trivial and getting easier.

Two facts follow. The capability is not contingent on anyone’s permission; it exists and will be exercised. And the field cannot therefore treat consent as a formality, because its own results establish that the alternative to consented profiling is not the absence of profiling but its invisible form.

3. Six Architectural Commitments

Each is stated as a constraint on system design rather than a value, because values that are not encoded in architecture are renegotiated under pressure.

1. The profile belongs to the person. The subject holds the authoritative copy, can view it in full, and can withdraw it. Institutions consult profiles; they do not own them. This inverts the prevailing arrangement, in which the assessing organization holds the file and the subject may, at best, request a summary.

2. Consent is layered, purpose-bound, and revocable. Consent to be assessed is not consent to every use of the result. Each purpose is separately explained and separately authorized, in language a non-specialist can act on, and every authorization can be withdrawn without penalty. A single blanket authorization at onboarding is the mechanism by which the first scenario in Section 1 becomes the second.

3. Scopes govern uses, not merely access. Access controls answer who can see this. Scopes answer what may be done with it, which is the question that matters when the holder of the data is an institution with authority over the subject. The same profile may be legitimately consulted to improve how a system explains something and illegitimately consulted to decide whether someone keeps their job.

4. Consultation is visible to the subject. The subject can see that their profile was consulted, when, and under which purpose category. This is the commitment most often omitted, because it is the one that transfers real oversight rather than the appearance of it. It also has a scientific benefit: subjects who can audit use have less reason to manage their self-presentation during assessment.

5. Some limits are not configurable. A category of uses is foreclosed by architecture rather than policy, and cannot be enabled by any party — including the system’s operator, the institution, and the subject. Consent cannot license everything, because consent is not always free (Section 6), and because a system whose limits are all configurable has no limits.

6. Profiles are portable. The person carries the profile between institutions. Portability is what makes the ownership claim in commitment 1 real rather than nominal, and it also disciplines institutions: a profile that leaves with the person cannot become an asset held against them.

For research specifically, three further requirements follow from the same logic: pseudonymization at intake with linkage keys held under independent access control; provider arrangements that exclude the use of submitted data for training; and publication practices that never reproduce raw material, since a corpus of private conversation cannot be de-identified by excerpt selection.

4. Consent Improves the Science

The ethical case is familiar. The scientific case is less often made and is, for this program, decisive.

It produces ground truth. Unconsented research must take its criterion where it can find it — a self-report scraped from somewhere, or none. Consented participation permits administering a full battery of validated instruments to the very people whose data is being analyzed, which is the only way to define a criterion by convergence across independent instruments rather than by assumption. The participation consent requires is the same participation that produces a criterion worth scoring against.

It permits the longitudinal, private substrate. The material on which deployed systems actually form impressions — extended private dialogue — exists only in people’s own accounts. It cannot be scraped. Either people donate it or the field continues studying proxies.

It reduces distortion. Assessment under coercion or surveillance invites impression management. A subject who owns the result, controls its uses, and can see its consultations has less incentive to game the instrument — which is a measurement-quality argument, not a comfort argument.

It survives replication and regulation. A finding built on unconsented collection is fragile: it may become unrepeatable as legal and platform conditions change. Comprehensive data-protection law already governs processing of this kind, and dedicated AI regulation now sits alongside it (Regulation [EU] 2016/679; Regulation [EU] 2024/1689). Architecture built to the higher standard remains runnable.

5. The Same Architecture in Deployment

A research ethic that stops at the study boundary is not an ethic; it is a protocol. The six commitments transfer to deployed systems with one addition and one warning.

The addition is that scope enforcement must live below the layer that can be reconfigured by the deploying institution. A limit implemented as an instruction to a model is a limit that erodes under sustained pressure from the party paying for the system; a limit implemented in the connective infrastructure holds. This is the deployment analog of pre-registration: the point of both is to remove a decision from the party with an interest in it, in advance.

The warning concerns the workplace, where most institutional deployment occurs, and where algorithmic systems have already been documented restructuring the terrain of managerial control — expanding the reach of direction, evaluation, and discipline in ways that outrun existing oversight (Kellogg, Valentine, & Christin, 2020). Psychological profiling deployed into that context does not arrive as a neutral capability. It arrives as an addition to an existing asymmetry, which is why commitments 3, 4, and 5 — scopes, visibility, and non-configurable limits — do the heaviest work there.

6. What This Does Not Solve

Two problems remain, and an architecture paper that concealed them would be advertising.

Consent under employment pressure. When assessment is requested by an employer, the voluntariness of consent is compromised in ways no interface fixes. Declining may be visible; declining may be inferred; declining may be costly. Data-protection regimes recognize the general problem in employment relationships, and it is not solved by better disclosure language. Partial mitigations exist — making declining genuinely costless and non-visible, foreclosing the highest-stakes uses by architecture rather than by choice, and separating the assessing party from the employing one — but they are mitigations. The honest position is that employment-context consent is weaker than research consent and should carry correspondingly narrower permissible uses.

People who never participate. The architecture governs profiles of consenting subjects. It says nothing about a system that infers a profile of a non-participant from ambient text, which the literature establishes is technically easy (Kosinski et al., 2013; Peters & Matz, 2024). This is the field’s central unresolved problem, and consent architecture is a partial answer at best: it can establish a norm that individual-level reads require participation, and it can make consented profiles the higher-quality option — but it cannot prevent unconsented inference by parties who decline the norm. What it can do is remove the excuse that no alternative existed.

A third limit is worth naming precisely because it is easy to overclaim. Consent legitimates the collection and use of a profile. It does not make the profile accurate. An enthusiastically consented invalid measurement is still an invalid measurement, which is why the architecture in this article and the validation program elsewhere in this series are complements rather than substitutes: consent without validity produces authorized error, and validity without consent produces accurate surveillance.

7. Conclusion

The field’s own results establish that psychological profiling from ordinary behavioral data is easy, accurate enough to matter, and available to anyone who wants it. That is the condition under which the question of architecture becomes unavoidable: not whether profiles will exist, but whether the people they describe will hold them.

The six commitments above are not a compliance regime. They are a specification for a system in which assessment originates with the assessed — where the profile is the person’s, the uses are enumerated and revocable, the consultations are visible, some limits cannot be lifted by anyone, and the whole thing travels with its subject.

The scientific dividend is real and worth stating without embarrassment: this architecture produces better criteria, better substrate, and less distorted measurement than the alternative. But the primary argument is simpler. A field that has demonstrated how easily people can be read without asking has an obligation to demonstrate the other thing — that the same capability can be built the other way round, with the person at the origin of it rather than the object of it.

That demonstration has to be built by someone. It should be built by parties who have to live under it, including the ones who stand to profit from getting it wrong.


References

Kellogg, K. C., Valentine, M. A., & Christin, A. (2020). Algorithms at work: The new contested terrain of control. Academy of Management Annals, 14(1), 366–410.

Kosinski, M., Stillwell, D., & Graepel, T. (2013). Private traits and attributes are predictable from digital records of human behavior. Proceedings of the National Academy of Sciences, 110, 5802–5805.

Peters, H., & Matz, S. C. (2024). Large language models can infer psychological dispositions of social media users. PNAS Nexus, 3(6), pgae231.

Regulation (EU) 2016/679 of the European Parliament and of the Council of 27 April 2016 (General Data Protection Regulation). Official Journal of the European Union, L 119, 1–88.

Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 (Artificial Intelligence Act). Official Journal of the European Union.

Schwartz, H. A., Eichstaedt, J. C., Kern, M. L., Dziurzynski, L., Ramones, S. M., Agrawal, M., Shah, A., Kosinski, M., Stillwell, D., Seligman, M. E. P., & Ungar, L. H. (2013). Personality, gender, and age in the language of social media: The open-vocabulary approach. PLoS ONE, 8(9), e73791.

Youyou, W., Kosinski, M., & Stillwell, D. (2015). Computer-based personality judgments are more accurate than those made by humans. Proceedings of the National Academy of Sciences, 112(4), 1036–1040.


Citation compliance note: all 7 references above were verified against live sources on August 16, 2026, in the session that produced this draft series. No reference is cited from memory.