STUDY
The AI Interpreter
A Criterion-Referenced Study of the Validity and Reliability of Large Language Models in Interpreting Personality Profiles.
The research question
Given identical standardized personality-assessment scores, how accurately and consistently do major large language models reproduce the criterion-supported behavioral interpretation of those scores?
Five large language models interpreted an identical panel of ten standardized DISC Plus personality profiles — natural and adaptive DISC scores together with seven Values Index scores. Before model administration, each profile was associated with a predetermined behavioral criterion describing the interpretations that should follow from the score configuration. Models were therefore evaluated not on whether their descriptions sounded plausible or were favorably received, but on the degree to which their interpretations agreed with the predefined criterion.
The instrument is held constant. The object being evaluated is the interpreter.
What was measured
First is criterion-referenced interpretive validity: whether a model's interpretation agrees with the predetermined behavioral meaning assigned to the supplied score configuration.
Second is cross-model consistency, or inter-model agreement: whether different language models presented with identical assessment data reach substantively similar conclusions.
Third is discriminant sensitivity: whether models preserve meaningful differences between profiles, particularly when those differences are relatively small.
Fourth is epistemic restraint: whether a model recognizes when a requested inference cannot reasonably be supported by the supplied assessment data.
A secondary question concerns negative-trait attenuation: whether models preserve unfavorable but criterion-supported interpretations or soften them to the point that their substantive meaning changes.
Findings
Exact agreement varied substantially across the five models, ranging from 51.3% to 75.0%. The unweighted mean of the five model-level accuracy rates was 63.7%. H1 predicted that no model would exceed 70% exact criterion agreement. H1 was disconfirmed.
Models distinguished strongly dissimilar profiles but reproduced 117 of 120 judgments identically across deliberately similar profile pairs — a rate of 97.5% — indicating limited sensitivity to small criterion-relevant differences.
Models generally preserved negatively valenced interpretations rather than reversing them into positive traits, although their narrative language frequently softened those interpretations.
The proportion of unsupported items answered above the expected confidence floor ranged from 0% for GPT to 100% for Gemini. The unweighted mean across models was 42%. Confidence on items that could not legitimately be inferred from the supplied personality scores varied sharply by model.
The findings demonstrate why apparent plausibility should not be treated as evidence of accurate psychometric interpretation.