Calibration and Sycophancy
Confidence that tracks accuracy, and the failure modes — sycophantic drift, halo compression, unearned intimacy — that flattery introduces.
- VIDEO
Calibration as Co-Equal Criterion
A spoken walkthrough of why confidence–accuracy coupling belongs beside accuracy as a headline criterion, never composited into a single number.
- VIDEO
The D0 Problem: Fabricated Familiarity at Zero Acquaintance
A spoken walkthrough of the D0 condition — negligible evidence, full elicitation, confidence required — and why the fabricated-familiarity rate belongs in every public report.
- VIDEO
Sycophancy as Construct Invalidity
A spoken walkthrough of why sycophancy is a measurement pathology rather than a tone problem: an instrument that drifts toward what the subject wants to hear is not measuring the subject.
- ARTICLE
The AI as Candidate Instrument: A Full Validity Battery for Machine Judgments of Persons
If an AI system judges the people it talks to, it is a candidate instrument. This article specifies the five-component validity battery — convergence, discrimination, dose–response, calibration, and profile accuracy — needed to test one.
- ARTICLE
Calibration as Co-Equal Criterion: Confidence–Accuracy Coupling in Machine Judgments of Humans
A system that is confidently wrong about people is a different kind of object than one that is uncertainly wrong. This article argues calibration belongs beside accuracy as a headline criterion, never composited into one number.
- ARTICLE
Sycophancy as Construct Invalidity: When the Instrument Measures Its Own Incentives
Sycophancy is not a tone problem but a measurement pathology: an instrument whose readings bend toward what the subject wants to hear is no longer measuring the subject.
- ARTICLE
The Halo in the Machine: Discriminant-Validity Failure and Impression Compression in LLM Person-Judgments
Why model assessments of different people look more alike than the people do — and a measurement framework for halo, compression, and discriminant-validity failure.
- ARTICLE
Unearned Intimacy: Anthropomorphism, Trust Miscalibration, and Systems That Claim to Know You
Fluent, remembering, speaking systems generate trust through mechanisms independent of their accuracy about you — a structural, measurable risk.