The Fluency Demand Taxonomy: Scoring Work Tasks FD0–FD4
Abstract
Labor economics has spent two decades sorting work by whether it can be codified. That sorting predicted the last wave of automation well. It is now producing strange results — high-credential analytical work looks exposed while much lower-credentialed interpersonal work does not — because the dimension that mattered for software does not obviously carry over to systems that write, summarize, and reason.
This article proposes a second axis. The Fluency Demand taxonomy scores work tasks on five ordinal levels, FD0 through FD4, by how much of a task’s outcome depends on communication being fitted to a specific person. FD0 tasks require no fitting: the output is correct or incorrect independent of the recipient. FD4 tasks are constituted almost entirely by the fitting — the outcome is whether one particular person understood, accepted, decided, or was reached.
The taxonomy is an ordinal approximation of a continuous underlying quantity developed in a companion article, and it is designed to be applied to the same task-level occupational data the existing exposure literature uses. Three things follow. Scoring rules and anchors are specified so the taxonomy can be applied by independent raters and its reliability reported. The taxonomy’s central prediction is stated: within occupations, FD3–FD4 task content should be systematically less exposed to language-model substitution than FD0–FD2 content, and the value of the former should appreciate as the latter deflates. And its limits are named, including the substantial risk that a high FD score marks work that is protected today and merely harder to automate rather than impossible to automate.
Keywords: task content, occupational exposure, automation, communication, labor economics, artificial intelligence
1. Why the Old Axis Is Straining
The task framework transformed how economists think about technological change. Its central insight was that technology substitutes for tasks, not occupations, and that the tasks most readily substituted are those following explicit, codifiable rules — with computers replacing routine cognitive and manual work while complementing non-routine analytical and interactive work (Autor, Levy, & Murnane, 2003). The framework organized a generation of research and predicted the pattern of labor-market polarization that followed.
Language models strain it in a specific way. Exposure analyses of large language models find that substantial shares of work tasks across the economy could be affected, with exposure concentrated in higher-wage, higher-education occupations — the opposite of the pattern earlier automation produced (Eloundou, Manning, Mishkin, & Rock, 2024). Occupation-level exposure measures built from AI capability descriptions show similarly unfamiliar distributions (Felten, Raj, & Seamans, 2021). Non-routine analytical work, the category the earlier framework marked as complemented, is now among the most exposed.
Something in the sorting has changed. The routine/non-routine axis captured whether a task’s procedure could be specified. Language models do not require procedures to be specified; they perform tasks whose procedures were never articulable, which is precisely what made those tasks non-routine.
Meanwhile a second empirical pattern has been accumulating from a different direction. The labor market has increasingly rewarded social skill, with growth concentrated in occupations requiring high levels of social interaction and the returns to social skill rising over decades (Deming, 2017). That finding sits adjacent to the exposure literature without being integrated with it.
This article proposes that the two are the same story viewed from opposite ends, and that the axis connecting them is how much a task’s outcome depends on communication fitted to a particular person.
2. What FD Measures
Fluency Demand is an ordinal scoring of tasks by communicative load — the share of a task’s outcome variance attributable to the quality of communication between the parties, a continuous quantity defined in a companion article. FD is the practical instrument: five levels, anchored, applicable by trained raters to task descriptions, and designed for use with the task-level occupational data the exposure literature already employs.
Three properties carry over from the underlying construct and constrain how FD may be used.
FD scores tasks, not occupations or people. Every occupation is a bundle of tasks at different FD levels. A physician’s task inventory contains FD0 elements (ordering a standard panel), FD2 elements (documenting a visit), and FD4 elements (a goals-of-care conversation). Occupational FD is a distribution, and reporting a single occupational FD score discards the information that matters most.
FD is relative to a stated outcome. “Explaining a diagnosis” scores differently against the patient can repeat the treatment plan than against the note is complete. Every FD score travels with its outcome specification.
FD is not difficulty, seniority, or emotional intensity. High-FD work is often low-credentialed; low-FD work is often highly technical. Conflating FD with skill level would reproduce exactly the assumption the taxonomy exists to test.
3. The Five Levels
FD0 — No fitting required. The task’s outcome is independent of who receives it. Format conversion, arithmetic, retrieval of a specified record, execution of a defined procedure. Communication is required only to initiate the task; beyond basic intelligibility, phrasing contributes nothing to the outcome. Anchor: the output would be equally correct if delivered to a different recipient, or to no one.
FD1 — Generic clarity. Outcomes depend on the message being clear, but not on it being fitted to a particular person. Standard documentation, routine notifications, well-formed instructions for a general audience. Communication quality matters and is optimizable in the abstract — one best version exists for all recipients. Anchor: a good version can be written without knowing who will read it.
FD2 — Audience fitting. Outcomes depend on the message being fitted to a category of recipient: expertise level, role, familiarity with context. Technical explanation to a lay audience, onboarding material for a defined group, segment-targeted communication. Fitting is required, but knowing the person as an individual adds little beyond knowing which group they belong to. Anchor: knowing the recipient’s category is nearly as good as knowing the recipient.
FD3 — Individual fitting. Outcomes depend materially on the message being fitted to this specific person: how they process information, what they will resist, what they need reassured or left unsaid. Advisory work, coaching, most management, complex sales, professional counsel. Two recipients in the same demographic category require materially different handling for the same outcome. Anchor: a version optimized for the category would fail for this individual in predictable ways.
FD4 — Constitutive communication. The communication is the task. Outcomes are wholly or nearly wholly determined by whether one particular person understood, accepted, decided, disclosed, or was reached. Delivering serious diagnoses, negotiating under conflict, therapeutic work, crisis conversations, high-stakes conflict resolution. There is no separable deliverable that could be evaluated apart from how it landed with the person. Anchor: if the communication fails, the task has failed entirely, with no residual value.
The boundary requiring the most care is FD2/FD3, since it is the boundary between category-level and individual-level fitting, and it is the boundary at which person-level knowledge begins to matter. Scoring rules must supply worked examples on both sides of it.
4. Application
Unit. Task statements, drawn from occupational task inventories of the kind used in existing exposure work (Eloundou et al., 2024; Felten, Raj, & Seamans, 2021). Using the same units makes FD scores directly joinable to published exposure measures.
Procedure. Trained raters assign FD levels against anchored definitions and a written codebook containing worked examples, with each task scored against an explicitly stated outcome. A substantial subsample is double-coded, agreement is reported, and disagreements are adjudicated under written rules. Inter-rater reliability is a property of the taxonomy and belongs in every report that uses it; a taxonomy whose reliability is unreported is a set of opinions with numbers attached.
Aggregation. Occupational FD profiles are reported as distributions across levels — the share of an occupation’s task time at each FD level — never as a single mean. Two occupations with identical mean FD may have entirely different distributions, and it is the FD3–FD4 share that carries the theoretical predictions.
Validation. The taxonomy’s construct validity is established by convergence with independent estimates of communicative load: expert ratings, experimental manipulations of communication quality where ethically available, and field variance decompositions. A taxonomy that cannot be checked against something outside itself is a classification scheme, not a measurement.
5. The Central Prediction
FD3–FD4 task content is less substitutable by current language models than FD0–FD2 content, and its relative value should rise as lower-FD content deflates.
The prediction’s mechanism is specific rather than mystical. High-FD tasks require an accurate model of one particular person, sustained across the interaction, together with the standing to act on it. Existing exposure measures assess whether a model can perform the task’s visible operation — draft the message, produce the explanation — and language models perform those operations well. What they do not establish is whether the operation, performed without an accurate individual model, achieves the outcome the task exists to produce. At FD4 those are the same thing, which is what makes the level distinct.
Two corollaries are testable.
Exposure measures should systematically overstate substitutability at high FD. Because published exposure ratings assess capability to perform the operation rather than to achieve the outcome, the gap between rated exposure and realized substitution should widen with FD level. This is directly checkable by joining FD scores to existing exposure data and to subsequent employment and wage outcomes.
Returns to high-FD work should rise relative to low-FD work. This connects the taxonomy to the social-skill findings already in the literature (Deming, 2017) and gives them a mechanism: social-skill returns are the observable shadow of FD3–FD4 task content becoming scarce relative to demand.
A third implication concerns systems rather than labor markets, and it is stated as a boundary rather than a prediction: FD identifies where accurate person-knowledge could matter, not whether any particular system possesses it. Whether machine person-perception is accurate enough to convert FD3–FD4 demand into performance is an empirical question addressed elsewhere in this program, and the taxonomy is deliberately silent on it.
6. Objections
“This is a defensive taxonomy — it defines the last refuge of human labor.” The objection has force, and the taxonomy is stated to be falsifiable rather than reassuring. If systems achieve accurate individual person-models and deploy them competently, FD3–FD4 work becomes exposed too, and FD will have identified the sequence of substitution rather than a permanent boundary. That would still be a useful contribution and a considerably less comfortable one. Any taxonomy of what machines cannot do should be written expecting to be overtaken, and its authors should say so in advance.
“High-FD work is protected by regulation and trust, not by communicative difficulty.” Partly true and worth separating empirically. Licensure, liability, and institutional trust do protect several high-FD occupations, and those protections are analytically distinct from communicative load. The separation is testable: FD3–FD4 tasks in unregulated occupations — sales, coaching, management, customer resolution — should show the predicted pattern if FD is doing the work, and should not if regulation is.
“The FD2/FD3 boundary is arbitrary.” It is the taxonomy’s most consequential judgment and requires the most careful anchoring, which is why the criterion is stated operationally: whether knowing the recipient’s category is nearly as good as knowing the recipient. Where raters disagree, the disagreement is measurable and reportable rather than hidden — which is the most that can be asked of an ordinal scheme.
“Ordinal levels lose information relative to a continuous measure.” Yes, deliberately. Cardinal estimation of communicative load requires experimental manipulation that is expensive and, for the highest-load tasks, frequently unethical. The ordinal scheme is what can actually be applied at the scale of an occupational task inventory, and the theoretical predictions it supports are ordinal predictions.
7. What the Taxonomy Is For
For labor research, FD supplies a second axis to join to existing exposure data, with the immediate empirical question being whether FD predicts realized substitution better than rated exposure alone.
For education and training, the FD2/FD3 boundary identifies where curricula should concentrate if the prediction holds: not on communication in general, which is FD1–FD2 territory and increasingly machine-assisted, but on the individual-fitting capacities that FD3–FD4 work requires.
For system design and honest capability claims, FD marks where person-level accuracy could plausibly pay and where it is decoration. A personalization capability deployed on FD0–FD1 work is theater regardless of how good it is, and a taxonomy that says so protects buyers from claims the underlying science does not support.
8. Conclusion
The routine/non-routine axis answered the question technology posed for forty years: can this procedure be specified? Language models have made that question less decisive, because they perform tasks whose procedures were never specified.
The question FD poses is different: how much of this task’s outcome depends on reaching one particular person? At FD0 the answer is none, and the work is fully specifiable. At FD4 the answer is nearly all, and there is no deliverable separable from how it landed.
Whether that axis carries the predictive weight this article claims is an empirical question, joinable to existing data, and stated here with its falsification conditions attached. If FD3–FD4 work turns out to be exposed at the same rate as everything else, the taxonomy will have been wrong in a way that is worth knowing quickly.
References
Autor, D. H., Levy, F., & Murnane, R. J. (2003). The skill content of recent technological change: An empirical exploration. Quarterly Journal of Economics, 118(4), 1279–1333.
Deming, D. J. (2017). The growing importance of social skills in the labor market. Quarterly Journal of Economics, 132(4), 1593–1640.
Eloundou, T., Manning, S., Mishkin, P., & Rock, D. (2024). GPTs are GPTs: Labor market impact potential of LLMs. Science, 384(6702), 1306–1308.
Felten, E., Raj, M., & Seamans, R. (2021). Occupational, industry, and geographic exposure to artificial intelligence: A novel dataset and its potential uses. Strategic Management Journal, 42(12), 2195–2217.
Citation compliance note: all 4 references above were verified against live sources on August 16, 2026, in the session that produced this draft series. No reference is cited from memory.
TOPICS
RELATED
- VIDEO
Scrubbing the Answer Key
A spoken walkthrough of contamination control and pre-registration discipline: four classes of leakage, auditable removal rules, and how findings stay confirmatory.
- VIDEO
The Fluency Demand Taxonomy: Scoring Work Tasks FD0–FD4
A spoken walkthrough of the second axis of occupational exposure: five ordinal levels scoring how much a task’s outcome depends on communication fitted to a specific person.
- VIDEO
Calibration as Co-Equal Criterion
A spoken walkthrough of why confidence–accuracy coupling belongs beside accuracy as a headline criterion, never composited into a single number.