| Anthropic Finding | AISA Criterion | Alignment | What AISA Does Differently |
|---|---|---|---|
| Iteration — the #1 fluency signal | P2 Iterative Dialogue — Tier 1 | Strong validation P2 is Tier 1 (highest priority), confirmed by Anthropic's data. |
Scores iteration quality: generic "try again" = 3-4, strategic multi-turn sequences = 7-8. Anthropic measured presence only. |
| Directive behaviours (clarifying goals, format, examples) | P1 Prompt Design — Tier 1 | Direct match These are the core P1 skills. |
Distinguishes using one technique inconsistently (3-4) from adapting structure to task type (7-8). Anthropic counted binary presence. |
| Setting collaboration terms (only 30% of users) | P3 Context & Memory — Tier 2 | Partial match P3 covers context architectures and memory management, which subsumes collaboration terms. |
Goes deeper — persistent memory layers, multi-conversation workflows, system prompts. Anthropic's data suggests these skills are even rarer than expected. |
| Evaluative gap — polished outputs suppress fact-checking and reasoning | T1 Output Evaluation + T2 Limitation Awareness — Tier 1/2 | AISA's strongest validation The artifact finding maps directly to the Copy-Paster persona (uses AI regularly, accepts output at face value). |
T1 probes how they verify. T2 asks whether they predict failure before it happens. Anthropic could only observe whether it happened. |
| Questioning reasoning (5.6× more likely with iteration) | T1 Output Evaluation | Strong overlap | Rubric requires a named methodology for 7-8 and architectural verification for 9-10. Much more granular than binary observation. |
| Identifying missing context (4× more likely with iteration) | T2 Limitation Awareness | Good overlap | Scores domain-specific failure prediction (7-8) and reasoning from AI architecture principles (9-10). |
| Artifact & production behaviours | W1 Workflow Integration + W2 Task Decomposition | Partial Anthropic measured whether people produced artifacts; AISA measures how deeply AI is integrated into workflow. |
W1 goes from "experimental use" (1-2) to "AI-native workflow" (9-10). W2 measures decomposition quality. Anthropic didn't assess either. |
| Not measured by Anthropic | U1 AI Fundamentals, U2 Tool Landscape, W3 Domain Application, S1 Safety | Gap in Anthropic's framework 4 of AISA's 11 criteria have no equivalent in the Fluency Index. |
These require conversational depth — you can't observe tool landscape knowledge or safety reasoning from chat logs alone. |
Anthropic's researchers explicitly acknowledged that their methodology could only capture what happens inside a chat window. They identified 13 additional "unobservable" behaviours that matter for AI fluency but are invisible in chat logs. AISA's conversational assessment probes four of them directly — AI Fundamentals, Tool Landscape, Domain Application, and Safety all require the kind of structured conversational probing that chat-log observation cannot perform.
Anthropic observed behaviour passively at scale. AISA assesses it actively through structured conversation — a dual-track AI system where one model talks to the candidate while a separate model evaluates independently, scoring 11 criteria on a calibrated 1–10 rubric with a final full-transcript calibration pass by a more capable model.
The most rigorous empirical AI fluency research published to date independently confirms that AISA's framework captures the right skills, measures dimensions that observation alone cannot reach, and does so through a method purpose-built for depth and accuracy.
AISA reports and certifications reflect real, validated AI proficiency — grounded in the same behaviours that large-scale research identifies as defining AI fluency.
No framework is perfect. Anthropic's dataset surfaced areas where we can be even more precise: