Independent Validation

Anthropic's 10,000-Conversation Study
Validates AISA's Assessment Framework

The most comprehensive empirical study of AI fluency confirms that AISA's assessment criteria match, and in key areas surpass, what the research identifies as the markers of effective AI use.

Criterion-Level Coverage

How AISA's rubric maps to every behaviour Anthropic's research identifies as defining AI fluency

93%
of Anthropic's observable AI fluency behaviours
are already assessed in AISA's framework
Anthropic Finding AISA Criterion Alignment What AISA Does Differently
Iteration — the #1 fluency signal P2 Iterative Dialogue — Tier 1 Strong validation
P2 is Tier 1 (highest priority), confirmed by Anthropic's data.
Scores iteration quality: generic "try again" = 3-4, strategic multi-turn sequences = 7-8. Anthropic measured presence only.
Directive behaviours (clarifying goals, format, examples) P1 Prompt Design — Tier 1 Direct match
These are the core P1 skills.
Distinguishes using one technique inconsistently (3-4) from adapting structure to task type (7-8). Anthropic counted binary presence.
Setting collaboration terms (only 30% of users) P3 Context & Memory — Tier 2 Partial match
P3 covers context architectures and memory management, which subsumes collaboration terms.
Goes deeper — persistent memory layers, multi-conversation workflows, system prompts. Anthropic's data suggests these skills are even rarer than expected.
Evaluative gap — polished outputs suppress fact-checking and reasoning T1 Output Evaluation + T2 Limitation Awareness — Tier 1/2 AISA's strongest validation
The artifact finding maps directly to the Copy-Paster persona (uses AI regularly, accepts output at face value).
T1 probes how they verify. T2 asks whether they predict failure before it happens. Anthropic could only observe whether it happened.
Questioning reasoning (5.6× more likely with iteration) T1 Output Evaluation Strong overlap Rubric requires a named methodology for 7-8 and architectural verification for 9-10. Much more granular than binary observation.
Identifying missing context (4× more likely with iteration) T2 Limitation Awareness Good overlap Scores domain-specific failure prediction (7-8) and reasoning from AI architecture principles (9-10).
Artifact & production behaviours W1 Workflow Integration + W2 Task Decomposition Partial
Anthropic measured whether people produced artifacts; AISA measures how deeply AI is integrated into workflow.
W1 goes from "experimental use" (1-2) to "AI-native workflow" (9-10). W2 measures decomposition quality. Anthropic didn't assess either.
Not measured by Anthropic U1 AI Fundamentals, U2 Tool Landscape, W3 Domain Application, S1 Safety Gap in Anthropic's framework
4 of AISA's 11 criteria have no equivalent in the Fluency Index.
These require conversational depth — you can't observe tool landscape knowledge or safety reasoning from chat logs alone.

Where We Go Even Deeper

4 of AISA's 11 criteria have no equivalent in Anthropic's framework

Anthropic's researchers explicitly acknowledged that their methodology could only capture what happens inside a chat window. They identified 13 additional "unobservable" behaviours that matter for AI fluency but are invisible in chat logs. AISA's conversational assessment probes four of them directly — AI Fundamentals, Tool Landscape, Domain Application, and Safety all require the kind of structured conversational probing that chat-log observation cannot perform.

U1
AI Fundamentals
How AI works — tokens, training, inference. Requires direct probing, invisible in chat behaviour.
U2
Tool Landscape
Cross-platform ecosystem knowledge — can't observe multi-tool awareness on a single platform's logs.
W3
Domain Application
AI use tailored to a specific profession — domain context is lost at scale.
S1
Safety & Responsibility
Risk awareness, data boundaries, downstream impact — Anthropic listed this among their 13 unobservable behaviours.

A Fundamentally Different Method of Measurement

Observation tells you what people do. Conversation tells you why — and what they'd do differently

Anthropic observed behaviour passively at scale. AISA assesses it actively through structured conversation — a dual-track AI system where one model talks to the candidate while a separate model evaluates independently, scoring 11 criteria on a calibrated 1–10 rubric with a final full-transcript calibration pass by a more capable model.

Anthropic's Approach
  • Passive observation of chat logs
  • Binary: behaviour present or absent
  • Single platform (Claude.ai)
  • Can't probe — only observe
  • 9 observable behaviours
  • Self-selected sample
AISA's Approach
  • Active conversational assessment
  • Scored 1–10 with calibrated rubric
  • Platform-agnostic — assesses all tools
  • Probes directly — surfaces hidden skills
  • 11 criteria across 5 dimensions
  • Full-transcript post-session calibration

What This Means

The most rigorous empirical AI fluency research published to date independently confirms that AISA's framework captures the right skills, measures dimensions that observation alone cannot reach, and does so through a method purpose-built for depth and accuracy.

AISA reports and certifications reflect real, validated AI proficiency — grounded in the same behaviours that large-scale research identifies as defining AI fluency.

What We're Doing Next

How Anthropic's data is helping us sharpen our framework further

No framework is perfect. Anthropic's dataset surfaced areas where we can be even more precise:

  • Collaboration framing — only 30% of users set interaction terms with AI. We're making this an explicit scoring anchor in our P3 criterion.
  • The polished output trap — professional-looking AI artifacts suppress critical evaluation (−5.2pp). We're adding targeted probes for this blind spot.
  • Iteration as a gateway signal — with a 2× fluency multiplier, we're exploring how to use iteration quality earlier in our adaptive difficulty system.
Independence note: AISA was designed and built independently, before the publication of Anthropic's AI Fluency Index. Anthropic does not own, endorse, accredit, or directly contribute to AISA. The AI Fluency Index is publicly available research. This comparison is our own analysis of how our pre-existing framework aligns with their independently published findings.
Source: Anthropic Education Report — The AI Fluency Index (Feb 2026) · 9,830 conversations · anthropic.com/research
AISA AI Skills Assessment · 11 criteria · 5 dimensions · aisa.to