Why Conversational AI Assessment Works
Self-reported AI skills correlate poorly with actual ability (r = 0.07–0.24). Conversational assessment measures real behaviour, not checkboxes.
The hardest thing about measuring AI skills is that the people who need measurement most are the least equipped to assess themselves. This isn't a character flaw — it's a measurement problem. And the solution isn't a better quiz.
The Self-Report Problem
In January 2026, researchers Zhang, Xiao, Botelho, Liao, Chiu, Stamper, and Koedinger published a study analysing 288 professionals' AI literacy using both self-reported and objective measures. The correlation between what people said they could do and what they could actually do was r = 0.07 to 0.24 — barely above random.
This isn't surprising. Self-assessment requires calibrated self-awareness, which requires experience with the thing you're assessing yourself on. AI fluency is new enough that most professionals have no reliable internal benchmark.
AISA's own data confirms this at scale. Across 307 professionals who predicted their AI fluency score before taking a conversational assessment, the average prediction was 64 out of 100. Their average actual score was 45. A gap of 19 points — not in the direction of modesty.
Any measurement system that relies on self-report inherits this problem. The question is what replaces it.
The Evidence Hierarchy
Not all evidence of AI skill is equal. There is a hierarchy, and the method you choose determines what level of evidence you can collect.
Demonstrated — the candidate performs a task, makes a decision, or walks through a real example from their work. You see the behaviour, not a description of it. This is the strongest evidence tier.
Described — the candidate tells you what they do. "I always cross-check AI outputs against primary sources." Credible, but unverifiable in the moment. Ceiling: moderate confidence.
Managed — the candidate explains what their team or organisation does. "We have a RAG pipeline for our knowledge base." This tells you about their environment, not their personal proficiency. Ceiling: low confidence for individual measurement.
Recognised — the candidate selects the correct answer from a list. This tells you they can identify the right answer when presented with it. It does not tell you whether they would generate that answer unprompted, apply it in context, or notice when it's needed. This is what multiple-choice tests measure.
A quiz operates at the recognition level. A self-report survey doesn't even reach that — it measures the candidate's belief about their own ability. A conversational assessment operates at the demonstrated and described levels — the candidate talks through real examples, responds to follow-up probes, and demonstrates reasoning in real time.
Why Quizzes Don't Work for AI Skills
Multiple-choice and knowledge-based quizzes have a defensible place in education and compliance testing. They don't work well for AI skills measurement for three specific reasons.
The staleness problem
AI tools, capabilities, and best practices change monthly. A question about GPT-4's context window written in January is factually wrong by March. Claude's tool use capabilities expanded three times in 2025 alone. A question bank that was accurate at publication is misleading within weeks.
A conversation about how you actually use AI is always current — because the candidate describes their current practice, not a frozen snapshot of the field.
The AI-answering-AI problem
Any question that can be pasted into a chatbot and answered correctly is not measuring the candidate — it's measuring whether they have a second device. Proctoring mitigates this but doesn't eliminate it, and creates friction that degrades the assessment experience.
A conversation is harder to outsource because it's adaptive. The next question depends on what the candidate just said. There is no answer key to look up because the questions don't exist until the conversation generates them.
The recognition-application gap
The distance between "I know what prompt chaining is" and "I routinely chain prompts when the task has sequential dependencies" is the entire point of fluency measurement. A quiz can test the first. Only a conversational probe — "Tell me about a time you broke a complex task into sequential AI steps" — can surface the second.
AI fluency is a practice, not a body of knowledge. You measure practices by observing them or hearing someone articulate them under probing. You don't measure them by asking which definition is correct.
How Conversational Assessment Works
AISA's implementation uses a dual-track architecture that separates conversation from evaluation — solving the bias problem that occurs when a single system both asks questions and judges answers.
Track A (the conversationalist) manages the dialogue. It's warm, adaptive, and peer-level. It follows up on what the candidate says, pivots when a topic is exhausted, and introduces exercises that probe specific competencies. The candidate talks to Track A and only Track A.
Track B (the silent evaluator) runs on every candidate message. It scores the response against an 11-criterion rubric, extracts verbatim quotes as evidence, classifies evidence quality (demonstrated vs. described vs. managed), and generates steering notes that guide Track A's next question. The candidate never knows Track B exists, which prevents evaluation anxiety from distorting performance.
Calibration pass. After the session, a more capable model (Claude Opus) reviews the entire transcript with full context and adjusts scores where the per-turn evaluator accumulated biases. Every adjustment requires disconfirming evidence — the calibrator must explain why a score should change, not just assert it.
The result: every score is tied to a specific quote from the conversation. A hiring manager reading the report can see exactly what the candidate said that produced each score — and judge for themselves whether the score is fair.
What Conversational Assessment Measures That Others Can't
AISA's framework covers 5 dimensions and 11 criteria. When cross-referenced against Anthropic's AI Fluency Index (which analysed 9,830 real conversations), AISA covers 93% of their observable fluency markers — plus four criteria that Anthropic's passive observation method cannot measure:
- AI Fundamentals (U1) — requires direct probing; invisible in chat behaviour logs
- Tool Landscape (U2) — cross-platform ecosystem knowledge; unobservable on a single platform
- Domain Application (W3) — profession-specific AI use; domain context is lost at scale
- Safety & Responsibility (S1) — risk awareness and downstream impact; Anthropic listed this among their unobservable behaviours
These four criteria require a structured conversation to surface. A quiz can test knowledge of AI safety concepts. A conversation can probe whether someone actually applies safety thinking in their work — and follow up when their initial answer is vague.
When Not to Use Conversational Assessment
Conversational assessment is not the right tool for every situation.
Quick team pulse checks. If you need a 2-minute snapshot of how your team feels about AI readiness, a self-report survey is faster and cheaper. You're measuring attitude, not ability — and that's fine for the question you're asking.
Compliance verification. If you need to verify that 500 employees completed AI safety training and can identify the key concepts, a multiple-choice test is appropriate. You're verifying knowledge transfer at scale, not measuring fluency.
Specific tool proficiency. If the question is "can this developer use GitHub Copilot effectively?", a scenario simulation with the actual tool is more direct than a conversation about it.
Budget-constrained bulk screening. Conversational assessment costs more per candidate than a quiz because it runs two AI models for 20–40 minutes per person. At high volume with tight budgets, a lighter screening method followed by conversational assessment for finalists may be more practical.
The right method depends on the question you're trying to answer. Conversational assessment answers "how fluent is this person at working with AI?" with evidence. Other methods answer different, narrower questions — and that's sometimes exactly what you need.
The Measured Reality
Across 1,800 completed AISA conversational assessments, the data paints a consistent picture:
- Average AI fluency score: 48/100
- 67% of professionals score below the Proficient tier
- Prediction gap: +19 points average overestimation
- Weakest skill: Tool Landscape (4.8/10) — most professionals cannot differentiate AI tools or explain when to use which
These numbers are measured from demonstrated behaviour in conversation — not from self-reported surveys or quiz scores. The methodology is published, the rubric is auditable, and every score maps to a verbatim quote.
Full methodology: AISA Methodology | Full rubric: The AISA Rubric | Full comparison of assessment methods: AI Assessment Methods Compared

Curious about your AI Fluency?
AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.

Curious about your AI Fluency?
AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.