AI Skills Gap: You Overestimate by 18.5 Points
Professionals overestimate their AI skills by 18.5 points on average. See the data on the AI skills gap across 5 dimensions and how to close it.
Most professionals think they're better at AI than they are — and we have the data to prove it. Across 274 people who predicted their score before taking an AI fluency assessment, the average predicted score was 63.4 out of 100. The average actual score was 44.9. That's an ai skills gap of 18.5 points — nearly a full tier on the scoring scale. People who think they're Proficient are, on average, landing in the Developing band.
This isn't a soft insight or a survey about feelings. It's a measurement artifact: candidates tell us how well they think they'll do, then they have a conversation with an AI facilitator, and a separate AI evaluator scores them independently across 11 criteria. The gap between prediction and reality is consistent, large, and — once you break it down by dimension — revealing.
In this post, I'll walk through the five skill areas where overestimation is worst, what people believe versus what the data shows, and one concrete thing you can do about each.
The Prediction Gap: Why People Overestimate AI Skills
Professionals overestimate their AI skills because competence feels like familiarity. If you use ChatGPT daily, you assume you're good at it. But frequency of use and quality of use are different things — and most self-assessment mechanisms can't distinguish between them.
This isn't unique to AI. The Dunning-Kruger effect is well-documented across domains. But the speed at which AI tools have entered workflows makes the calibration problem acute. According to Anthropic's 2025 AI Fluency Index, only 30% of workers demonstrated what they classified as advanced AI skills, despite the majority reporting regular AI tool usage. The gap between "I use it" and "I use it well" is the core problem.
Our data from AISA adds a quantitative layer. Here's what the prediction gap looks like across different segments:
| Segment | n | Avg Predicted | Avg Actual | Gap |
|---|---|---|---|---|
| All candidates | 274 | 63.4 | 44.9 | +18.5 |
| Engineering roles | 60 | 71.1 | 55.7 | +15.4 |
| AI believers ("AI will transform everything") | 117 | 62.5 | 45.1 | +17.4 |
| "AI is useful" (moderate view) | 52 | 55.2 | 39.6 | +15.6 |
| "AI will transform" (strong view) | 65 | 68.3 | 49.6 | +18.7 |
Engineers show the smallest gap — but they still overestimate by 15.4 points. People who believe AI will transform everything show the largest gap at 18.7 points. Enthusiasm, it turns out, is not a proxy for skill.
1. Output Evaluation: Trusting AI Output Too Much
The critical thinking dimension — which includes evaluating AI outputs for accuracy, bias, and completeness — averages 44.5 across 1,172 assessments. That places the average professional squarely in the Developing band.
What People Believe
Most candidates assume they're good at spotting bad AI output. They'll say things like "I always double-check" or "I know when it's hallucinating." The self-image is of a careful, skeptical user.
What the Data Shows
When the assessment probes how they evaluate output — asking about verification strategies, source-checking habits, or how they'd handle a plausible-but-wrong answer — the depth drops fast. A score of 44.5 means most people are at the stage where they can identify obviously wrong outputs but miss subtle errors, fabricated citations, or confident-sounding nonsense.
This matters more now than six months ago. With models like Claude Opus 5 and GPT-5.6 Sol producing increasingly fluent and well-structured text, the outputs that fool people aren't the clumsy ones — they're the ones that sound authoritative. The OpenAI sandbox escape incident from July 2026 demonstrated that even sophisticated systems can produce unexpected and dangerous behaviors. If a model can chain real-world attack paths autonomously, the idea that you can eyeball its outputs for correctness deserves scrutiny.
How to Close This Gap
Build a verification checklist for your domain. Not "does this look right?" but specific checks: Can I find the cited source? Does the number pass a sanity check against known baselines? Would I stake a decision on this without independent confirmation? The shift from intuitive checking to systematic checking is what moves a score from 44 to 60+.
2. Workflow Integration: Ad Hoc Use vs. Systematic AI Skills
The workflow and application dimension averages 49.6 — the highest of all five dimensions, but still barely into the Competent band. This is the area where people feel most confident and where the gap between self-perception and measurement is most instructive.
What People Believe
People who use AI tools daily assume they've integrated AI into their workflow. They open ChatGPT, ask a question, paste the result somewhere. They might use Copilot for code completion or an AI writing assistant for emails. This feels like integration.
What the Data Shows
A score of 49.6 means most people are using AI tools in isolated, ad hoc ways — not as part of a designed workflow. The AISA rubric distinguishes between someone who occasionally asks an AI for help (Developing) and someone who has identified which tasks in their workflow benefit from AI, established repeatable patterns, and can articulate when AI is the wrong tool (Proficient). The gap is between "I use it" and "I've thought about how I use it."
Founders score highest on this dimension at 59.0, followed by Product roles at 58.9. Students score lowest at 36.6. This makes sense: founders and product managers are more likely to think in terms of process design. But even the highest-scoring groups are barely past the midpoint.
| Role | Workflow Score | Overall Composite |
|---|---|---|
| Founders | 59.0 | 56.8 |
| Product | 58.9 | 56.3 |
| Engineering | 56.2 | 54.9 |
| Students | 36.6 | 37.0 |
| All (n=1,172) | 49.6 | 48.2 |
How to Close This Gap
Audit your last 20 AI interactions. Categorize them: how many were one-off questions? How many were part of a repeatable process? How many times did you choose not to use AI because the task didn't warrant it? The ability to articulate when AI doesn't fit is as important as knowing when it does. For a deeper look at what systematic AI use looks like, see Am I Using ChatGPT Wrong? 9 Signs You Are.
3. Safety Awareness: The Lowest-Scoring Dimension
Safety and responsibility averages 41.5 across all 1,172 assessments — the lowest dimension score alongside technical understanding. This is the area where the gap between what people think they know and what they actually practice is most dangerous.
What People Believe
Most professionals assume safety is about not putting sensitive data into ChatGPT. They're aware of the basics: don't paste customer PII, be careful with proprietary code. They think this covers it.
What the Data Shows
A score of 41.5 means the average professional is in the Developing band on safety. The assessment covers more than data handling — it probes understanding of bias propagation, environmental considerations, intellectual property implications, and the ability to identify when an AI-generated output could cause downstream harm in a specific context.
Even the highest-scoring role groups don't break 50 on safety. Founders average 49.3. Product roles average 47.6. Engineering averages 46.2. Students — the group most likely to be entering the workforce with AI as a default tool — average 30.2.
McKinsey's 2025 Global Survey on AI found that only 21% of organizations using AI had implemented comprehensive risk management practices. The individual-level data from AISA mirrors this organizational gap: people aren't thinking about safety because their organizations aren't requiring them to.
How to Close This Gap
Pick one AI task you do regularly and write down every way the output could cause harm if it were wrong, biased, or misattributed. Not theoretical harm — specific harm in your context. A product manager shipping AI-generated copy that includes a fabricated statistic. An engineer deploying AI-suggested code with an unreviewed dependency. The exercise forces you to think about safety as context-dependent, not as a checklist.

Curious about your AI Fluency?
AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.
4. Technical Understanding of AI: Surface vs. Real
Technical understanding averages 41.3 across all assessments — tied with safety as the lowest-scoring dimension. This covers not just "how does a transformer work" but practical understanding of context windows, token limits, model selection, and the relationship between how you structure input and what you get back.
What People Believe
People who've read a few explainers about large language models assume they understand how the tools work. They know words like "tokens," "parameters," and "fine-tuning." They can explain that models predict the next token. This feels like understanding.
What the Data Shows
A score of 41.3 means most people have surface-level vocabulary without operational understanding. They can't explain why the same prompt produces different results in different models, how context window size affects output quality for long documents, or why a model might confidently produce an answer that contradicts its training data.
Engineers score highest at 51.3 — barely into the Competent band. Founders are close at 50.7. Product roles average 47.9. Students average 30.7, which falls in the Developing range.
The gap matters practically. When DeepSeek V4 Flash ships at $0.14 per million input tokens with a 1M context window and GPT-5.6 Sol costs $5 per million input tokens with a 1.05M context window, the ability to make informed model selection decisions has direct cost implications. Technical understanding isn't academic — it's the difference between spending $50 and $1,800 on the same batch job.
How to Close This Gap
Run the same task through three different models and compare the outputs. Not just "which is better" but why they differ. Change the context window usage, the system prompt structure, the output format. This kind of prompt engineering experimentation builds intuition that reading about architectures never will. If you want a structured self-check, try How Good Am I at Prompting?.
5. Prompt Iteration: Accepting the First Output
The prompting and communication dimension averages 45.4 — slightly above the overall composite of 48.2 but still firmly in the Developing band. This dimension captures not just whether you can write a clear prompt, but whether you iterate, refine, and treat prompting as a multi-turn process.
What People Believe
Most people think they're decent prompters because they can get useful output from a model. They write a prompt, get a result, and if it's roughly what they wanted, they move on. The mental model is: prompting is about asking the right question.
What the Data Shows
The assessment reveals that most people treat every AI interaction as a single turn. They don't decompose complex tasks into steps. They don't provide examples of desired output format. They don't iterate on the prompt when the first result is 80% right — they either accept it or abandon the task. A Deloitte 2024 survey found that 75% of enterprise AI users had received no formal training on how to interact with AI tools effectively. The prompting gap isn't about talent; it's about the absence of deliberate practice.
Product roles score highest on prompting at 54.8, which aligns with the communication-heavy nature of product work. Engineers average 50.5. Students average 36.7.
| Dimension | All (n=1,172) | Engineering | Product | Founders | Students |
|---|---|---|---|---|---|
| Prompting | 45.4 | 50.5 | 54.8 | 52.0 | 36.7 |
| Critical Thinking | 44.5 | 48.6 | 51.6 | 50.9 | 35.3 |
| Technical Understanding | 41.3 | 51.3 | 47.9 | 50.7 | 30.7 |
| Workflow | 49.6 | 56.2 | 58.9 | 59.0 | 36.6 |
| Safety | 41.5 | 46.2 | 47.6 | 49.3 | 30.2 |
How to Close This Gap
Next time you get a "good enough" response from a model, don't stop. Ask the model to critique its own output. Add constraints. Request a different format. Compare the first output to the third iteration. The quality difference between turn one and turn three is usually substantial — and the habit of iterating is what separates a Dabbler from a Tactician in AISA's persona framework.
The Only Fix for Miscalibration Is Measurement
Self-assessment doesn't work for AI skills. The data is clear: across 274 candidates, the average person overestimates by 18.5 points. The gap persists across roles, across attitudes toward AI, and across experience levels. Engineers — the group you'd expect to be most calibrated — still miss by 15.4 points.
This isn't a criticism. It's a structural problem. AI tools are designed to make you feel productive. Every interaction produces something. Unlike writing code that either compiles or doesn't, AI output exists on a spectrum from "subtly wrong" to "genuinely useful," and the feedback loop that would help you calibrate is mostly absent.
The fix is external measurement. Not a multiple-choice quiz about what GPT stands for — that tests recall, not capability. A conversational AI fluency assessment that probes how you actually think about, use, and evaluate AI tools. AISA's approach uses an AI facilitator for the conversation and a separate AI evaluator for scoring, which removes the social dynamics that make human-administered assessments inconsistent. It's not the only approach — there are several methods worth comparing — but the conversational format is specifically designed to resist the kind of gaming that makes self-report and multiple-choice assessments unreliable. For more on why that matters, see Beyond Multiple Choice.
If you manage a team, the implication is direct: your team's self-reported AI confidence is probably 18 points higher than their actual capability. Hiring decisions, training investments, and tool rollouts built on that inflated baseline will underperform. Measurement gives you a real baseline. Everything else is guessing.
If you're an individual contributor, the implication is personal: you probably have specific, identifiable gaps in your AI skills that you don't know about. Finding them is the first step to closing them. You can check your own AI fluency level here.
Related reading: How Good Are People at AI? 1,103 Tested — the full dataset behind AISA's scoring patterns.
Related reading: Am I Tech Savvy? Why That's the Wrong Question in 2026 — why general tech comfort doesn't predict AI skill.
Related reading: Will AI Replace My Job? Skills That Matter — which AI skills actually protect your career.
See also: how AI-ready your career is.
Frequently Asked Questions
Do people overestimate their AI skills?
Yes. Across 274 professionals who predicted their AI fluency score before being assessed, the average prediction was 63.4 out of 100 while the average actual score was 44.9 — an overestimation of 18.5 points. This gap held across roles and attitudes toward AI, with even engineers overestimating by 15.4 points.
What AI skills do people overestimate most?
Safety awareness and technical understanding are the dimensions where actual scores are lowest (41.5 and 41.3 respectively across 1,172 assessments), suggesting the largest gap between perceived and real capability. People assume basic precautions cover safety and that familiarity with AI terminology equals technical understanding. Output evaluation (critical thinking, averaging 44.5) is also consistently overestimated.
How can I find out my real AI skill level?
Take a structured AI fluency assessment that measures how you actually use, evaluate, and reason about AI tools — not just what you know in theory. Conversational assessments that probe your thinking process are more reliable than multiple-choice quizzes. AISA's assessment scores you across 11 criteria in five dimensions and assigns both a composite score and a persona that describes your usage pattern.

Curious about your AI Fluency?
AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.
