Product Managers AI Skills: Most Accurate Self-Assessors

Product managers predict their AI skills most accurately among professionals. Data from 694 assessments reveals calibration gaps by role.

By Ozan Dagdeviren··12 min read
datarolesproduct-managementpredictionproduct managers ai skillsai self assessment accuracycalibrationprediction gapai fluencyteam ai readiness

Product managers' predicted AI scores land closest to their actual scores — closer than engineers, founders, or students. Across 694 completed AI fluency assessments on AISA where candidates estimated their own score before starting, product managers overestimated by an average of 12 points (n=30), making them the most accurately calibrated professional group in our dataset. The overall average overestimation across all roles was 17.8 points (n=694). This post breaks down the AI self-assessment accuracy data by role, explores why calibration might be a core PM competency, identifies the worst calibrators, and argues that calibration matters more than raw score when building AI-capable teams.

Self-Assessment Accuracy Ranking by Role

Founders are the most accurately calibrated group in our data, followed closely by product managers. Engineers overshoot by nearly 15 points, and students miss by a staggering 36. Here is the full ranking from our prediction gap data:

RolenAvg PredictedAvg ActualAvg Gap (Overestimation)
Founders4761.755.1+6.6
Product3067.655.6+12.0
Engineering13570.755.9+14.8
Students4764.628.6+36.0
All roles combined69462.744.9+17.8

A few things jump out immediately.

Founders and PMs Share a Calibration Advantage

Founders actually edge out product managers on raw gap size (+6.6 vs. +12.0). But both groups share something important: their actual scores are nearly identical (55.1 and 55.6 respectively), and both predict in a realistic range. Neither group wildly inflates their self-image.

Engineers Predict Highest, But Overshoot

Engineers predicted the highest scores of any group (70.7) and did score highest in actuality (55.9). But the +14.8 gap means they consistently believe they're roughly 15 points better than they are. That's the difference between landing in the Developing tier and the Proficient tier — a meaningful miscalibration.

Students Are in a Different Category Entirely

A +36.0 gap is not a rounding error. Students predicted 64.6 (solidly Proficient) but scored 28.6 (Emerging tier). This is the Dunning-Kruger effect in its purest measurable form.

Why Calibration Might Be a Product Manager Skill

Product managers professionally estimate effort, scope risk, and receive rapid feedback on the accuracy of their estimates. This feedback loop — predict, measure, adjust — is the exact mechanism that builds calibration over time. It's not surprising that this transfers to AI self-assessment accuracy.

The Estimation Muscle

PMs estimate story points, forecast launch timelines, size markets, and prioritize backlogs. Every sprint retrospective is a calibration event: "We predicted X velocity, we delivered Y." Over hundreds of these cycles, PMs develop a refined sense of what they know versus what they think they know. When asked "How well do you understand AI?" they apply the same estimation discipline.

Founders share a version of this — they pitch investors, forecast revenue, and get corrected by reality constantly. Their +6.6 gap reflects even tighter calibration, likely because the stakes of miscalibration (running out of money) are existential.

Connection to Critical Thinking

AISA's assessment framework weights Critical Thinking at 22% of the composite score. This dimension measures exactly the skills that drive calibration: evaluating AI outputs against expectations, recognizing the limits of one's own understanding, and adjusting confidence based on evidence.

Product managers score 51.7 on Critical Thinking (n=95) — the second-highest of any role group, behind only Data professionals at 50.6 (n=34). For context, the population average is 43.8 (n=1,598). PMs don't just predict their scores well; they also perform well on the dimension most closely associated with accurate self-assessment.

Feedback Loops vs. Echo Chambers

Engineers often work in environments where AI tools provide immediate, tangible output — code compiles or it doesn't. This creates a sense of mastery that may not transfer to broader AI fluency. A developer who uses GitHub Copilot daily might reasonably feel expert-level, but AISA's assessment covers prompting strategy, hallucination detection, safety awareness, and workflow design — areas where daily Copilot use doesn't automatically build skill.

PMs, by contrast, tend to interact with AI across a wider surface area: writing specs with LLMs, evaluating AI features for their products, assessing vendor claims. This breadth may produce more realistic self-models.

The Worst Calibrators: What Overshoot Looks Like

The largest prediction gaps don't just indicate overconfidence — they reveal specific patterns of misunderstanding about what AI fluency actually requires.

Students: +36.0 Gap (n=47)

Students predicted an average score of 64.6, which would place them in the Proficient tier (60-79). They actually scored 28.6, which is Emerging (0-27) — barely clearing the threshold into Developing. This is the largest gap in our data by a wide margin.

The dimension scores tell the story. Students average 34.7 on Prompting, 33.1 on Critical Thinking, 29.1 on Technical Understanding, 34.3 on Workflow, and 27.5 on Safety (n=124 for dimension data). Their weakest area — Safety & Responsibility at 27.5 — is also the area where overconfidence carries the most real-world risk.

Why the massive gap? Students likely conflate AI usage with AI fluency. Using ChatGPT to write essays is not the same as understanding context windows, evaluating output reliability, or designing multi-step workflows. The gap between "I use AI every day" and "I use AI effectively and safely" is where the +36 points live.

Engineers: +14.8 Gap (n=135)

Engineers' overshoot is more nuanced. They genuinely are the highest-scoring group on Technical Understanding (51.6, n=325) and Workflow (56.4). But they predict 70.7, which would place them in the Proficient-to-Advanced range. Their actual 55.9 is solidly Developing-to-Proficient.

The gap concentrates in two areas. First, Safety & Responsibility: engineers average 47.3, well below their self-image as technically rigorous professionals. Second, the gap between knowing how AI works technically and knowing how to deploy it responsibly in a team context. An engineer might understand transformer architectures but still lack fluency in AI ethics frameworks or bias detection.

This pattern echoes findings from Anthropic's research. Their AI Fluency Index — which AISA's framework overlaps with at 93% — similarly finds that technical depth doesn't automatically produce well-rounded AI competence.

The Overall Population: +17.8 Gap (n=694)

Across all 694 candidates who provided predictions, the average gap is +17.8 points. The average predicted score (62.7) would place the typical candidate in the Proficient tier. The average actual score (44.9) places them in Developing. Nearly the entire population believes they're a tier higher than they actually are.

This is consistent with broader research on self-assessment accuracy. A 2024 study from the National Bureau of Economic Research found that workers overestimate their skills by 20-30% on average across domains, with the gap widening in newer fields where external benchmarks are scarce. AI fluency is exactly such a field.

AISA

Curious about your AI Fluency?

AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.

Why Calibration Matters More Than Raw Score

A team of engineers scoring 55 who think they score 55 is safer than a team scoring 55 who think they score 70. Miscalibrated confidence is where AI risk lives. This is the core argument for measuring calibration alongside competence.

The Risk of Confident Incompetence

When someone overestimates their AI skills by 15+ points, they make predictable mistakes:

  • They skip verification. If you believe you're Proficient at evaluating AI output, you're less likely to double-check a response that "looks right." This is the mechanism behind cognitive surrender — deferring to AI output without critical evaluation.
  • They deploy without guardrails. Overconfident teams ship AI features without adequate testing, monitoring, or fallback mechanisms. The recent CoSnitch vulnerability in Microsoft Copilot (CVE-2026-24301, patched August 18) is a reminder that even mature AI integrations carry security risks that require vigilant, calibrated teams.
  • They underestimate training needs. A manager who believes their team scores 70 won't invest in the training that a team actually scoring 55 needs.

Calibration as a Hiring Signal

When evaluating candidates for AI-adjacent roles, the prediction gap may be more informative than the raw score. Consider two candidates:

  • Candidate A: Predicts 65, scores 60. Gap: +5. This person knows what they know.
  • Candidate B: Predicts 80, scores 65. Gap: +15. This person scores higher but has a blind spot about their own limitations.

For roles involving AI deployment decisions — product managers, engineering leads, anyone touching production AI systems — Candidate A's calibration may matter more than Candidate B's extra 5 points of raw skill.

AISA's AI skills assessment captures both the raw score and the prediction gap, giving hiring managers and team leads a two-dimensional view of AI readiness.

Team-Level Calibration

The implications scale to teams. If your engineering org has an average prediction gap of +14.8 (as our data shows), your team collectively believes it's operating at a level it hasn't reached. This affects:

  • Resource allocation: Teams that overestimate their AI fluency under-invest in training and over-invest in ambitious AI projects they're not ready to execute.
  • Risk assessment: Overconfident teams underestimate the probability and severity of AI failures.
  • Vendor evaluation: Teams that overrate their own skills are worse at evaluating AI tools and vendors — they don't know what questions to ask.

The EU AI Act's transparency requirements, enforced since August 2, add regulatory weight to this argument. Organizations deploying AI systems now face fines up to €15M or 3% of global turnover for non-compliance. Miscalibrated teams are compliance risks.

What to Do With This Data

If you're a product manager reading this, the data suggests your professional habits serve you well in AI self-assessment. But a +12.0 gap still means you're overestimating — just less than everyone else.

If you're an engineering manager, the +14.8 gap in your function is worth addressing directly. Consider:

  1. Baseline your team. Use a conversational assessment (not a multiple-choice quiz — those are gameable and shallow) to get actual scores alongside self-predictions.
  2. Share the gap data. Making the prediction gap visible is itself a calibration intervention. People adjust when they see the evidence.
  3. Focus on the weak dimensions. Engineers' Safety & Responsibility scores (47.3) lag their Technical Understanding (51.6). Targeted development in safety and ethics closes both the skill gap and the calibration gap.

If you manage students or early-career professionals, the +36.0 gap demands a fundamentally different approach. These individuals need structured exposure to what AI fluency actually entails before they can begin to self-assess accurately. The AI certification for product managers path, for example, builds the kind of structured feedback loops that improve calibration over time.

Building Better Calibration Across Your Organization

Calibration isn't fixed. It's a skill that improves with deliberate practice and feedback. The same mechanisms that make PMs and founders better calibrators can be replicated across roles.

Create Prediction-Feedback Loops

Before any AI initiative, ask team members to predict outcomes: "How accurate will this LLM-generated summary be?" "How long will this prompt engineering task take?" Then measure. Over time, the gap between prediction and reality shrinks. This is the same principle behind iterative refinement in prompt engineering — and it works for self-assessment too.

Make Calibration Visible

AISA's assessment shows each candidate their predicted score alongside their actual score. This moment of confrontation — seeing the gap — is the single most powerful calibration tool we've observed. McKinsey's 2024 report on AI adoption found that organizations with formal AI skill assessments were 1.6 times more likely to capture value from AI initiatives. Part of that value capture comes from teams that accurately understand their own capabilities.

Normalize the Gap

The overall population overestimates by 17.8 points. This isn't a personal failing — it's a human universal, amplified by a domain where external benchmarks barely exist. Normalizing the gap ("everyone overshoots, here's by how much") reduces defensiveness and opens the door to genuine skill development.

For teams looking to build a structured baseline, AISA's AI readiness assessment provides both individual and team-level calibration data, including dimension-by-dimension breakdowns that pinpoint exactly where overconfidence concentrates.


Related reading: AI Certification for Product Managers — structured paths for PMs building verified AI skills.

Related reading: AI Engineer Certification: 5 Worth It [2026] — why engineers' calibration gap makes credential choice harder.

Related reading: AI News of the Week: Sol Price Cut (Aug 23) — this week's model updates and what they mean for AI fluency.

Frequently Asked Questions

Which professionals judge their AI skills most accurately?

In AISA's dataset of 694 self-predictions, founders are the most accurately calibrated group with an average overestimation of just +6.6 points (n=47), followed by product managers at +12.0 (n=30). Engineers overshoot by +14.8 (n=135), and students by +36.0 (n=47). Both founders and PMs share professional habits — constant estimation, scoping, and feedback — that build calibration as a transferable skill.

Why do people overestimate their AI skills?

The average candidate overestimates their AI fluency by 17.8 points (n=694), predicting a Proficient-tier score while actually landing in Developing. This happens because AI is a domain with few external benchmarks — most people gauge their skills by comparing themselves to peers or by equating daily AI usage with deep competence. Without structured feedback (like a scored assessment), there's no mechanism to correct the gap.

Does confidence predict AI competence?

Not reliably. In our data, students predicted the second-highest scores (64.6) but achieved the lowest actual scores (28.6, n=47). Engineers predicted the highest (70.7) and did score highest (55.9, n=135), but their +14.8 gap means confidence still significantly outpaces competence. The prediction gap — the distance between confidence and actual performance — is a more useful signal than confidence alone, especially for hiring and team planning decisions.

Learn more about how AISA assesses product managers.

Ozan Dagdeviren

Ozan Dagdeviren

Founder of AISA — the AI skills assessment platform used by professionals worldwide to measure, certify, and develop their AI fluency. More about AISA

AISA

Curious about your AI Fluency?

AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.

The Science Behind AISA

Metropolitan PoliceHarvard UniversityCrowdboticsE.S.E.

In 2026, Anthropic published the AI Fluency Index — the largest empirical study of AI fluency to date, analysing nearly 10,000 conversations. AISA covers 93% of the behaviours Anthropic identified as markers of AI fluency and goes even deeper with 4 additional dimensions. The U.S. Department of Labor's AI Literacy Framework (TEN 07-25) defines what every worker needs to know about AI — AISA covers 100% of its 25 sub-competencies.Read our analysis: Anthropic's AI Fluency Study & AISA · DOL AI Literacy Framework & AISA

AISA's framework is developed by a team with deep roots in tech, behavioural science, and AI product leadership — the rubric is informed by backgrounds spanning the Metropolitan Police, Harvard, Crowdbotics (Silicon Valley), and the European School of Economics.