How Good Is My Team at AI? [2026 Data]
How good is my team at AI? Data from 1,710 assessments reveals team AI skills benchmarks by role, hidden gaps, and how to measure team AI fluency.
How good is my team at AI? If you're an engineering manager, L&D lead, or CTO, you've probably asked this question — and probably answered it with gut feel. Maybe you ran a survey. Maybe you looked at who's using Copilot. Maybe you just assumed your engineers are fine because they're engineers. Across 1,710 completed AI fluency assessments on AISA, the data tells a different story than most leaders expect.
The gap between what teams think they know and what they actually know is large, consistent, and invisible without measurement. This post breaks down what "good" looks like at the team level, where the hidden problems sit, and how to stop guessing.
The Question Every L&D Lead and CTO Is Asking
Every leader building an AI adoption strategy eventually hits the same wall: you can see tool usage metrics, but you can't see skill. Seat licenses tell you who has access. They don't tell you who's prompting effectively, who's verifying outputs, or who's blindly pasting AI-generated code into production.
The question "how good is my team at AI" is really three questions:
- Where does my team sit relative to the market? Not relative to what they claim in a self-assessment — relative to a scored, validated benchmark.
- How wide is the variance? A team average of 55 means something very different if everyone's between 50-60 versus half the team at 30 and half at 80.
- Where are the specific gaps? A team might be strong at prompting but weak on critical thinking and safety — which is arguably worse than being uniformly mediocre.
McKinsey's 2024 Global Survey on AI found that 72% of organizations had adopted AI in at least one business function, up from 55% the prior year. But adoption isn't proficiency. Having tools deployed says nothing about whether your people can use them well.
Why Self-Assessment Fails at Team Level
Across 800 AISA test-takers who predicted their own scores before the assessment, the average predicted score was 62.2 against an average actual score of 44.2 — an overestimation gap of 18 points. That's not a rounding error. That's people consistently placing themselves a full tier above where they actually land.
The gap varies by role, and the pattern is instructive:
| Role | Avg Predicted | Avg Actual | Overestimation Gap |
|---|---|---|---|
| Engineering (n=146) | 69.0 | 54.6 | 14.4 |
| Product (n=32) | 66.2 | 53.9 | 12.3 |
| Founders (n=65) | 59.9 | 54.0 | 5.9 |
| Students (n=57) | 66.8 | 31.0 | 35.8 |
Founders are the most calibrated — still off by nearly 6 points, but at least in the right ballpark. Engineers overestimate by 14.4 points. Students overestimate by a staggering 35.8 points. If you're relying on self-reported AI confidence surveys to plan your training budget, you're building on sand.
The Dunning-Kruger Problem in AI Skills
This isn't surprising if you've read the research on confidence calibration. People who know the least about a domain are the worst at estimating their own ability. In AI specifically, the problem is amplified because the tools feel easy to use. Anyone can type a prompt and get a response. The gap between getting a response and getting a good, verified, production-ready response is where real skill lives — and it's invisible to the person who doesn't know what they're missing.
What "Good" Looks Like for Team AI Skills
A team with strong AI skills scores in the Proficient tier (60-79 composite) across most dimensions, with no single dimension below Developing (28-59). The overall average across all 1,710 AISA assessments is 46.8 — solidly in the Developing tier. The median is 47. Most teams, if assessed honestly, would land here.
But the average obscures the real story. Role-level variation is significant:
| Role | n | Avg Composite | Tier |
|---|---|---|---|
| Product | 98 | 55.6 | Developing |
| Founders | 172 | 55.6 | Developing |
| Engineering | 340 | 54.8 | Developing |
| Data | 36 | 50.3 | Developing |
| Design | 41 | 49.8 | Developing |
| Students | 134 | 35.3 | Developing |
Product and Founders tie at 55.6. Engineering is close behind at 54.8. Every single role group lands in the Developing tier. Nobody — not even the strongest cohort — reaches Proficient on average.
Dimension-Level Breakdown by Role
The composite score hides dimension-level weaknesses. Here's where it gets actionable:
| Role | Prompting | Critical Thinking | Technical Understanding | Workflow | Safety |
|---|---|---|---|---|---|
| Engineering | 51.1 | 48.2 | 51.1 | 56.1 | 47.1 |
| Product | 54.5 | 50.9 | 46.0 | 58.5 | 48.0 |
| Founders | 51.7 | 50.4 | 49.3 | 57.3 | 50.0 |
| Data | 48.3 | 49.2 | 40.9 | 50.8 | 47.2 |
| Design | 49.8 | 43.6 | 42.3 | 53.7 | 36.2 |
| Students | 35.5 | 33.8 | 29.3 | 34.7 | 27.6 |
A few patterns jump out:
- Workflow is consistently the strongest dimension across every role. People are better at integrating AI into their processes than they are at thinking critically about its outputs.
- Safety is consistently weak, especially for Design (36.2) and Students (27.6). This is the dimension that covers responsible AI deployment, bias awareness, and data privacy — the stuff that creates organizational risk.
- Technical Understanding is the weakest dimension overall at 39.5 across all test-takers. Even engineers only hit 51.1. Data professionals — who you'd expect to be strongest here — score just 40.9.
For a deeper look at what the composite score actually measures, see AI Fluency Score: What It Measures.
What Proficient Actually Requires
To reach the Proficient tier (60-79), a team member needs to demonstrate more than basic tool usage. They need to show they can construct effective prompts with constraints, verify AI outputs against domain knowledge, understand model limitations, integrate AI into multi-step workflows, and apply appropriate stakes-based review to different use cases.
The AI Fluency Index provides the full benchmark data. The short version: Proficient is where someone becomes a net positive with AI tools rather than a liability.
The Hidden Problem: 30% of Your Team May Be Dabblers
Across all 1,710 assessments, 30.1% of test-takers are classified as Dabblers — the second-lowest persona, with an average score of just 26.3. That's the Emerging tier. These are people who have tried AI tools, maybe use ChatGPT occasionally, but lack the skills to use them effectively or safely in a professional context.
Here's the full persona distribution:
| Persona | Share | Avg Score | Tier |
|---|---|---|---|
| Bystander | 2.8% | 10.8 | Emerging |
| Dabbler | 30.1% | 26.3 | Emerging |
| Copy-Paster | 7.0% | 30.2 | Developing |
| Sceptic | 5.8% | 42.1 | Developing |
| Enthusiast | 21.8% | 52.2 | Developing |
| Tactician | 7.8% | 59.9 | Developing |
| Conductor | 4.1% | 70.9 | Proficient |
| Builder | 15.7% | 71.4 | Proficient |
| Architect | 4.3% | 87.5 | Advanced |
Let that sink in. If your team mirrors the general population — and without measurement, you have no reason to assume otherwise — roughly 1 in 3 people are Dabblers. Add Bystanders and Copy-Pasters, and nearly 40% of your workforce is in the bottom three personas.
Meanwhile, only 8.4% reach Proficient-tier personas (Conductor + Builder), and just 4.3% are Architects.
The World Economic Forum's Future of Jobs Report 2025 identified AI and big data skills as the fastest-growing skill demand globally, with 86% of employers expecting AI to transform their business by 2030. The gap between that expectation and the actual skill distribution above is where organizational risk lives.
Why Dabblers Are Dangerous
A Dabbler isn't someone who refuses to use AI. That's a Bystander or Sceptic — and at least those people aren't generating outputs they can't evaluate. A Dabbler uses AI tools but lacks the critical thinking and verification skills to catch errors. They're the ones who'll paste hallucinated code into a PR, send a client an AI-drafted email with fabricated statistics, or feed sensitive data into a consumer-tier tool without checking the data retention policy.
The Dabbler problem is invisible in usage metrics. They show up as active users. They look like adoption success stories. They're not.
For more on how this maps to training investment, see AI Training Needs Assessment.

Curious about your AI Fluency?
AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.
Five Signs Your Team Has an AI Skills Gap You Can't See
Most AI skills gaps don't announce themselves. They hide behind enthusiasm, tool adoption metrics, and the assumption that smart people figure things out. Here are five patterns we observe consistently across team assessments.
1. High Tool Adoption, Low Output Quality
Your team has Copilot seats, ChatGPT Enterprise, Claude — and usage is up. But code review catch rates for AI-generated code haven't changed. Customer-facing content still needs heavy editing. The tools are being used; the outputs aren't getting better.
This pattern typically maps to a team heavy on Dabblers and Enthusiasts (combined 51.9% of the general population). They're eager but lack the evaluation frameworks to assess what the AI gives them.
2. No One Can Explain Why They Chose a Specific Model
Ask your team why they used GPT-6 Astra instead of Claude Fable 5.1 for a specific task. If the answer is "it's what I always use" or a shrug, that's a Technical Understanding gap. The overall average for Technical Understanding across all AISA assessments is just 39.5 — the lowest of all five dimensions.
Model selection matters. With GPT-6 Astra now at $10/$50 per 1M tokens and Claude Fable 5.1 at the same price point but with 75% cheaper cache reads, choosing the wrong model for a given task has real cost implications at scale.
3. AI Outputs Go Straight to Production Without Verification
This is the safety gap in action. Design teams scored just 36.2 on Safety in AISA assessments. If your team doesn't have a consistent practice of verifying AI outputs — checking citations, testing generated code, reviewing for bias — you have a safety problem that won't show up until it causes an incident.
4. The Same Three People Answer Every AI Question
In most teams, AI knowledge concentrates in a few individuals. Everyone else defers to them. This creates a bus-factor problem and masks the actual skill distribution. Those three people might be Builders or Architects. The rest of the team might be Dabblers. You won't know until you measure.
5. Your "AI Training" Was a One-Time Workshop
A single workshop doesn't build fluency. It builds awareness at best. If your last AI training was a lunch-and-learn six months ago, your team's skills reflect whatever they've picked up informally since then — which, based on the data, probably isn't enough.
For a structured approach to identifying these gaps, the AI Skill Gap Analysis with AISA guide walks through the methodology.
How to Measure Team AI Fluency with AISA
Measuring team AI skills requires more than a quiz. Multiple-choice tests measure recognition, not application. Self-assessments measure confidence, not competence (as the 18-point overestimation gap demonstrates). You need a method that captures how people actually work with AI.
AISA for Teams uses a conversational assessment format: each team member talks to an AI facilitator that adapts to their responses, while a separate AI evaluator scores independently across 11 criteria in 5 dimensions. The conversation format makes it extremely difficult to game — AISA detects copy-paste, style shifts, and suspicious response speed.
Step 1: Baseline Your Team
Roll out the assessment to your full team. This takes about 25-30 minutes per person. You'll get individual scores across all five dimensions (Prompting & Communication, Critical Thinking, Technical Understanding, Workflow & Application, Safety & Responsibility), a persona classification for each person, and a team-level dashboard showing distribution and gaps.
Step 2: Identify the Distribution
The team average matters less than the shape of the distribution. A team of 20 with an average of 50 could be:
- Clustered: everyone between 45-55. Uniform skill level, easy to train as a group.
- Bimodal: half at 30, half at 70. You need different interventions for each group.
- Long-tailed: most people at 40-50, with two or three outliers at 80+. Your experts are carrying the team.
The persona breakdown makes this concrete. If 30% of your team are Dabblers and 20% are Builders, you know exactly who needs foundational training and who needs advanced workflow optimization.
Step 3: Map Gaps to Dimensions
Don't just look at composite scores. A team that's strong on Workflow (48.1 overall average) but weak on Safety (41.3 overall average) needs a very different intervention than a team that's weak across the board.
The dimension-level data also reveals role-specific training needs. Your designers might need Safety-focused training (36.2 average). Your data team might need Technical Understanding support (40.9 average). One-size-fits-all training wastes budget on skills people already have.
Step 4: Set Targets and Reassess
Pick a realistic target. Moving a team from Developing (28-59) to Proficient (60-79) is achievable with focused training over 3-6 months. Moving from Emerging to Proficient in the same timeframe is not.
Reassess quarterly. AI tools and capabilities change fast — Claude Fable 5.1, GPT-6 Astra, and Gemini 3.8 Flash all shipped in a single week this September. Skills that were sufficient three months ago may not be sufficient now. Continuous measurement catches regression and validates training ROI.
The assessment is validated against Anthropic's AI Fluency Index (93% overlap) and the U.S. Department of Labor's AI Literacy Framework (100% coverage), so you're measuring against established standards, not a proprietary rubric with no external validation.
From Guessing to Knowing: What Happens Next
The difference between teams that adopt AI effectively and teams that just have AI tools comes down to whether leadership treats AI fluency as a measurable competency or an assumed byproduct of tool access.
The data is clear: the average score is 46.8. Nearly a third of all test-takers are Dabblers. Every role group — including engineers and product managers — lands in the Developing tier on average. Self-assessments overestimate by 18 points.
You can keep guessing, or you can measure. The AISA for Teams assessment gives you the baseline, the distribution, and the dimension-level detail to make training investments that actually move the needle.
Founders who've been through the individual assessment already know what this looks like — 170 founders assessed so far, averaging 55.6 with the smallest overestimation gap of any role. If your leadership team has taken the assessment, extending it to the full organization is the logical next step.
Related reading: AI Skill Gap Analysis with AISA — how to turn assessment data into a targeted training plan.
Related reading: AI Training Needs Assessment — framework for mapping skill gaps to training investments.
Related reading: AI Fluency Score: What It Measures — deep dive into the 11 criteria and 5 dimensions behind the composite score.
Frequently Asked Questions
How do you benchmark a team's AI skills?
You benchmark team AI skills by assessing each member individually against a validated rubric, then analyzing the distribution — not just the average. AISA scores across 11 criteria in 5 dimensions (Prompting, Critical Thinking, Technical Understanding, Workflow, Safety), producing both individual scores and a team-level view. The overall benchmark across 1,710 assessments is a composite of 46.8, placing the average in the Developing tier (28-59).
What is the average AI score for a team?
Across 1,710 AISA assessments, the overall average composite score is 46.8 out of 100, with a median of 47. Role-level averages range from 35.3 (Students) to 55.6 (Product and Founders). No role group averages above the Developing tier. Most teams, if assessed honestly, will land somewhere in the 40-55 range depending on their role mix.
How many people on my team are AI beginners?
Based on AISA's data across 1,710 assessments, 30.1% of all test-takers are classified as Dabblers (average score 26.3), and an additional 2.8% are Bystanders (average score 10.8). Combined with Copy-Pasters at 7.0%, roughly 40% of the assessed population falls into the three lowest skill personas. Without measurement, there's no reliable way to know your team's actual distribution — self-assessments overestimate by an average of 18 points.

Curious about your AI Fluency?
AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.

