AI Competency Framework: Build One [2026]
Learn how to build an AI competency framework that maps to measurable skills. Covers dimensions, role mapping, proficiency levels, and AI competency assessment.
An AI competency framework defines what "good at AI" actually means inside your organization — broken into dimensions, mapped to roles, and tied to measurable proficiency levels. Without one, you're stuck with self-reported confidence surveys and vendor certifications that test recall, not application.
Most frameworks fail because they conflate knowing about AI with knowing how to use it. This guide walks through building one that doesn't.
What Is an AI Competency Framework?
An AI competency framework is a structured model that defines the specific skills, behaviors, and proficiency levels required for effective AI use across roles in an organization. It answers three questions: what competencies matter, how good does someone need to be, and how do you measure it.
Think of it as the rubric your organization uses to move from vague statements like "we need to upskill on AI" to concrete, role-specific expectations. A good framework does three things:
It Separates Knowledge from Application
Knowing what a transformer architecture is and knowing when to switch from GPT-6 Astra to Claude Fable 5.1 for a specific task are different competencies. A useful AI skills framework captures both — but weights application more heavily, because that's where value shows up.
It Scales Across Roles
A designer, an engineer, and a founder all need AI fluency, but the shape of that fluency differs. The framework needs to accommodate role-specific expectations without becoming a separate document for every job title.
It Connects to Measurement
A framework that can't be assessed is a poster on a wall. Every competency in the model should map to observable behaviors that can be evaluated — ideally through something more rigorous than a multiple-choice quiz.
Why Generic AI Competency Frameworks Fail
The most common failure mode is conflating knowledge with application. A framework that asks "Can this person define retrieval-augmented generation?" is testing vocabulary. A framework that asks "Can this person decide when RAG is the right architecture for a given problem, and articulate the tradeoffs?" is testing competency.
According to the World Economic Forum's Future of Jobs Report 2025, 59% of employers expect to adjust hiring criteria to prioritize AI skills by 2030 — but the report also notes that most organizations lack clear definitions of what those skills are. That gap is where bad frameworks thrive.
The Academic Trap
Academic frameworks tend to organize competencies around AI subdisciplines: machine learning, natural language processing, computer vision. This makes sense for a computer science curriculum. It makes less sense for a product manager who needs to evaluate whether an AI feature is ready for production.
The Vendor Trap
Vendor-specific frameworks — "certified in Tool X" — test proficiency with a particular product, not transferable AI competency. When the tool changes (and it will — we've seen four major model releases in a single week), the certification becomes stale. The underlying skill of model comparison and selection doesn't expire the same way.
The Self-Report Trap
Self-assessment is the most common approach and the least reliable. Across 828 AISA assessments where candidates predicted their own scores, the average prediction was 62.2 against an actual average of 44.2 — an overestimation gap of 18 points. For Students, the gap was 35 points (predicted 66.7, actual 31.7). People don't know what they don't know, and frameworks built on self-report inherit that blindness.
The Five Dimensions of Applied AI Competency
A practical AI competency framework needs dimensions that reflect how people actually use AI at work — not how AI works in theory. AISA's model, validated against the U.S. Department of Labor's AI Literacy Framework (100% coverage) and Anthropic's AI Fluency Index (93% overlap), organizes competency into five dimensions. You can adopt these directly or use them as a reference architecture for your own.
Prompting & Communication (23% weight)
This covers the ability to formulate effective instructions, provide appropriate context, and iterate on outputs. It's not just "write a good prompt" — it includes context engineering, knowing when to use few-shot examples versus system-level instructions, and recognizing when a prompt strategy isn't working.
Across 1,738 AISA assessments, the population average for this dimension is 44.8 out of 100. Product roles score highest at 54.7; Students score 35.7.
Critical Thinking (22% weight)
The ability to evaluate AI outputs, detect errors, identify hallucinations, and maintain intellectual independence. This is the dimension that separates someone who uses AI from someone who uses AI well. It includes source triangulation, recognizing when a model is confidently wrong, and knowing what to verify.
Population average: 43.1. The gap between Founders (50.3) and Students (33.9) is substantial — roughly 16 points.
Technical Understanding (20% weight)
Not "can you build a model" but "do you understand enough about how models work to make good decisions?" This includes understanding context windows, token economics, training data cutoffs, and why a model might behave differently on the same prompt. For technical roles, it extends to API integration patterns and evaluation frameworks.
This is the lowest-scoring dimension across the population at 39.5. Even Engineering roles average only 51.0.
Workflow & Application (25% weight)
The highest-weighted dimension because it's closest to business value. Can someone integrate AI into their actual work? Do they know which tasks to delegate to AI and which to keep manual? Can they build multi-step workflows that combine AI tools effectively?
Population average: 48.0 — the highest of the five dimensions. Product roles lead at 58.7, suggesting that people who think in workflows naturally extend that thinking to AI.
Safety & Responsibility (10% weight)
Covers data privacy, bias awareness, appropriate use policies, and understanding the limitations of AI systems. Lower weight doesn't mean lower importance — it means the competency surface area is narrower. But it's the dimension with the most variance: Design roles average 36.2 while Founders average 49.9, a 13.7-point spread. (For more on this pattern, see Designers & AI: The Safety Blind Spot.)
How to Map AI Competencies to Roles
A framework that applies the same expectations to every role is a framework that helps no one. The data shows significant variation in where different roles are strong and where they're weak — and your framework should reflect that.
Start with Role Clusters, Not Individual Titles
Group roles by how they interact with AI, not by department. AISA data across 1,738 assessments shows clear clustering:
| Role Cluster | n | Composite | Strongest Dimension | Weakest Dimension |
|---|---|---|---|---|
| Engineering | 348 | 54.7 | Workflow (55.9) | Safety (46.8) |
| Product | 99 | 55.8 | Workflow (58.7) | Technical (46.2) |
| Founders | 175 | 55.6 | Workflow (57.2) | Safety (49.9) |
| Data | 37 | 51.0 | Workflow (51.5) | Technical (42.1) |
| Design | 41 | 49.8 | Workflow (53.7) | Safety (36.2) |
| Students | 140 | 35.5 | Prompting (35.7) | Safety (28.1) |
The spread between Engineering (54.7) and Students (35.5) is 19.2 points — nearly a full tier difference. A single competency bar for both groups would be meaningless.
Define Role-Specific Proficiency Targets
For each role cluster, set a target proficiency level per dimension. An engineering team might need Proficient-level Technical Understanding but only Competent-level Safety & Responsibility. A design team might need the inverse.
The key insight from the data: Workflow & Application is the strongest dimension across every role cluster. People learn to apply AI before they learn to think critically about it or understand its mechanics. Your framework should account for this — don't assume that high workflow scores mean high overall competency.
Account for the Confidence Gap
When setting targets, remember that self-assessment will skew high. Engineering roles overestimate by 14.1 points on average (predicted 68.5, actual 54.4). Founders are the most calibrated at 6.3 points of overestimation. Build your framework around assessed competency, not reported competency. For a deeper look at role-level data, see How Good Is My Team at AI?

Curious about your AI Fluency?
AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.
Setting Proficiency Levels That Actually Mean Something
Most frameworks use three to five levels with labels like "Beginner / Intermediate / Advanced." These are too coarse to be useful and too vague to be actionable.
Use Behavioral Anchors, Not Labels
Each level should describe what someone does, not what they know. AISA's persona system illustrates this with ten behavioral archetypes. Here's a simplified five-tier model you can adapt:
| Level | AISA Persona Equivalent | Composite Range | Behavioral Description |
|---|---|---|---|
| 1 — Novice | Bystander / Dabbler | 0–27 | Rarely uses AI tools. When they do, prompts are vague and outputs are accepted without review. |
| 2 — Developing | Copy-Paster / Sceptic | 28–45 | Uses AI for simple tasks but treats it as a search engine. Limited iteration. May distrust outputs without knowing how to verify them. |
| 3 — Competent | Enthusiast / Tactician | 46–65 | Integrates AI into regular workflows. Can iterate on prompts, recognizes obvious errors, and selects appropriate tools for common tasks. |
| 4 — Proficient | Conductor / Builder | 66–85 | Designs multi-step AI workflows. Evaluates outputs critically, understands model tradeoffs, and can teach others. |
| 5 — Expert | Architect | 86–100 | Architects AI systems and processes. Deep understanding of capabilities and limitations. Sets organizational AI strategy. |
From AISA's data, the population distribution tells you where most people actually land: 30.2% are Dabblers (avg score 26.2), 21.7% are Enthusiasts (avg 52.2), and only 4.3% are Architects (avg 87.4). If your framework assumes most people are at Level 3 or above, you're building for a population that doesn't exist yet.
Don't Over-Segment
Stanford's 2024 AI Index Report found that organizations with more than seven proficiency levels in their competency models saw lower adoption of the framework itself — people couldn't reliably distinguish between adjacent levels. Five tiers is the practical sweet spot: granular enough to track growth, coarse enough to be consistently applied.
Tie Levels to Business Outcomes
Level 3 (Competent) might mean "can independently complete AI-assisted tasks without supervision." Level 4 (Proficient) might mean "can design AI workflows that reduce team cycle time." These outcome anchors make the framework useful for managers, not just L&D teams.
Measuring Against the Framework: AI Competency Assessment
A framework without measurement is a wish list. The question is how you measure — and the options vary dramatically in reliability.
Self-Assessment: Fast but Unreliable
Self-report surveys are cheap and easy to deploy. They're also systematically wrong. The 18-point average overestimation gap from AISA's data isn't an outlier — Gartner's 2024 research on digital skills found that employees overestimate their proficiency by 20-30% on average across technology domains. Self-assessment is useful for gauging engagement and confidence, not competency.
Multiple-Choice Tests: Better but Incomplete
Quiz-based assessments can test knowledge reliably, but they can't test application. Knowing the definition of chain-of-thought prompting and knowing when to deploy it in a real workflow are different skills. Multiple-choice also has a ceiling effect — it can identify who doesn't know something, but it struggles to differentiate between competent and expert practitioners.
Conversational Assessment: Closest to Real Performance
The most reliable approach is assessing people in conditions that resemble actual AI use — open-ended, multi-turn, requiring judgment. AISA's AI skills assessment uses this model: a conversational AI facilitator explores competency across all five dimensions, while a separate AI evaluator scores independently against a rubric. This approach catches things that quizzes miss: Can someone recover when their initial approach fails? Do they verify outputs or accept them blindly? Can they articulate why they chose a particular strategy?
Combine Methods for Coverage
The practical recommendation: use self-assessment for initial baselining and engagement tracking, use a structured assessment like AISA for accurate measurement, and use project-based observation for ongoing calibration. No single method covers everything.
For teams looking to connect assessment results to training priorities, see AI Training Needs Assessment.
Template: A Starter AI Competency Matrix
Below is a starter matrix you can adapt. It maps the five dimensions to three role clusters with suggested proficiency targets. Adjust the targets based on your organization's AI maturity and strategic priorities.
| Dimension | Weight | Engineering Target | Product Target | All-Staff Minimum |
|---|---|---|---|---|
| Prompting & Communication | 23% | Level 4 — Proficient | Level 4 — Proficient | Level 2 — Developing |
| Critical Thinking | 22% | Level 3 — Competent | Level 4 — Proficient | Level 2 — Developing |
| Technical Understanding | 20% | Level 4 — Proficient | Level 3 — Competent | Level 1 — Novice |
| Workflow & Application | 25% | Level 4 — Proficient | Level 4 — Proficient | Level 2 — Developing |
| Safety & Responsibility | 10% | Level 3 — Competent | Level 3 — Competent | Level 2 — Developing |
| Composite Target | Level 4 (66–85) | Level 4 (66–85) | Level 2 (28–45) |
How to Use This Matrix
- Baseline your team. Run an AI competency assessment to see where people actually are — not where they think they are.
- Identify gaps by dimension. The data will likely show that Workflow scores are higher than Technical Understanding or Safety scores. That's normal — it matches the population pattern.
- Set 90-day targets. Moving someone from Level 2 to Level 3 in a single dimension is a realistic quarterly goal. Moving them two levels is not.
- Reassess quarterly. Competency isn't static, especially in a field where four frontier models can launch in a single week. Regular measurement keeps the framework honest.
For team-level assessment, consider running a cohort through the same assessment to get comparable data across the group. Individual scores are useful; distribution data is more useful.
Customization Guidance
This matrix is a starting point. Your organization might need additional dimensions (e.g., domain-specific AI applications) or different weight distributions. The principle to preserve: every cell in the matrix should connect to a measurable behavior, and every target should be validated against actual assessment data, not assumptions.
For role-specific data to inform your targets, AI Skills by Job Role: 2026 Data breaks down dimension scores across a broader set of functions.
Related reading: How Good Is My Team at AI? [2026 Data] — Benchmark your team's AI skills against 1,738 assessed professionals.
Related reading: AI Fluency Score: What It Measures [2026] — What the composite score actually captures and how it's calculated.
Related reading: AI Skills for Founders: 170 Assessed [2026] — Where founders overperform and where they have blind spots.
Frequently Asked Questions
What should an AI competency framework include?
An AI competency framework should include defined dimensions of competency (such as prompting, critical thinking, technical understanding, workflow application, and safety), role-specific proficiency targets, behavioral anchors for each level, and a measurement methodology. The framework should distinguish between knowledge and application — testing whether someone can use AI effectively, not just define AI concepts.
How many AI competency levels should I define?
Five levels is the practical sweet spot. Fewer than four levels makes it hard to track meaningful growth; more than seven levels leads to inconsistent classification, as evaluators struggle to distinguish between adjacent tiers. Each level should be anchored to observable behaviors — what someone does at that level — rather than abstract descriptors like "intermediate" or "advanced."
How do you assess AI competency?
The most reliable approach combines multiple methods. Self-assessment captures confidence and engagement but overestimates actual skill by an average of 18 points based on AISA data across 828 predictions. Structured conversational assessments — where candidates demonstrate competency through open-ended, multi-turn interactions — provide the most accurate measurement. Complement these with project-based observation for ongoing calibration. Avoid relying solely on multiple-choice tests, which measure recall but not application.

Curious about your AI Fluency?
AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.

