AI Landscape Snapshot — Week 29

Kimi K3 debuts as the largest open-weight model, Thinking Machines ships Inkling, and six labs now field frontier models.

By AISA Team··5 min read
ai-landscapeweeklyindustrykimi-k3gpt-5.6inklingopen-weight-models

The Frontier Widens: Six Labs Above 50

The most striking pattern this week isn't any single release — it's the shape of the leaderboard. Artificial Analysis reported that four frontier launches landed in eight days (Grok 4.5, GPT-5.6, Muse Spark 1.1, and Kimi K3), bringing the count of labs with a model scoring above 50 on the Intelligence Index from two in early June to six. The top three models — Claude Fable 5 (60), GPT-5.6 Sol (59), and Kimi K3 (57) — come from three different labs and span just three points.

Claude Fable 5 still holds #1, but its lead has narrowed from four points to one. The price of near-frontier intelligence, meanwhile, has collapsed.

Kimi K3: The Largest Open-Weight Model Ships

Moonshot AI released Kimi K3 on July 16 — a 2.8-trillion-parameter mixture-of-experts model with a 1-million-token context window. It's the largest open-weight model ever released, though only 16 of its 896 experts activate per token, keeping inference tractable.

K3 debuted at #3 on the AA Intelligence Index (score 57), ahead of Claude Opus 4.8 and GPT-5.5. It took #1 on Arena.ai's Frontend Code benchmark with 1,679 Elo points, beating Claude Fable 5 in blind developer testing. API pricing sits at $3/$15 per million input/output tokens. Full open weights are promised by July 27.

Two things to watch: the weights release date (which would make this the first open 3T-class model anyone can self-host), and Anthropic's accusation from February that Moonshot used 3.4 million Claude exchanges for distillation training.

Thinking Machines Releases Inkling

Mira Murati's Thinking Machines Lab shipped Inkling on July 15 — its first model, built from scratch in roughly nine months. Inkling is a 975B-parameter MoE (41B active), trained on 45 trillion tokens across text, image, audio, and video. It has a 1M-token context window and ships under Apache 2.0.

Thinking Machines is refreshingly direct about positioning: they state explicitly that Inkling is "not the strongest overall model available today, open or closed." It scored 41 on the AA Intelligence Index, making it the leading U.S. open-weight model. Benchmark highlights: SWE-bench Verified 77.6%, GPQA Diamond 87.2%, and HLE 46% with tool access.

Inkling is designed as a customizable base — the real product is Tinker, the company's fine-tuning platform. Day-0 support in transformers 5.14.0, SGLang, vLLM, and llama.cpp.

GPT-5.6: Now Generally Available

OpenAI's GPT-5.6 family hit general availability on July 9 after a government-reviewed limited preview. The three-tier structure — Sol ($5/$30), Terra ($2.50/$15), Luna ($1/$6) — replaces OpenAI's old single-model approach. All three share a 1.05M-token context window and 128K max output.

Sol leads the AA Coding Agent Index at 80 and scored 91.9% on Terminal-Bench 2.1 in Ultra mode. But Claude Fable 5 still leads on SWE-Bench Pro (80.0% vs 64.6%). Terra delivers GPT-5.5-class quality at half the price, which may be the bigger practical story for most teams.

AISA

The AI Fluency Assessment

Get Your Free AI Certificate in a 20-minute conversation with Aisa.

Free AI CertificationAI Fluency Score & PersonaAction Plan & Learning BoxGlobal Leaderboard

Grok 4.5: SpaceXAI's Cursor-Trained Contender

Grok 4.5 launched July 8 — SpaceXAI's first model co-developed with Cursor after the $60B acquisition agreement. Priced at $2/$6 per million tokens with a 500K context window, it scores 54 on the AA Intelligence Index. SpaceXAI positions it as "Opus-class, but faster, more token-efficient and lower cost." Independent testing confirms the cost story but shows it trailing Opus 4.8 on raw benchmarks. Not yet available in the EU.

Gemini 3.5 Pro: Still Waiting

Google's Gemini 3.5 Pro reportedly targeted July 17 for GA after missing its June deadline. Reports cite a full base-model rebuild. Reported specs — 2M context window, Deep Think reasoning, ~$15/$60 pricing — remain unconfirmed. Google has published no official benchmarks, pricing, or model card. Gemini 3.5 Flash ($1.50/$9, 1M context) remains the shipping 3.5 model.

Framework & Tool Updates

  • MCP spec heading toward a 2026-07-28 release candidate: stateless architecture, MCP Apps (UI iframes), Tasks extension, enterprise auth. This removes the biggest production scaling pain points.
  • Microsoft Agent Framework: GitHub Copilot SDK promoted to 1.0.0rc1. Declarative agents hit RC.
  • PyTorch 2.13: FlexAttention for Apple Silicon with ~12x sparse-attention speedup.
  • Transformers 5.14.0: Day-0 Inkling support.

Policy & Geopolitics

China launched the WAICO AI alliance on July 17. Five Eyes agencies published joint guidance on securing agentic AI in critical infrastructure. U.S. lawmakers are considering restrictions on Chinese AI model adoption by American companies. The open-weight debate is intensifying — prominent voices predict potential U.S. restrictions on frontier open-weight models within months.

What This Means for Practitioners

Multi-model routing is now table stakes. With six labs at the frontier and pricing ranging from $1/$6 (Luna) to $10/$50 (Fable 5), the teams getting the best results are routing tasks to the right model, not picking one model for everything. GPT-5.6's Sol/Terra/Luna tiers make this explicit, but the pattern applies across providers.

Open weights are catching up fast. Kimi K3 at #3 overall and Inkling as a strong customizable base mean teams can build on open models without the performance penalty that existed even six months ago. If the K3 weights land on July 27 as promised, self-hosting a frontier-competitive model becomes viable for well-resourced teams.

Test your actual workloads. Benchmarks show different winners depending on the task — Sol leads coding agents, Fable 5 leads SWE-Bench Pro, K3 leads frontend code, Grok 4.5 wins on cost-per-capability. The only benchmark that matters is your benchmark.

Assess where you stand with the AISA assessment, check the AI skills rubric to understand what skills matter most in this multi-model era, and explore role-specific guidance for developers building with these tools.

Ozan Dagdeviren

Ozan Dagdeviren

Founder of AISA — the AI skills assessment platform used by professionals worldwide to measure, certify, and develop their AI fluency. More about AISA

AISA

The AI Fluency Assessment

Get Your Free AI Certificate in a 20-minute conversation with Aisa.

Free AI CertificationAI Fluency Score & PersonaAction Plan & Learning BoxGlobal Leaderboard

The Science Behind AISA

In 2026, Anthropic published the AI Fluency Index — the largest empirical study of AI fluency to date, analysing nearly 10,000 conversations. AISA covers 93% of the behaviours Anthropic identified as markers of AI fluency and goes even deeper with 4 additional dimensions. The U.S. Department of Labor's AI Literacy Framework (TEN 07-25) defines what every worker needs to know about AI — AISA covers 100% of its 25 sub-competencies.Read our analysis: Anthropic's AI Fluency Study & AISA · DOL AI Literacy Framework & AISA