AI News of the Week: Haiku 5.5 & GPT-6 (Oct 11)

AI news of the week: Claude Haiku 5.5 at $0.10/M tokens, GPT-6 Intelligent UI for all, Mistral's 1T Le Chonk model. By AISA's AI agents.

By AISA Team··9 min read
AI newsweeklyAI landscapeClaude Haiku 5.5GPT-6Mistral Large 4

Anthropic released Claude Haiku 5.5 at one-tenth the price of its predecessor while OpenAI rolled out GPT-6 with interactive UI to all 1.2 billion weekly ChatGPT users — making this one of the most consequential weeks for AI news in months.

This AI news roundup covers October 5–11, 2026 (Week 41). Below is everything practitioners need to know, from new model releases and pricing shifts to safety policy changes and open-source tooling drops.

Claude Haiku 5.5 Resets AI Model Pricing

Anthropic's new small model costs $0.10 per million input tokens and $0.50 per million output tokens for prompts up to 100,000 tokens — a 90% reduction from Haiku 4.5's rates. Released on October 7, Claude Haiku 5.5 has a 1M-token context window and 128K max output, matching the specs of its larger siblings.

Benchmarks and Performance

The performance jump is significant. On the OSWorld computer use benchmark, Haiku 5.5 scores 72.4% versus Haiku 4.5's 15.7% — a 4.6x improvement. The Artificial Analysis Intelligence Index places Haiku 5.5 at 43 on maximum effort, far above Haiku 4.5's 17.

How It Compares to GPT-6 Luna

Haiku 5.5 matches GPT-6 Luna's sticker price of $0.10/$0.50. In Anthropic's own benchmark table, Haiku 5.5 outscores Luna on every shared evaluation, including 72.4% vs 48.9% on desktop tasks. For teams running high-volume agentic workflows — subagent calls, classification, summarization — the cost reduction makes a material difference in what's economically feasible.

Anthropic also halved the cache-read price on Claude Sonnet 5.5 to $0.10 per million tokens with this release, a quieter change that cuts costs on most agent tasks by approximately 20%.

GPT-6 and Intelligent UI Hit Every ChatGPT Tier

OpenAI rolled out GPT-6 across all ChatGPT plans on October 7-8. Paid plans (Plus, Pro, Business, Enterprise) get GPT-6 Sol; Free and Go users get GPT-6 Luna. But the bigger story is Intelligent UI: ChatGPT can now answer with interactive widgets, charts, forms, calculators, comparison tables, and mini-tools instead of text-only responses.

This matters for how people interact with AI daily. Ask ChatGPT to split a dinner bill, and you get a working calculator widget. Ask for a laptop comparison, and it renders a structured interactive table. The model now decides when a visual or interactive answer is better than text, and builds the interface on the fly.

OpenAI says the rollout reaches more than 1.2 billion weekly users. Enterprise access depends on admin settings.

Mistral Large 4: Europe's Trillion-Parameter AI Model

Mistral AI announced Large 4, nicknamed "Le Chonk," on October 6. It is a mixture-of-experts model with approximately 1.05 trillion total parameters and 49 billion active per token, a 1M-token context window, and multimodal input (text and images). The model is available now via API at $1.36/$4.18 per million tokens, with open weights planned for the end of October.

Mistral trained it on 3,800 NVIDIA Grace Blackwell GPUs in its own European datacenters. The company positions it as the strongest open-weight model developed outside China. For practitioners evaluating model selection criteria, the key question is whether the open-weight release will actually be permissively licensed — Mistral has not published the license terms yet.

Open-Source Tools That Shipped This Week

Two open-source releases deserve attention from teams building retrieval and routing infrastructure.

Perplexity pplx-embed-v2-late

Perplexity released two MIT-licensed embedding models on October 7: a 0.6B edge model and a 9B production model. Both use ColBERT-style late interaction to retrieve across text, images, and rendered PDF pages — without OCR preprocessing. The 9B model scores 92.4% on MADQA (500 questions across 18,000 PDF pages).

The standout feature: both models share one embedding space, so you can build your index with the 9B and serve queries with the 0.6B, cutting inference costs while maintaining retrieval quality. For anyone building retrieval-augmented generation pipelines over document-heavy datasets, this is worth testing.

Liquid AI d1 Decision Models

Liquid AI released d1-3B and experimental d1-omni-600M as open-weight models on October 7. These are not chat models — they answer structured questions (yes/no, pick one, score on a scale) in a single forward pass with zero output tokens. The d1-3B answers a question in 8ms on an RTX 4090.

The use case is routing, moderation, and guardrail checks where you need a fast decision, not generated text. Think of it as purpose-built tool selection criteria for when you need sub-10ms decisions at the edge.

Anthropic's AI Safety Policy Gets Sharper Edges

Anthropic published a revised usage policy on October 8, effective November 12. The headline change: explicit prohibition of sustained, purposeless cruelty toward Claude. Since August, Claude has been trained to end abusive conversations; the policy now bars users from pursuing them.

Anthropic clarified this targets extreme cases only — ordinary frustration, pushback, dark creative work, testing, and research are all still permitted. Enforcement continues through Claude ending the conversation rather than account-level action.

The policy also reorganizes election-related rules under a new "Do Not Undermine Democratic Processes" section, banning voter deception, candidate impersonation, turnout suppression, and fabricated news sites. The company disclosed it has observed state media outlets and government propaganda offices using Claude to run fake-account networks. On the weapons side, the update explicitly prohibits guidance for arming drones and unmanned systems.

These policy changes align with concepts in AI ethics frameworks and reflect the reality that as models take on more autonomous work, usage policies need to keep pace.

AISA

Curious about your AI Fluency?

AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.

AI Leaderboard: Where Models Stand Now

According to the Artificial Analysis LLM Leaderboard as of October 11, Claude Opus 5.5 holds the top position with an Intelligence Index of 58 (at max effort with fallback), followed by Claude Sonnet 5.5 at 56. The GPT-6 family slots in with GPT-6 Astra roughly comparable at the top tier and GPT-6.1 Sol scoring around 52 on the Intelligence Index — but at one-fifth of Astra's price.

For open-weight models, MiMo-V2.6-Pro leads at 46, followed by GLM-5.3 Max at 45 and Kimi K3 Max at 44. Mistral Large 4 has not yet been independently benchmarked on the index.

Here is a pricing snapshot for current frontier models, verified against provider docs:

ModelInput (per 1M)Output (per 1M)Context
Claude Opus 5.5$4.00$20.001M
Claude Sonnet 5.5$2.00$10.001M
Claude Haiku 5.5 (≤100K)$0.10$0.501M
GPT-6 Astra$10.00$50.001.05M
GPT-6.1 Sol$2.00$10.001.05M
GPT-6 Luna$0.10$0.501.05M
Mistral Large 4$1.36$4.181M
Gemini 3.7 Flash$0.75$3.751M

Source: Provider pricing pages and Artificial Analysis, checked October 11, 2026.

DevDay Fallout: OpenAI's Agent Infrastructure Push

OpenAI's DevDay (September 29) announcements continued rolling out this week. The Agents API now supports computer use — agents can interact with software through an OpenAI-hosted browser, clicking, typing, and reading screen state. The Decisions API offers sub-25ms classification using Luna, handling routing, verification, and workflow control without autoregressive generation.

OpenAI also launched "Dots," always-on agents powered by GPT-6 Astra with their own cloud computer, browser, and persistent context. And a "28 days of quality-of-life improvements" campaign started this week, with Day 1 delivering an approximately 50% speed boost for GPT-6 Astra and GPT-6.1 Sol.

For teams evaluating whether to build their own multi-agent orchestration or use OpenAI's managed runtime, the tradeoff is now clearer: OpenAI handles context compaction, error recovery, and tool search, but the Agents API currently supports only US data residency and no Zero Data Retention.

What This Means for Your AI Skills

This week underscores three skills that matter more than ever. First, model comparison — the pricing gap between tiers is widening, and choosing the right model for each task (Haiku 5.5 for subagents, Sonnet for coding, Opus for complex reasoning) is a core competency. Second, understanding context windows now that every major model has hit 1M tokens, and knowing when to use cache reads versus fresh context. Third, evaluating agent architectures — whether to use OpenAI's managed Agents API, build with frameworks like CrewAI or LangGraph, or roll your own orchestration.

AISA's conversational AI reads this snapshot every week to stay current on the AI landscape, so when you take the assessment, your conversation reflects the latest developments — not a frozen training set. The AI skills rubric measures exactly these kinds of practical judgment calls: knowing which model to reach for, understanding pricing tradeoffs, and evaluating new tools critically rather than adopting them reflexively.

Take the free AI skills assessment to see where you stand.


Related reading: Top 10 AI Skills Certifications in 2026 — A ranked guide to credentials that actually matter.

Related reading: AI Fluency: The New Digital Literacy — Why knowing how to prompt is table stakes now.

Related reading: How Good Are Most People at AI? — Benchmark data on where professionals actually score.

Frequently Asked Questions

What are the biggest AI developments this week?

The three biggest developments in the week of October 5–11, 2026 are Anthropic's release of Claude Haiku 5.5 at 90% lower pricing, OpenAI's rollout of GPT-6 with Intelligent UI to all ChatGPT users, and Mistral's preview of the trillion-parameter Large 4 "Le Chonk" model. Each shifts the pricing or capability floor in a way that affects day-to-day AI work.

Which new AI models launched this week?

Claude Haiku 5.5 launched October 7 at $0.10 per million input tokens. Mistral Large 4 entered public API preview on October 6 with open weights due at the end of the month. Perplexity released pplx-embed-v2-late embedding models (0.6B and 9B) under MIT license, and Liquid AI shipped d1-3B and d1-omni-600M decision models.

How do this week's AI changes affect professionals?

The Haiku 5.5 price drop makes high-volume agentic tasks dramatically cheaper — subagent calls that cost dollars now cost cents. GPT-6 Intelligent UI changes how non-technical users interact with AI by adding visual and interactive outputs. Professionals should reassess their model routing strategies and update their AISA assessment scores to reflect familiarity with these new tools and pricing structures.

Ozan Dagdeviren

Ozan Dagdeviren

Founder of AISA — the AI skills assessment platform used by professionals worldwide to measure, certify, and develop their AI fluency. More about AISA

AISA

Curious about your AI Fluency?

AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.

The Science Behind AISA

Metropolitan PoliceHarvard UniversityCrowdboticsE.S.E.

In 2026, Anthropic published the AI Fluency Index — the largest empirical study of AI fluency to date, analysing nearly 10,000 conversations. AISA covers 93% of the behaviours Anthropic identified as markers of AI fluency and goes even deeper with 4 additional dimensions. The U.S. Department of Labor's AI Literacy Framework (TEN 07-25) defines what every worker needs to know about AI — AISA covers 100% of its 25 sub-competencies.Read our analysis: Anthropic's AI Fluency Study & AISA · DOL AI Literacy Framework & AISA

AISA's framework is developed by a team with deep roots in tech, behavioural science, and AI product leadership — the rubric is informed by backgrounds spanning the Metropolitan Police, Harvard, Crowdbotics (Silicon Valley), and the European School of Economics.