AI News of the Week: GPT-6 Astra Arrives (Sep 06)

AI news of the week: GPT-6 Astra launches with computer use, Claude Fable 5.1 tops leaderboards, Claude proves Fermat's Last Theorem. By AISA's AI agents.

By AISA Team··9 min read
AI newsweeklyAI landscape

OpenAI released GPT-6 Astra on September 3 — its most capable model yet, featuring computer use and a new "recurrent depth" reasoning architecture — capping a week that also saw Anthropic's Claude Fable 5.1 claim the #1 spot on the Artificial Analysis leaderboard and Claude autonomously formalize a proof of Fermat's Last Theorem in 11 days.

This week's AI news delivered the densest 72 hours of frontier model releases since the August wave. Three major labs shipped new flagships, Anthropic published a landmark research result, and the copyright battle over training data escalated. Here is what practitioners need to know.

GPT-6 Astra: OpenAI's New Flagship AI Model

GPT-6 Astra launched September 3 as OpenAI's most powerful model, built for end-to-end computer use and complex multi-step work across coding, science, and professional tasks. It is available through the API as gpt-6-astra and is rolling out to ChatGPT paid tiers.

The headline specs: a 1.05 million token context window with 128K max output, priced at $10 per million input tokens and $50 per million output tokens — 2.5x the cost of GPT-5.6 Sol. There is a catch for heavy context users: prompts exceeding 272K input tokens trigger a surcharge of 2x input and 1.5x output rates across the entire request, not just the tokens past the threshold.

The most technically significant change is what OpenAI calls "recurrent depth." Astra routes tokens repeatedly through the same transformer layers, reasoning in latent space rather than producing readable chain-of-thought. The API returns only a paraphrased summary of reasoning, not the raw trace. Safety researchers have flagged this as a transparency concern — you can no longer inspect the model's actual reasoning steps. This matters for anyone building verification checklists or stakes-based review workflows.

On benchmarks, Astra scores 96.3% on OpenAI's long-context MRCR at the 512K-1M range (up from 73.8% for Sol) and 72.6% on OSWorld 2.0 for computer use tasks. It is the first OpenAI model rated "Critical" for cybersecurity under their Preparedness Framework, meaning advanced cyber capabilities are gated behind the Daybreak program for vetted organizations.

Claude Fable 5.1 Takes the AI Leaderboard Crown

Anthropic shipped Claude Fable 5.1 and Claude Mythos 5.1 on September 1, and Fable 5.1 now ranks #1 on the Artificial Analysis Intelligence Index with a score of 57, ahead of GPT-6 Astra at 55. Both models are the same underlying architecture — Mythos 5.1 has lighter safeguards for vetted cybersecurity and life-sciences organizations.

The pricing story matters more than the benchmarks for most teams. Headline rates hold steady at $10/$50 per million tokens, but Anthropic cut cache-read prices by 75%, from $1 to $0.25 per million tokens. For agentic workflows where the model repeatedly revisits the same codebase, tool definitions, or conversation history, Anthropic estimates this reduces effective costs by 25-45%. That is a meaningful shift for anyone running long-horizon coding agents or multi-step research pipelines.

Fable 5.1 posts a 73.4% on CursorBench 3.2.0 and a 31.4% on AutomationBench (up from 17.1% for Fable 5). The model is callable as claude-fable-5-1 across the Anthropic API, AWS, Google Cloud, and Azure. Teams upgrading from Fable 5 should note three documented breaking API changes in tool choice and thinking-block behavior.

Gemini 3.8 Flash and Meta Muse Spark 1.3

Google and Meta both released updated models on September 2, targeting the fast-iteration tier of the market.

Gemini 3.8 Flash

Google's third Flash release in six weeks, Gemini 3.8 Flash costs the same as 3.7 Flash — $0.75 per million input, $3.75 per million output — with those rates doubling on January 1, 2027. It is built on 3.7 Flash (not a new base model) and improves benchmarks by deliberately burning more thinking tokens on complex tasks. Google explicitly recommends staying on 3.7 Flash for efficiency-first workloads. The 3.8 Flash Cyber variant launched alongside it for security teams through Google's Fairwind program.

Meta Muse Spark 1.3

Meta's closed-weights frontier model shipped its fourth version since April, priced at $1.25/$4.25 per million tokens with a Contributor tier at roughly $0.10/$0.20 where Meta trains on your data. It posts 75.4% on DeepSWE 1.1 and 98.5% on long-context retrieval, with a 1M-token context window. Meta reports 20% fewer tool calls and 25% fewer tokens than Spark 1.2 for equivalent tasks. Open weights are still planned but have not shipped.

Claude Formalizes Fermat's Last Theorem — AI Research Milestone

The week's most striking research result: Anthropic announced that Claude autonomously produced the first complete, computer-checked proof of Fermat's Last Theorem in the Lean programming language. The run took 11 days, generated 13 million lines of Lean code, proved 29,500 intermediate theorems, and consumed roughly 6 billion output tokens.

The work used Prove2Me, an open collaborative platform built by Anthropic researcher Tianyi Peng's group at Columbia University. Dozens of Claude agents worked in parallel, navigating a directed acyclic graph of theorem statements. Kevin Buzzard, the Imperial College London mathematician who had been leading a community effort to formalize the same theorem, reviewed the artifact and endorsed it. This completes Freek Wiedijk's famous list of 100 formalization challenges — a benchmark that has been open for 20 years.

What this means practically: formal verification of mathematical proofs has gone from a multi-year project to an 11-day autonomous run. Anyone working on agent orchestration should study the architecture — the parallel agent approach with shared dependency graphs is a pattern that generalizes beyond mathematics.

AISA

Curious about your AI Fluency?

AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.

Sony Music Publishing and Warner Chappell Music filed a 48-page joint complaint against Anthropic and co-founders Dario Amodei and Benjamin Mann on August 29. The publishers allege Anthropic used "tens of thousands" of copyrighted musical compositions to train Claude without authorization, calling it "one of the largest and most blatant ongoing thefts of intellectual property in history."

The lawsuit seeks up to $150,000 per work willfully infringed and references Anthropic's $1.5 billion settlement with book authors from September 2025. This is now the fifth music-related copyright suit against Anthropic, following earlier cases from Universal, Concord, BMG, and Round Hill Music. Anthropic disputes the claims and says it will defend itself in court.

For practitioners, the expanding litigation reinforces the importance of understanding AI data privacy and the provenance of training data — particularly if you are building products that generate or reference copyrighted content.

Apple's New CEO and the AI Gap

John Ternus officially became Apple's CEO on September 1, succeeding Tim Cook after 15 years. He inherits what multiple analysts have described as the only large technology company without a frontier AI model of its own — Apple's Siri experience is powered by Google's Gemini. Ternus, a hardware engineering veteran who joined Apple in 2001, faces his first major public test on September 9 at the annual iPhone launch event, where Apple is expected to unveil a foldable device.

The leadership change matters for the AI landscape because Apple's massive installed base represents a distribution channel that frontier labs compete fiercely to access. Whether Ternus accelerates Apple's own AI development or deepens partnerships with existing model providers will shape how hundreds of millions of users interact with AI.

What This Means for Your AI Skills

This week highlights several skills that separate effective AI practitioners from casual users.

First, model selection just got harder. With Claude Fable 5.1, GPT-6 Astra, Gemini 3.8 Flash, and Muse Spark 1.3 all launching within 48 hours, the ability to evaluate tradeoffs across price, quality, transparency, and context handling — what AISA measures as model comparison skill — is now a weekly necessity, not an annual review.

Second, GPT-6 Astra's opaque reasoning changes how you verify outputs. If you cannot inspect the chain of thought, your hallucination detection and verification workflows need to adapt. The AI skills rubric covers exactly these evaluation competencies.

Third, the Fermat's Last Theorem result demonstrates that multi-agent orchestration is moving from experimental to production-grade. Understanding how to decompose complex problems across parallel agents is becoming a core skill.

AISA's conversational AI reads this snapshot every week to stay current on the AI landscape, so when you take the assessment, your conversation reflects the latest developments — not a frozen training set.

Take the free AI skills assessment to see where you stand.


Related reading: Top 10 AI Skills Certifications in 2026 — The definitive ranking of credentials that actually matter.

Related reading: AI Fluency: The New Digital Literacy — Why knowing how to use AI is becoming as fundamental as knowing how to use a spreadsheet.

Related reading: How Good Are Most People at AI? — What AISA's data reveals about the actual skill distribution across thousands of assessments.

Frequently Asked Questions

What are the biggest AI developments this week?

The three biggest developments are OpenAI's launch of GPT-6 Astra with its new recurrent depth reasoning architecture and computer use capabilities, Anthropic's Claude Fable 5.1 taking the #1 position on the Artificial Analysis leaderboard, and Claude autonomously formalizing a complete proof of Fermat's Last Theorem in Lean in 11 days. Sony and Warner also filed a major copyright lawsuit against Anthropic.

Which new AI models launched this week?

Four frontier models launched between September 1-3, 2026: Claude Fable 5.1 and Mythos 5.1 from Anthropic (Sep 1), Gemini 3.8 Flash and 3.8 Flash Cyber from Google (Sep 2), Muse Spark 1.3 from Meta (Sep 2), and GPT-6 Astra from OpenAI (Sep 3). All four target long-horizon agentic work and share 1M-class context windows.

How do this week's AI changes affect professionals?

Professionals need to re-evaluate their model choices given four new frontier options with different price-performance tradeoffs. GPT-6 Astra's opaque reasoning requires updated verification practices, Claude Fable 5.1's cache price cut makes long-running agents significantly cheaper, and the Fermat's Last Theorem result shows multi-agent orchestration is ready for complex autonomous work.

Ozan Dagdeviren

Ozan Dagdeviren

Founder of AISA — the AI skills assessment platform used by professionals worldwide to measure, certify, and develop their AI fluency. More about AISA

AISA

Curious about your AI Fluency?

AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.

The Science Behind AISA

Metropolitan PoliceHarvard UniversityCrowdboticsE.S.E.

In 2026, Anthropic published the AI Fluency Index — the largest empirical study of AI fluency to date, analysing nearly 10,000 conversations. AISA covers 93% of the behaviours Anthropic identified as markers of AI fluency and goes even deeper with 4 additional dimensions. The U.S. Department of Labor's AI Literacy Framework (TEN 07-25) defines what every worker needs to know about AI — AISA covers 100% of its 25 sub-competencies.Read our analysis: Anthropic's AI Fluency Study & AISA · DOL AI Literacy Framework & AISA

AISA's framework is developed by a team with deep roots in tech, behavioural science, and AI product leadership — the rubric is informed by backgrounds spanning the Metropolitan Police, Harvard, Crowdbotics (Silicon Valley), and the European School of Economics.