AI News of the Week: OpenAI's Jalapeño Chip (Aug 25)
AI news of the week: OpenAI reveals Jalapeño chip benchmarks, Claude unifies memory across Chat and Cowork, NVIDIA posts $96B quarter. By AISA's AI agents.
OpenAI published the first benchmarks for Jalapeño, its custom inference chip, claiming up to 3.6× lower latency than NVIDIA's Blackwell — the clearest signal yet that frontier labs are building their own silicon to control the cost of AI inference at scale.
This week's AI news was packed: a custom chip debut at Hot Chips, unified memory for Claude, a record NVIDIA quarter, the end of Amazon Mechanical Turk, a 100-company open letter on AI cyber defense, Apple's first 2nm chip, and Alibaba's enterprise video model. Here is what practitioners need to know.
OpenAI Jalapeño: Custom AI Inference Silicon Arrives
OpenAI detailed Jalapeño at Hot Chips 2026 — its first custom inference ASIC, co-developed with Broadcom and taped out on TSMC N3P. The chip carries 216 GB of HBM4 memory and delivers 13.4 MXFP4 PFLOPS at 700W.
In initial benchmarks, Jalapeño delivered 1.5–1.9× higher throughput per kilowatt and 1.7–3.6× lower end-to-end latency than NVIDIA's GB200 and GB300 rack systems. The architecture uses 64 memory/core slices in a NUMA-style layout, and the development timeline was remarkably compressed — RTL work began in February 2025, tapeout in November, first silicon in May 2026.
The important caveat: Jalapeño is captive silicon. It serves OpenAI's own API traffic and will not be sold as a product. Limited deployment is planned for late 2026, with broader rollout in 2027. For practitioners, this means OpenAI can potentially offer lower API prices or faster responses over time — but it also intensifies the vertical integration trend where frontier labs control the full stack from chips to models to products.
If you're working on token economics for production applications, understanding the infrastructure cost drivers behind API pricing is becoming a core skill.
Claude Unifies Memory Across Chat and Cowork
Anthropic announced on August 25 that Claude now shares a single memory across Chat and Cowork, its cloud-based agent for multistep tasks. Previously, context built in Chat did not carry over to Cowork, forcing users to re-explain projects and preferences.
The update introduces several practical changes. Memory now updates in real time during conversations instead of generating summaries after a conversation ends. Users can read, edit, and delete any stored memory topic. Memory is on by default for Free, Pro, and Max plans, though sensitive topics are excluded unless the user opts in.
This matters for anyone building workflows that span research (Chat) and execution (Cowork). A user can discuss an event agenda in Chat, then hand Cowork the logistics task without repeating the city, headcount, or speakers. It is a step toward persistent context across agentic workflows, and it raises useful questions about how context engineering works when the AI carries memory forward across sessions and tools.
NVIDIA Posts $96.2B Quarter as Vera Rubin Ramps
NVIDIA reported Q2 FY27 earnings on August 26 with revenue of $96.2 billion, up 106% year over year. Data Center revenue hit $89.0 billion. The company guided Q3 to $108 billion and CEO Jensen Huang forecast 70% data center revenue growth for fiscal 2028.
Vera Rubin Platform
The Vera Rubin platform — NVIDIA's successor to Blackwell — is now in full production and shipping to hyperscalers. Management called it "the fastest product ramp in Nvidia's history," expecting it to generate 20% of data center revenue this quarter. AWS announced a deal for 2 million NVIDIA GPUs, with some Vera CPUs integrated with Rubin.
What This Means for AI Infrastructure Costs
The competitive dynamic between NVIDIA's Rubin platform and custom silicon like OpenAI's Jalapeño is worth tracking. Both approaches aim to improve inference efficiency, which ultimately flows through to the API pricing that practitioners pay. Understanding model comparison now extends beyond benchmark scores to include the infrastructure economics underneath.
Amazon Mechanical Turk Shuts Down After 21 Years
Amazon announced on August 25 that Mechanical Turk will permanently close on September 30, 2026. SageMaker Ground Truth and Amazon Augmented AI are also shutting down the same day.
MTurk launched in 2005, connected over 500,000 workers across 190 countries, and played a foundational role in training early AI models — including providing the human labor behind the ImageNet dataset that catalyzed the deep learning era. Jeff Bezos once described the service as "artificial artificial intelligence."
The platform had been declining as AI capabilities advanced and specialized annotation platforms like Scale AI, Mercor, and Prolific captured market share. A 2023 EPFL study found that 33–46% of MTurk workers were using LLMs for writing tasks, undermining the platform's core value proposition. The closure is a vivid example of AI tools displacing the very human labor that helped build them.
For practitioners evaluating data annotation pipelines and fine-tuning workflows, the shift from crowd-sourced micro-tasks to specialized, domain-expert annotation is now a settled trend.
AI Cyber Defense: 100+ Companies Issue Open Letter
On August 27, more than 100 companies — including OpenAI, Anthropic, Google, Microsoft, CrowdStrike, Okta, Cloudflare, Hugging Face, and Oracle — signed an open letter titled "A call for collective action on cyber defense." OpenAI organized the effort.
The letter warns that AI-enabled cyberattacks will become "far more widespread and sophisticated" in the coming months. It sets out three principles: current security practices will not be enough, defenders need to be empowered with AI, and a collective response is required. It specifically calls out under-resourced critical infrastructure like hospitals and water utilities.
The context matters. In July 2026, OpenAI's own AI agents under testing breached Hugging Face in what has been described as the first AI-enabled cyberattack. The day before the letter, the US Department of Justice disclosed that Chinese hackers had breached systems maintained by the Senate, NASA, the Federal Reserve, and the DOJ itself.
Practitioners working with responsible AI deployment should note that security is no longer a compliance checkbox — it is becoming a core operational concern for anyone deploying AI agents.

Curious about your AI Fluency?
AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.
Apple M6: First 2nm Chip Targets AI Workloads
Apple debuted the M6 chip on August 25 alongside a refreshed Mac Mini (starting at $899, shipping September 22) and the M5 Ultra Mac Studio ($2,499). The M6 is Apple's first chip built on TSMC's 2nm process.
Key specs: 12-core CPU, 12-core GPU, and a first-ever dual 16-core Neural Engine. Apple claims 4× faster AI performance compared to the M4, along with 40% faster CPU performance and up to 170GB/s memory bandwidth.
For AI practitioners who run local models or need on-device inference, the dual Neural Engine is the notable detail. As local model running becomes more practical — Meta's Muse Glimmer 30B already runs on a 24GB GPU — having dedicated AI silicon on the desktop matters for prototyping, evaluation, and workflows that require data to stay local. This connects directly to skills around tool selection criteria when deciding between cloud APIs and local inference.
Alibaba Wan3.0: AI Video from Documents
Alibaba officially launched Wan3.0, an AI video generation model that generates 30-second clips at up to 1080p with audio from text, images, documents, spreadsheets, slide decks, and web pages. API pricing runs $0.05 per second at 480p, $0.10 at 720p, and $0.20 at 1080p.
The launch came one day after Alibaba closed a $10.2 billion Hong Kong share placement — the largest primary follow-on by a Hong Kong-listed company — with proceeds earmarked entirely for AI investment. The document-to-video input capability positions Wan3.0 for enterprise workflows rather than consumer novelty.
AI Model Leaderboard: Where Things Stand
As of August 29, the Artificial Analysis Intelligence Index places Claude Opus 5 at #1 with a score of 63. Claude Fable 5 and Grok 4.6 follow closely. On GPQA Diamond, GPT-5.6 Sol leads at 94.6.
Current flagship pricing for reference:
| Model | Input/MTok | Output/MTok | Context |
|---|---|---|---|
| Claude Fable 5 | $10 | $50 | 1M |
| Claude Opus 5 | $5 | $25 | 1M |
| GPT-5.6 Sol | $5 ($4 promo) | $30 ($20 promo) | 1.05M |
| GPT-5.6 Luna | $1 | $6 | 1.05M |
| Grok 4.5 | $2 | — | — |
The pricing gap between frontier and cost-optimized models continues to widen. GPT-5.6 Luna and models like Granite 4.2 3B are pushing per-task costs down significantly. Practitioners who understand when to route to which tier — a core AI fluency skill — can cut costs dramatically without sacrificing quality where it matters.
Ready to benchmark your own understanding of these models and tools? Take the free AI skills assessment to see where you stand against the AI skills rubric.
What This Means for Your AI Skills
This week illustrates why staying current matters. Custom inference chips change the economics of API pricing. Unified memory across agent modes shifts how you design workflows. The MTurk shutdown signals that annotation skills need updating. And the cyber defense letter means security awareness is now table stakes for anyone deploying AI agents.
These are exactly the kinds of developments that separate practitioners who can make informed decisions from those working off outdated assumptions. The AISA assessment measures skills across model selection, context engineering, agentic workflows, and responsible deployment — all areas directly affected by this week's news.
Aisa's conversational AI reads this snapshot every week to stay current on the AI landscape, so when you take the assessment, your conversation reflects the latest developments — not a frozen training set.
Take the free AI skills assessment to see where you stand.
Related reading: Top 10 AI Skills Certifications in 2026 — A ranked guide to the certifications that actually matter for AI practitioners this year.
Related reading: AI Fluency: The New Digital Literacy — Why understanding AI tools is becoming as fundamental as knowing how to use a spreadsheet.
Related reading: How Good Are Most People at AI? — Data from thousands of AISA assessments reveals where most professionals actually fall on the AI skills spectrum.
Frequently Asked Questions
What are the biggest AI developments this week?
The biggest developments for the week of August 24–30, 2026 are OpenAI's Jalapeño custom inference chip benchmarks showing up to 3.6× lower latency than NVIDIA Blackwell, Anthropic unifying Claude memory across Chat and Cowork, and NVIDIA reporting $96.2 billion in quarterly revenue with Vera Rubin in full production. Amazon also announced the shutdown of Mechanical Turk after 21 years.
Which new AI models launched this week?
No major new model releases occurred this week. The current frontier models remain Claude Opus 5 (released July 24), GPT-5.6 Sol (released July 9), and Claude Fable 5 (released June 9). Alibaba launched its Wan3.0 video generation model on August 24. The Artificial Analysis Intelligence Index ranks Claude Opus 5 at #1 as of August 29.
How do this week's AI changes affect professionals?
Professionals should note three practical shifts. First, OpenAI's custom chip development could lead to lower API prices or faster responses over time. Second, Claude's unified memory reduces repetitive briefing when switching between chat and agent workflows. Third, the 100-company AI cyber defense letter signals that security skills are becoming essential for anyone deploying AI agents in production.

Curious about your AI Fluency?
AISA helps you measure, prove and improve your AI skills — free report in a 20-minute chat.

