ai-powered-markdown-translatorArticle translated from fr to en with gpt-5.4-mini.
August 24 is as much a hardware day as a model day. NVIDIA puts out three infrastructure announcements in the same Hot Chips wave โ first silicon measurements for Vera Rubin NVL72, Groq 3 LPX entering production, and opening the platform to third-party chips via NVLink Fusion. Alibaba launches Wan 3.0, which generates 30 seconds of native video in a single pass. And two non-American labs appear on the same day alongside national institutions: Mistral with HUMAIN in Saudi Arabia, Sakana AI with the Japanese Ministry of Defense. Added to that are the general availability of enterprise authorization for MCP connectors at Anthropic, GPT-5.6 arriving in Kiro, and a series of announcements on voice agents.
Wan 3.0 generates 30 native seconds in a single pass
August 24 โ Alibaba launches Wan 3.0, the new generation of its video model. The most visible shift is in duration: the model produces 30 native seconds in a single pass, without stitching successive shots together (stitching). Until now, going beyond about ten seconds meant generating several segments and then joining them, with the coherence drift that this causes in lighting, textures, and faces.
The second axis is multimodal input. Wan 3.0 accepts up to 20 reference assets and analyzes documents and web pages in addition to images, audio, and video โ what the official page calls omnimodal creation (Omni-Creation). In practice, you feed it a brand folder rather than a text description.
The model also generates native synchronized audio: voices, effects, and ambience are produced at the same time as the images, not added afterward. On the still-image side, Wan 3.0 merges up to 9 images into a single composition and generates 12 sequential images with a constant style and subject, which explicitly targets storyboarding and ad variations.
Two less spectacular capabilities matter in professional use: long-text rendering across 12 languages, covering charts, formulas, and infographics where video models usually fail, and precise color control, which is essential for matching brand guidelines.
The ecosystem followed the same day: Wan 3.0 was available on day zero at fal, Leonardo.Ai, Scenario, RunningHub, Venice, TapNow, DeeVid, Pixmax, and via the Pika API Club.
| Announced capability | Value for Wan 3.0 |
|---|---|
| Native duration | 30 s in one pass, without stitching |
| Reference assets | Up to 20, with analysis of documents and web pages |
| Audio | Synchronized, generated natively with the images |
| Text rendering | Long text across 12 languages (charts, formulas, infographics) |
| Image fusion | Up to 9 images in one composition |
| Image sequences | Up to 12 coherent sequential images |
๐ Wan 3.0 โ official page
NVIDIA measures up to 30x more work per watt on Vera Rubin NVL72
August 24 โ At the Hot Chips conference, NVIDIA publishes its first silicon performance measurements for Vera Rubin NVL72. The headline figure is aimed directly at data-center efficiency: up to 30x more throughput per megawatt than the GB300 NVL72 generation on the DeepSeek V4 Pro model, and a cost per million tokens up to 35x lower.
The methodology matters as much as the result. NVIDIA excludes classic inference measurements, calibrated on sequences of 1,000 to 8,000 tokens: in an agentic session, context accumulates and reaches hundreds of thousands of tokens as input. The measurements were therefore taken with SemiAnalysis AgentX, a workload of recorded real-world agentic coding sessions, with tool calls and subagents preserved. Two caveats: the figures are awaiting SemiAnalysis review and do not yet include the contribution of the Vera CPU.
The rationale rests on an OpenRouter statistic: an agentic workload consumes 15x more tokens than a simple chat request. For an AI factory constrained by energy, the issue is stated bluntly โ throughput per megawatt determines revenue, and cost per million tokens determines margin. DSX MaxLPS technology also makes it possible to provision up to 40% more GPUs at the same megawatt budget.
The gains come from a stack of inference optimizations: disaggregated service separating context processing from response generation, distributed KV cache with staged offloading, fused CUDA kernels such as MegaMoE, and NVFP4 quantization that compresses weights to 4 bits. The full platform includes seven chips, including Rubin GPUs.
| Measured comparison | Announced gain |
|---|---|
| Vera Rubin NVL72 vs GB300 NVL72 (throughput per megawatt) | up to 30x |
| Vera Rubin NVL72 vs GB300 NVL72 (cost per million tokens) | up to 35x cheaper |
| GB300 NVL72 vs Hopper (throughput per megawatt) | up to 15x |
| DSX MaxLPS | +40% GPUs at equal megawatt budget |
๐ Vera Rubin NVL72 and agentic workloads
Groq 3 LPX enters full production, Nebius and SpaceXAI adopt Vera Rubin
August 24 โ Second part of the same Hot Chips wave: the NVIDIA Groq 3 LPX rack-scale system moves into full production, extending Vera Rubin NVL72.
The reference figure comes from an Artificial Analysis benchmark on Gemma 4 31B: 3,400 output tokens per second at a long context of 100,000 tokens, or 4x the nearest alternative platform. The problem it targets has a precise name โ decoding latency: an agent generates its response token by token, and the slightest per-token delay multiplies along the workflow chains.
The division of roles is explicit: Rubin GPUs handle context at scale, while LPX accelerates latency-sensitive decoding. A rack can bring together 256 LP30 accelerators connected by direct chip-to-chip links.
Three adopters are cited. Nebius is the first AI cloud to integrate Groq 3 LPX, in its Token Factory; CoreWeave is deploying Spectrum-X Multiplane in production to interconnect its Vera Rubin racks; and SpaceXAI is adopting NVIDIA Vera CPUs for orchestration, tool calling, and code execution behind agentic AI, with the intention of building its architecture around Vera Rubin, from terrestrial data centers to satellites in orbit.
On networking, exceeding the largest current clusters has until now required a third layer, costly in cabling, optics, power, and latency. Spectrum-X Multiplane instead splits each server connection into multiple independent paths, or planes, each operating its own two-level network.
| Measured element | Announced value |
|---|---|
| Groq 3 LPX on Gemma 4 31B, 100k context | 3,400 output tokens/s, 4x the nearest platform |
| Groq 3 LPX rack | up to 256 LP30 accelerators, chip-to-chip links |
| Spectrum-X Multiplane | up to 512,000 GPUs across two levels |
| Loss of one plane out of eight | about 90% bandwidth retained |
| Hardware recovery | 11x faster than software balancing, 1.6x output |
๐ Groq 3 LPX, Spectrum-X, and NVLink Fusion
NVLink Fusion connects custom chips to the NVIDIA platform
August 24 โ NVIDIAโs third article of the day: NVLink Fusion brings custom chips (XPU) into a 72-accelerator NVLink scaling domain. The sixth generation of NVLink shows end-to-end latency 3x lower than solutions based on standard Ethernet, with 10x higher packet throughput; the roadmap announces domains of up to 1,152 accelerators and co-packaged optics. NVLink-C2C connects XPUs to Vera CPUs or other CPUs with up to 6x the energy efficiency of a PCIe interface.
The argument is one of decoupling: planning an AI factory โ power, building, cooling, racks, networking โ starts long before the accelerator mix is fixed, and a data center locked to a single chip becomes a schedule risk. Adopters reuse the MGX rack architecture and its supply chain. Named partners: Intel, MediaTek, GUC, QCT, and Annapurna Labs.
Mistral and HUMAIN sign for sovereign AI in Saudi Arabia
August 24 โ Mistral announces a strategic collaboration with HUMAIN, the Saudi artificial intelligence company: compute infrastructure, advanced model development, and deployment of solutions in the kingdom and the region. According to Mistral, the commitment amounts to several hundred million euros.
The initial focus areas are cybersecurity and voice, with frontier models strong in Arabic. Mistral will explore using HUMAINโs datacenters for its local compute needs, and the two partners will jointly target the kingdomโs regulated industries. The announcement extends a sequence begun this summer โ an expanded partnership with Microsoft, then European Compute Units in early August โ and the press release notes that these are forward-looking statements, subject to later commercial agreements.
Control is becoming the defining topic in enterprise AI adoption. Across the world, and especially in regulated industries, organizations want the benefits of AI without giving up control over their sensitive data and systems. โ @MistralAI on X
๐ @MistralAI announcement ยท Mistral post
Sakana AI lands a contract from the Japanese Ministry of Defense
August 24 โ Sakana AI makes public a contract signed on July 29 with the Japanese Ministry of Defense (้ฒ่ก็), covering the study and demonstration of AI functions needed for integrated analysis operations. The project targets analysts at the Intelligence Directorate, across three areas: collection, analytical capabilities, and document management. The component being mobilized is the labโs AI agent technology.
This is not its first defense contract: in March 2026, the Tokyo-based lab had already been selected for the core technologies of command-and-control systems. For a company whose public identity rests on its open models, the trajectory is worth noting โ the blog now has a dedicated defense and intelligence section, and Sakana AI is hiring for this team.
ใSakana AIๆ ชๅผไผ็คพใฏใไปคๅ8ๅนด7ๆ29ๆฅใ้ฒ่ก็ใจใ็ทๅๅๆๆฅญๅใซๅฟ ่ฆใชAIๆฉ่ฝใฎ่ชฟๆปใปๅฎ่จผใใซ้ขใใๅฅ็ดใ็ท ็ตใใใใพใใใใ โ Sakana AI, official post
๐ Sakana AI โ integrated analysis for Defense
Anthropic makes enterprise authorization for MCP connectors generally available
August 24 โ Enterprise-managed auth for MCP connectors moves to general availability on Claude Team and Claude Enterprise, after an open beta in June. The problem it solves is concrete: until now, connecting an MCP connector required two steps โ the administrator enabled it for the organization, then each user went through an OAuth flow for every tool. The administrator now authorizes once, via the company identity provider, and users find their tools already connected.
One point matters to security teams: an administrator can require that a connector always go through the identity provider, preventing a personal account from being linked to a work tool. Ten MCP providers are supported at launch โ Asana, Atlassian, Canva, Datadog, Figma, Granola, Linear, Notion, Slack, and Supabase โ with Okta as the only IdP for now. The mechanism is not proprietary: this is the first implementation of the Enterprise-Managed Authorization extension of the Model Context Protocol.
๐ @ClaudeDevs announcement ยท Anthropic post
Boris Cherny recognizes cybersecurity refusals as a bug
August 24 โ The day after his remarks on the three capability tiers, the head of Claude Code answers two recurring questions from security developers. The first is suspicion of a privileged internal model: the team uses exactly the same Fable as end users.
The second is more direct: refusals triggered by cybersecurity topics are acknowledged as a bug, in terms that do not try to soften it, and the team is working to reduce them. That is useful information for practitioners of defensive security or CTFs, who regularly run into refusals on legitimate requests. No timeline is given, but the message extends the August update on Fable 5โs biology safeguards, which already aimed to reduce false positives.
We use the same exact Fable. Cyber security refusals suck, and we are working on reducing them. More to come. โ @bcherny on X
Anthropic Starts Automatically Maintaining Its Own Apps
August 23 โ Pushed to clarify what he means by coding, Boris Cherny draws a simple distinction: coding is the act of writing code, engineering is everything else on top of that. He acknowledges that the split between tactical work and strategic work is another valid lens.
The end of his message is what brings the new information. He says he sees parts of strategic work beginning to be automated, and gives a concrete example: Anthropic is starting to automatically maintain its apps, and a growing number of customers would be doing the same. The nuance matters โ this is no longer about generating code on demand, but about letting agents handle the routine maintenance of production applications: dependency updates, bug fixes, adaptations to API changes. The verb used describes an ongoing rollout rather than a settled fact, and no figures, named apps, or exact scope are given.
๐ @bcherny thread on coding and engineering
An Automated Weekly Sales Digest with Claude Code
August 24 โ Anthropic publishes a post about an internal use of Claude Code outside software development: a field marketer ัะฐััะบะฐles how she replaced the fifteen-minute Monday morning stand-up with a personalized weekly digest, delivered in a Slack message to each salesperson.
The architecture is simple: Claude is connected to BigQuery via MCP, with the warehouse serving as the source of truth for marketing data. The method is more instructive than the technique: the pilot started with a single volunteer team, and every correction raised by recipients was turned into an explicit prompt rule. By the end of the first week, it contained nine content rules, each traceable to a specific piece of feedback. Since each message is composed from the recipientโs account list, no two messages are identical; BDRs (business development representatives) then asked for their own variant.
๐ A personalized weekly digest with Claude Code
GPT-5.6 Arrives in Kiro, AWSโs Development Agent
August 24 โ The full GPT-5.6 family โ Sol, Terra, and Luna โ is now available in Kiro, AWSโs software development agent. The main selling point is not raw power but price/performance: on Terminal-Bench 2.1, a test run jointly by OpenAI and AWS, GPT-5.6 Terra completes successful tasks at about 82% lower cost.
The explanation given is Kiroโs method, spec-driven development: the tool turns a high-level intent into requirements, technical design, and executable tasks, anchoring the model in a precise context from the start and reducing false leads. OpenAI details the intended uses: structured implementation plans, multi-step coding tasks, repository context, and checkpoints before applying changes. For anyone following the code-agent ecosystem, the announcement mainly confirms that the GPT-5.6 lineup is spreading beyond OpenAI surfaces.
๐ GPT-5.6 in Kiro
Codex Gains a Developer Mode with Access to the Chrome DevTools Protocol
August 24 โ The ChatGPT and Codex changelog gets a wave of new features centered on the agent-driven browser. The main one is a developer mode for Browser use, in both Chrome and Codexโs built-in browser: it grants controlled access to the Chrome DevTools Protocol, the interface used by Chromeโs development tools. The agent can therefore profile performance and debug network traffic, console output, and page state, instead of being limited to what it reads on screen.
The same protocol also powers speed: Browser use becomes up to twice as fast, thanks to DOM snapshots that reduce back-and-forth with the browser. Scheduled automations now respect the selected approval mode, fixing an annoying mismatch for unattended runs.
| Changelog item | Announced scope |
|---|---|
| Developer mode for Browser use | Chrome and Codex built-in browser, controlled CDP access |
| Faster Browser use | Up to 2x faster (CDP optimizations and DOM snapshots) |
/init command | Compose the app, behavior identical to the Codex CLI |
| Computer Use | Businesses outside the EEA, UK, and Switzerland |
| Computer Use controls on Windows | Configurable application by application |
๐ ChatGPT and Codex changelog
The Gemma 4 Good Challenge Winners
August 24 โ Google publishes the results of its Gemma 4 Good Challenge, a Kaggle competition that invited developers to tackle real-world problems with its open models: more than 1,600 projects submitted in six weeks. The challenge was less about model quality than about deployment in constrained environments, with participants using LiteRT, Cactus, Ollama, llama.cpp, and Unsloth to run on ordinary hardware.
What the winning projects have in common is offline operation and privacy by design. GEM-4 takes first place with a robotic assistant for older and disabled people: a Gemma 4 31B annotates training video sequences, and a fine-tuned E2B controller translates observations and instructions into movement. Among the special prizes, Gem-Care fine-tunes Gemma 4 E2B to reconstruct non-standard speech and cuts the word error rate from 32.7% to 19.0%; PreVillage provides conversational routing in romanized Nepali on a Raspberry Pi 5 at 7.5 tokens per second.
๐ The winning projects of the Gemma 4 Good Challenge
ADK Natively Evaluates Voice Agents
August 24 โ Googleโs Agent Development Kit now supports evaluating live and voice agents, in the same loop as text agents. The principle: a user simulator speaks its turns, synthesized by Gemini TTS and then streamed to the agent, whose spoken responses are scored automatically. Live mode is enabled simply by the presence of a live_model_config block; without it, the same test cases run in text mode.
The test cases come in two styles: scenario, where a persona improvises its turns according to a plan, and fixed conversation, scripted word for word. Scoring uses LLM judges driven by natural-language rubrics, written once and then applied to the whole suite. Execution happens via adk eval or AgentEvaluator, which makes it possible to insert voice evaluations into a CI/CD pipeline. ADK Web adds a toggle between standard mode and live mode, rebuilds the transcript from the audio stream, and attaches a playable audio clip to each turn.
๐ Evaluating voice agents in ADK
Grok Voice Think Fast 2.0 Takes the Lead in the Speech-to-Speech Index
August 24 โ xAI announces that Grok Voice Think Fast 2.0 has taken first place in Artificial Analysisโs Speech-to-Speech index, which measures less the quality of the delivered voice than an agentโs ability to reason about what it hears, solve a real customer problem, and complete a task by calling tools.
The model is not new today: its launch post dates back to July 29. What xAI is sharing here is the ranking and production numbers. At Starlink, a subsidiary of the same group, Grok Voice handles more than 15,000 incoming support and sales calls per day and closes more than 3,000 orders per week, combining voice and chat. The Think Fast family runs its reasoning in parallel with speech synthesis, avoiding paying for intelligence in latency.
| Benchmark | Think Fast 2.0 | GPT-Realtime-2.1 (High) | Gemini 3.1 Flash (High) |
|---|---|---|---|
| AA Speech-to-Speech Quality Index | 82.9% | 79.1% | 69.5% |
| ฯ-voice Bench (agent performance) | 56.5% | 45.7% | 37.7% |
| Time to first sound | 0.70 s | โ | 2.98 s |
๐ @SpaceXAI announcement ยท x.ai post
ElevenLabs Brings Its Entire API into the Terminal
August 24 โ ElevenLabs releases version 1 of its CLI, exposing the full API in the terminal. The announcement is aimed equally at two audiences: developers and code agents.
That dual target shows up in the technical choices. The commands are documented to be discoverable, agent skills are built directly into the tool, and the output is available in structured JSON โ three conditions for an agent to control the tool reliably rather than interpret free text. Dry run mode completes the package: it lets you preview the result of an operation before executing it, a useful safeguard when an autonomous agent is making the call. The release follows the official MCP server launched by the vendor on August 17.
๐ ElevenLabs CLI v1
MiniMax Opens 14 Days of Unlimited Access on GMI Cloud
August 24 โ From August 24 to September 6, MiniMax is opening unlimited access to its models on GMI Cloud: M3 and M2.7 on the language side, Speech 2.8 and Music 3.0 on the audio side.
What stands out goes beyond the commercial promotion. On August 21, MiniMax had only issued an enigmatic announcement hinting at the arrival of an M3 model, without specs or a date. Three days later, M3 is explicitly available in a dated offer, alongside two audio models whose version numbers had not appeared in the vendorโs previous communications. No benchmark or technical specification accompanies the announcement: this is a release, not a product sheet.
๐ 14 days of unlimited access on GMI Cloud
Briefs
- Nemotron 3.5 Lightning enters the top 4 of PinchBench โ NVIDIAโs model ranks in the top 4 open-weight models with a 86.4% average success rate on standardized OpenClaw agent tests; Nemotron 3 Ultra keeps first place. ๐ @NVIDIAAI tweet
- Wan 3.0 is already available via the Pika API Club โ The aggregator, launched on August 5, provides access to more than a hundred models through a single API. The day-zero addition extends the logic already seen with Seedance 2.5 on August 17. ๐ @pika_labs tweet
- Runway extends Ruby to all models on its platform โ Outputs from Seedance 2.5, Gen-4.5, or MiniMax H3 can now be converted to 16-bit EXR or to ProRes and 10- and 12-bit HEVC, without a separate color pipeline per generator. ๐ @runwayml tweet
- Runway opens an in-person hackathon in San Francisco โ September 30, focused on building agents and applications; applications go through hackathon.runway.com. ๐ @runwayml tweet
- According to Boris Cherny, prompt injection turned out to be solvable โ He continues on bug-free code generation and cites the precedent: the combination of alignment training and mechanistic interpretability eventually managed to solve the problem, after many years. ๐ @bcherny thread
- GitHub webinar on August 25 about the Copilot app โ Live demo from 6 p.m. to 7 p.m. CEST with Pierce Boggan and James Clancey; the registration page details Agent Merge, which handles rebasing, reviews, and checks all the way to merge. ๐ @github tweet
- GitHub highlights season 2 of its podcast โ Presentation video of the hosts and their favorite episodes from the previous season, with no product announcement. ๐ @github tweet
- Grok 4.6 at half price in the Nous Research portal โ 50% discount for one week, usable with the Hermes open-source agent or through an existing SuperGrok or X Premium+ subscription. ๐ @SpaceXAI tweet
- Codex Live Build highlights voice-driven control โ Live session on X with creator Alex Finn, focused on keyboard-free work in Codex on desktop and mobile. Demo only, with no feature announcement. ๐ @OpenAIDevs tweet
- HeyGen publishes a prompt guide for real-estate videos โ Prompt writing for cinematic-quality avatar videos, continuing its real-estate series. ๐ @HeyGen tweet
- The 110th anniversary of U.S. national parks with Maps, Search, and Gemini โ Google combines its search trends (beginner hiking queries are up 165% month over month) and the use of AI Mode, Ask Maps, and Gemini Live in travel planning. No new feature. ๐ blog.google post
What it means
After several days centered on software use cases, hardware is back in the spotlight โ and what it measures has changed. NVIDIA explicitly sets aside inference benchmarks calibrated on sequences of 1,000 to 8,000 tokens in favor of real recorded agentic coding sessions, with growing context, tool calls, and sub-agents included. The unit of account follows the same shift: it is no longer raw throughput but work per megawatt, and cost per million tokens. The constraint is no longer available silicon, but the energy you can supply to it. The GPU/LPU split says the same thing from another angle: context processing on one side, latency-sensitive decoding on the other, because agentic inference has ceased to be a homogeneous workload.
The second signal of the day is platform openness. NVLink Fusion connects chips designed by others, SpaceXAI takes Vera CPUs, and Intel, MediaTek, and Annapurna Labs appear in the partner list. The business argument is about timing: an AI factory is planned โ power, building, cooling, network โ long before the accelerator mix is settled, and selling the rack, the network, and the tooling rather than just the accelerator itself amounts to becoming indispensable even to those who make their own silicon.
More unexpectedly, sovereignty takes center stage in two announcements on the same day. Mistral announces a collaboration with HUMAIN for Arabic-language models and the use of Saudi datacenters; Sakana AI signs with the Japanese Ministry of Defense for intelligence analysis. Two non-American labs, two national institutions, one logic: what Mistral calls control and what the Japanese government calls informational power both refer to the ability to run training and inference within a chosen jurisdiction. For Sakana, whose public identity rests on open work, the dedicated defense track and the associated hiring point to a durable line of activity rather than an isolated contract.
One discreet but consistent thread remains: tooling is being structured to be driven by an agent, not just a human. ElevenLabs releases a CLI whose selling points โ discoverable commands, structured JSON, dry-run mode โ are explicitly aimed at code agents. Codex gains access to the Chrome DevTools Protocol, meaning the ability to read a pageโs actual state rather than its render. ADK brings native voice evaluation so speaking agents stop being shipped on intuition. And Grok Voice publishes production volumes โ 15,000 calls per day at Starlink โ instead of sticking to a ranking score. Each time, the same shift: getting out of the prototype requires machinable interfaces and reproducible measurements.
Sources
- Wan 3.0 โ official page
- NVIDIA โ Vera Rubin NVL72 and agentic workloads
- NVIDIA โ Groq 3 LPX, Spectrum-X, and NVLink Fusion
- NVIDIA โ NVLink Fusion and XPU
- NVIDIA โ Nemotron 3.5 Lightning on PinchBench
- Mistral and HUMAIN โ official announcement
- Mistral โ announcement of the collaboration with HUMAIN on X
- Mistral โ control as a defining issue
- Sakana AI โ integrated analysis for Defense
- Sakana AI โ announcement on X
- Anthropic โ enterprise-managed auth for MCP connectors
- Anthropic โ enterprise-managed auth announcement on X
- Boris Cherny โ cybersecurity refusals
- Boris Cherny โ coding, engineering, and automatic maintenance
- Boris Cherny โ prompt injection as a solvable problem
- Anthropic โ a personalized weekly digest with Claude Code
- OpenAI โ GPT-5.6 in Kiro
- OpenAI โ ChatGPT and Codex changelog
- OpenAI โ Codex Live Build with Alex Finn
- Google โ the award-winning projects from the Gemma 4 Good Challenge
- Google โ evaluating voice agents in ADK
- Google โ the 110th anniversary of U.S. national parks
- xAI โ Grok Voice Think Fast 2.0
- xAI โ Speech-to-Speech ranking on X
- xAI โ Grok 4.6 in the Nous Research portal
- ElevenLabs โ CLI v1
- MiniMax โ 14 days of unlimited access on GMI Cloud
- Runway โ Ruby extended to all models
- Runway โ hackathon in San Francisco
- Pika โ Wan 3.0 in the Club API
- HeyGen โ prompt guide for real estate videos
- GitHub โ webinar on the Copilot app
- GitHub โ season 2 of the podcast