Search

Google announces Gemini 4 Argon, OpenAI says it thwarted a distillation campaign, Cohere launches Embed 5

ai-powered-markdown-translator

Article translated from French to English using gpt-6.1-sol.

View project on GitHub ↗

On Wednesday, September 30, Google announces Gemini 4 Argon, a frontier model initially rolling out to trusted cyber defenders in the Fairwind Program, with output expanded to 1 million tokens. OpenAI says it thwarted a coordinated campaign to extract its models’ protected reasoning, attributing the core of the activity to individuals associated with Moonshot AI. Cohere launches Embed 5 with its own evaluation method, Ideogram 4.5 promises successive edits without artifacts, and HeyGen releases a video model built on MiniMax H3; on the developer side, Claude Code 2.1.286, Zed 1.22.0 and VS Code 1.140 are released on the same day.


Gemini 4 Argon: Google’s new frontier model, initially for cyber defenders

September 30 — Google announces Gemini 4 Argon, its new frontier model, in a post by Koray Kavukcuoglu, Senior Vice President of Google DeepMind and Google’s Chief AI Architect (Chief AI Architect). Argon is designed to support deep reasoning across long, complex workflows (long-horizon workflows) in three areas: real-world software engineering, enterprise knowledge work (legal, finance) and cyber defense.

The model is not yet available to the public. It is initially rolling out to a group of trusted cyber defenders in the Fairwind Program, launched on September 2 with Gemini 3.8 Flash Cyber, who receive it, like Google’s internal teams, without cyber guardrails (cyber guardrails). Developers, businesses and the general public will follow “as soon as possible,” starting with paying API customers and Google AI Ultra subscribers, with no date announced. In the meantime, Google is strengthening refusals of harmful cyber and NRBC requests, monitoring the model’s chain of thought and actions, and isolating its high-risk training and evaluations in sealed sandboxes; the company also says it is participating in the US government’s voluntary process for access to models before their release.

Introducing Gemini 4 Argon – our new frontier model. It’s built for complex workflows across coding, enterprise knowledge work, and cybersecurity defense – rolling out today to a set of trusted testers through our Fairwind Program. — @GoogleDeepMind on X

Pricing is already set: introductory pricing is 2 dollars per million input tokens and 10 dollars per million output tokens, matching the price of GPT-6.1 Sol launched the previous day, then 4 and 20 dollars after a period whose duration is not specified; cached input costs 95 % less than input. The output limit increases to 1 million tokens, up from 64K previously, the highest in the industry according to Google: the model can produce hundreds of thousands of tokens in a single trajectory to solve a difficult problem.

Price per million tokensIntroductory pricePrice after the introductory period
Input2 dollars4 dollars
Output10 dollars20 dollars

The Google DeepMind page compares Argon with GPT-6 Astra, Claude Fable 5.1 and Claude Opus 5.5 across nineteen rows. Argon achieves the highest score on 13 of them, shares first place with GPT-6 Astra on CWE-bench v1 (68 %) and trails others on 5. Agentic coding is the most closely contested area: first on DeepSWE v1.1 (77,9 %) and Vibe Code Bench, Argon finishes last among the four on FrontierSWE v2 and Terminal-bench 4.0. Its lead is clear in knowledge work, particularly on Harvey’s Legal Agent Benchmark (19,6 % versus a maximum of 6,7 % for the other three), and in long context between 256K and 1M tokens. The highest score in each row is in bold.

Evaluated benchmark (domain)Gemini 4 ArgonGPT-6 AstraClaude Fable 5.1Claude Opus 5.5
Vals Index (knowledge work)68,9 %63,1 %65,8 %67,0 %
AutomationBench (knowledge work)51,3 %41,4 %31,4 %42,5 %
Vals Finance Agent v2 (knowledge work)65,4 %53,5 %58,9 %58,6 %
Harvey’s Legal Agent Benchmark (knowledge work)19,6 %5,4 %6,7 %3,8 %
DeepSWE v1.1 (agentic coding)77,9 %74,1 %67,4 %74,2 %
FrontierSWE v2 (agentic coding)55,0 %65,5 %56,3 %62,3 %
Vibe Code Bench (agentic coding)91,9 %89,6 %90,3 %90,3 %
Terminal-bench 4.0 (agentic coding)57,4 %58,2 %57,9 %66,4 %
PostTrainBench (ML engineering)45,3 %44,3 %40,2 %49,3 %
Terminal-Bench Science 0.1 (science and math)57,6 %68,1 %52,6 %63,3 %
LABBench 2 (science and math)88,8 %85,4 %68,6 %73,1 %
RiemannBench (science and math)76,0 %72,0 %65,6 %69,6 %
GraphWalks BFS up to 128K (long context, F1)99,7 %98,7 %91,4 %90,6 %
GraphWalks BFS from 256K to 1M (long context, F1)84,2 %71,8 %65,0 %66,8 %
Agent’s Last Exam (computer use)39,5 %34,2 %—38,2 %
OSWorld-2.0, offline subset, partial score (computer use)69,2 %72,6 %——
Chartography (multimodal understanding)71,6 %71,0 %46,2 %66,3 %
LVBench (multimodal understanding)91,7 %87,5 %79,7 %83,7 %
CWE-bench v1 (cybersecurity)68,0 %68,0 %58,0 %67,0 %

At Google, thousands of employees already use it. On libgav1, Google’s open source video decoder, Argon agents replaced 32 000 lines of SIMD code in an existing Rust port with safe Rust: the resulting decoder is 2,7 times faster than that port, with identical video output. Other agents have prepared memory optimizations for data centers that are expected to free up more than 300 TiB once deployed. In cybersecurity, Wiz uses it in its Scan for Good initiative: Argon discovered a critical vulnerability exposing sensitive personal data in healthcare software used by hospitals worldwide, which previous frontier models had missed.

🔗 Gemini 4 Argon: our next era of frontier intelligence (Google blog) 🔗 Google DeepMind’s Gemini page and benchmark table


OpenAI says it thwarted a distillation campaign and attributes its core to individuals associated with Moonshot AI

September 30 — In a post in its Security section, OpenAI says it identified and then neutralized a coordinated campaign aimed at extracting its models’ protected reasoning (protected reasoning), meaning the internal trace a model produces while working on a task. The company describes this as adversarial distillation (adversarial distillation): the systematic, unauthorized use of a model’s outputs or reasoning to train, replicate or improve another model.

According to OpenAI, the operators neither broke encryption, compromised a database nor directly accessed users’ stored conversations. They manipulated exchanges to make protected reasoning reappear in a visible form, for example by copying encrypted reasoning from one conversation and asking a model, in another conversation, to decrypt and transcribe it. Independent researchers had reported related vulnerabilities through responsible disclosure, described in the paper “Stealing Reasoning Traces from Proprietary LLM APIs,” submitted to arXiv on August 10; OpenAI confirms that these attack paths were real.

Campaign stageDate or volume recorded by OpenAI
Start of activity, at low volumeJuly 1
Activity peaksJuly 24 and 25
Requests matching the extraction pattern during peaks16 000, from more than 4 000 users
Group linked to this patternmore than 15 000 users
Complete neutralization of the groupJuly 28

On attribution, the post remains cautious: it has not been established that all the operators belong to the same actor, and the 16 000 requests during the peaks were extraction attempts, not necessarily successful ones. However, OpenAI attributes a core of the activity to individuals associated with Moonshot AI, the developer of Kimi.

… we attribute a core cluster of the activity to individuals associated with Moonshot AI … — OpenAI, Security section post

In response, OpenAI banned or restricted fraudulent accounts, tightened its signup and infrastructure controls, expanded monitoring of linked networks and strengthened protections for hidden reasoning across users, workspaces, organizations and model families. It closed a route that allowed another user’s encrypted reasoning to be replayed to recover its contents, and added controls that detect and withhold streaming output (streaming) that might expose it. The findings were shared through the Frontier Model Forum and government channels: according to OpenAI, the technique is not specific to its models, and any system that makes reasoning portable or replayable may be exposed to related risks. The company presents adversarial distillation as a safety and national security risk, since extracted reasoning could be used to train another model without the original’s guardrails. The issue is not new: in February, Anthropic attributed distillation campaigns to DeepSeek, Moonshot AI and MiniMax, and since September 24, Anthropic has been charging for blocked requests in several categories, including reasoning extraction.

🔗 Paper “Stealing Reasoning Traces from Proprietary LLM APIs” on arXiv


Embeddings: Cohere launches Embed 5 and the RCP-nDCG@10 metric, Perplexity releases pplx-embed-v2-context-9b-preview

September 30 — Cohere launches Embed 5, a family of embedding models for enterprise search, in two tiers that share the same embedding space: Embed 5 Pro, for maximum quality and offline indexing, and Embed 5 Fast, for interactive search, high-volume RAG and agentic search. A corpus can be indexed with Pro and then queried with Fast without reindexing anything, provided the output dimension stays the same: across 40 development datasets, this setup recommended by Cohere retains an average of 98,4 % of the quality of an all-Pro system, compared with 96,6 % for an all-Fast system.

Both models read up to 128K tokens, accept text, images or both fused into a single vector, cover more than 100 languages and produce vectors with 256 to 2 048 dimensions in float, int8 or binary formats. Cohere recommends 1 024 dimensions in int8, or 1 Ko per vector compared with 8 Ko in float32 at 2 048 dimensions. Pro costs 0,12 dollars per million text tokens and Fast 0,08 dollars, a third less, with an average document throughput 2,4 times higher; images are billed at 0,40 dollars per million tokens for both. Embed 5 is available today through the Cohere API, Model Vault, Microsoft Foundry and Amazon SageMaker, as well as in North, and can be deployed privately with vLLM.

Benchmark (metric)Embed 5 ProEmbed 5 FastOther models cited
ViDoRe V3, 8 domains (RCP-nDCG@10)85,884,5Voyage 4 Large 83,7 ; Gemini Embedding 2 83,2 ; OpenAI text-embedding-3-large 75,5
FinanceBench (RCP-nDCG@10)80,180,0Pro 21,4 points ahead of OpenAI text-embedding-3-large
FinQA (RCP-nDCG@10)90,088,8—
ViDoRe V3 Finance (RCP-nDCG@10)85,083,9—
Parsed PDFs, suite average (RCP-nDCG@10)84,883,4Voyage 4 Large 83,6 ; Gemini Embedding 2 80,8 ; Embed 4 78,6
Fused text and image, 5 datasets (nDCG@10)82,381,2Gemini Embedding 2 61,3
Page image, 5 financial datasets (nDCG@10)77,073,2Embed 4 71,1 ; Voyage Multimodal 3.5 70,1 ; Gemini Embedding 2 56,7
5 European languages (nDCG@10 or RCP-nDCG@10)77—Voyage 4 Large 76 ; Gemini Embedding 2 73

Most of these scores are expressed in RCP-nDCG@10, the evaluation method Cohere releases on the same day (see below): a note in the post explains that it reranks a fixed candidate set and therefore measures ranking quality rather than first-stage retrieval. The second multilingual table deserves a full reading: Cohere highlights its improvements over Embed 4, up to 13 points in Farsi, but across these ten languages, Gemini Embedding 2 outperforms Embed 5 Pro on nine, with Chinese the only exception, and Voyage 4 Large outperforms it on five. Cohere will present Embed 5 during a session on X on October 8.

🔗 Cohere’s Embed 5 post · Announcement from @cohere on X

RCP-nDCG@10, the metric Cohere uses to measure Embed 5

September 30 — Cohere also releases RCP-nDCG@10 (Rubric-Calibrated Preferences nDCG@10), its search evaluation method. The starting point: nDCG, the metric used by MTEB and BEIR, compares results against a necessarily incomplete list of documents annotated in advance, so a relevant document that has never been annotated earns no credit; in Cohere’s study, humans nevertheless judge 28 % of documents marked irrelevant to be useful. RCP-nDCG@10 assigns the assessment to a calibrated AI judge, which answers the same five yes/no questions for each query and compares documents in groups. Blind validation with 46 annotators across 289 head-to-head comparisons: when the two metrics identify different winners, humans agree with RCP-nDCG in 70 % of cases. The paper, code and data are published, and MTEB’s lead maintainer announces its integration. Cohere says it has optimized its fifth generation of Embed and Rerank against this metric, with the latter still having no release date, and cautions that these models will not always appear best under traditional nDCG alone.

🔗 Cohere’s RCP-nDCG@10 post · Code and data on GitHub

Perplexity releases pplx-embed-v2-context-9b-preview and the context-bench benchmark

September 30 — Perplexity releases pplx-embed-v2-context-9b-preview in preview under the MIT license on Hugging Face, a contextual embedding model: each document chunk is encoded with the entire document in view, so a passage that says “she” remains linked to the correct entity. Rather than learning from a single “gold” passage (gold passage) per query, the model learns from a teacher, Perplexity’s context compression model, which scores every token in the document: it thus learns to retrieve both the answer and the passages that support it. In a blind evaluation on context-bench, a private benchmark from turbopuffer, coauthor of the post (2 099 queries, 38 894 documents), it outperforms voyage-context-4; on ConTEB, it achieves the best average without winning every task. Perplexity cautions that the weights and interface may still change, and is preparing access through its API, with no date announced.

Metric published by Perplexitypplx-embed-v2-context-9b-previewDifference from voyage-context-4
Answer recall on context-bench, at K = 1045,5 %+14,4 points
Evidence recall on context-bench, at K = 1040,6 %+5,0 points
Storage per vector (1 024 dimensions, int8)1 Ko8 times less than voyage-context-4 in float32

🔗 Perplexity Research post · Model on Hugging Face


Ideogram 4.5: an image editing model designed for successive edits

September 30 — Ideogram launches Ideogram 4.5, which it presents as “the most precise editing model.” The starting observation: with each edit, image models introduce artifacts, pixel shifts and color drift, so an image quickly becomes unusable after several passes. Ideogram 4.5 is designed to eliminate this accumulation and make editing over multiple turns possible (multi-turn editing).

Introducing Ideogram 4.5, the most precise edit model. With each edit, leading models add artifacts, pixel shifts, and color changes. Ideogram 4.5 eliminates artifact buildup, making multi-turn editing possible. Live in Ideogram, the API, and launch partners. Open weights soon. — @ideogram_ai on X

To support this point, Ideogram publishes a side-by-side comparison of the same successive edits with GPT Image 2.5 Sunburst, Nano Banana Pro and Nano Banana 2, whose outputs it says become unusable within a few edits. The comparison is visual, with no numerical score. Examples cover color and lighting (shadows, reflections and highlights follow the change without shifting the composition), text modification, product photography, interior design, step-by-step restoration of old photos, turning a sketch or depth map into an image, and cropping.

The model also introduces editing at any resolution (Zoom editing): a region of a high-resolution image, 24,2 megapixels (4 016 × 6 016 pixels) in Ideogram’s example, can be edited without downscaling the image, with its edges preserved so it can be reintegrated without a visible seam. Although designed for targeted editing, the model also performs very well in general editing, according to Ideogram.

Announcement itemWhat is known
Availabilityfrom September 30 in Ideogram and through the API
Launch partnersPika and Luma, which announce it the same day
Open weightspromised “soon,” like those of Ideogram 4.0 released in June
Pricingnot published

🔗 Ideogram 4.5 model page


HeyGen Video: a video model built on MiniMax H3, available through an API

September 30 — HeyGen launches HeyGen Video, a video generation model aimed at businesses that want production-quality video without paying the associated cost. Built on MiniMax H3 and post-trained by HeyGen, it produces video from text or an image, with audio generated in the same call, and can be used through HeyGen’s API as well as on OpenRouter, Runware and ComfyUI.

Three pricing figures coexist. Pricing starts at 0,01 dollars per second until the end of October, a 50 % discount on the standard rate of 0,02 dollars, according to the catalog page. Its comparison table, however, uses the list price at 768p with audio, before launch promotions: 0,03 dollars per second. Even at this price, HeyGen Video is the cheapest of the models compared.

Model comparedList price per second (768p, audio)HeyGen’s internal Elo (HeyGen Video = 1 000)
HeyGen Video0,03 dollar1 000
Seedance 2.00,303 dollar955
H3 Max0,08 dollar952
H3 Max Turbo0,04 dollar920
H3 Max balanced0,08 dollar860
Kling 3.0 Pro0,168 dollar857
Veo 3.10,40 dollar744

For quality, HeyGen relies on an internal evaluation conducted using prompts from Artificial Analysis Arena (4 800 votes), with Elo normalized to 1 000 for its model. In a blind preference comparison against H3 Max, 55,8 % of the 650 votes go to HeyGen Video, within a range of 49 to 62 % that does not rule out a tie. For a 10-second image-to-video clip, diffusion transformer inference time falls to 3,7 seconds, compared with 4,1 for H3 Max Turbo and 8,3 for H3 Max, excluding captioning time.

MiniMax welcomed the launch and highlighted another H3-derived model on the same day: Creatify’s Boreal-H3, a video model optimized for advertising.

🔗 HeyGen Video catalog · Announcement from @HeyGen on X


General-purpose assistants and agents: Gemini skills, Manus Flex and Grok Bot

The Gemini app launches skills, which will replace Gems

September 30 — Google launches skills in the Gemini app: instructions saved once and recalled by typing a slash followed by their name. Gemini can also create them from conversations and launch them on its own when a prompt matches, multiple skills can be combined, and each can include reference files (plain text, PDFs, images). Already available in Gemini Spark, they are coming to Gemini chat worldwide, across all Google AI subscription tiers and, for now, for users aged 18 and over; Workspace customers will follow in the coming weeks. Skills will replace Gems, which will disappear starting in November for personal accounts, in March 2027 for Workspace business, enterprise and nonprofit, and in June 2027 for education, with automatic migration. Opal, the mini-app creation tool, will also shut down in November, as will “Gems by Google Labs,” which will not be migrated.

🔗 Let skills in Gemini tackle your most repetitive tasks (Google blog)

Manus Flex: your own API key in the Manus agent

September 29 — Manus launches Manus Flex, which allows Manus to be powered by an API key from a supported inference provider (bring your own API key). Until now, Manus selected the model itself and supplied the agent harness and infrastructure; with Flex, a provider’s key powers the same projects, tools, runtime environments and agent workflows, with a choice of primary model and reasoning effort level. Billing is split: inference is billed directly by the provider, while the other services and infrastructure used by a task continue to consume Manus credits. OpenRouter, Fireworks and Modal are the first three members of the inference partner program (Manus Flex Inference Partner Program). The post, published the day after the launch of Manus 2.0, provides no plan, pricing or model list.

🔗 Introducing Manus Flex

Grok Bot delegates coding to Cursor and manages pull requests

September 30 — Two days after Team Bots, SpaceXAI expands Grok Bot’s software development capabilities. In a thread from the @bot account, the company announces that its Bots can delegate coding tasks to Cursor, manage pull requests through GitHub and Origin plugins (the latter is not described), and share video demos of what they build. SpaceXAI also releases its own engineering team’s Bots as installable templates (templates), each supplied with its skills, routines and connectors. The thread specifies no plan, pricing or date, and no accompanying post appears on x.ai. The connection with Cursor is not new: Grok Bot for Enterprise was offered for two weeks to Grok and Cursor Enterprise customers in early September; what changes is that a Bot delegates the coding work itself.

🔗 Thread from @bot on X


Robotics and world models: Runway’s Praxis-1, Ai2’s MolmoAct 2 and NVIDIA’s Physis-Lang

Runway introduces Praxis-1, an open-weight world action model

September 30 — Runway introduces Praxis-1, its first open-weight world action model (world action model), designed to control real robots. It relies on the same video pretraining as its world models, such as Solaris or GWM Worlds 2, and aims to be a generalist policy model for robotics developers and researchers. Runway’s premise: robotic demonstrations are scarce, while video is virtually unlimited. In its demonstrations, a single policy moves from a studio to a kitchen without retraining. Nothing is available to download yet: Praxis-1 is being tested at Noble Machines, Standard Bots, and Ultra, each on its own hardware, and its public release, including weights, is announced for the coming months.

Metric published by RunwayReported result
Correlation between world model simulation and real-world results0,95
Final placement error after fine-tuning with web video16,1 cm
Same error with video of a teleoperated robot16,0 cm

🔗 Praxis-1 on Runway Research

Ai2’s MolmoAct 2 tops Reality Check after fine-tuning by the benchmark organizer

September 30 — Ai2 announces that MolmoAct 2, its fully open robotics model, ranks first for overall success on Reality Check, an independent benchmark for real-world manipulation, with a major caveat: the models are fine-tuned by the organizer, Poke & Wiggle, on the benchmark’s tasks. Launched the previous day by the company, Reality Check includes 14 400 runs on real robots, using FR3 Duo stations installed at the Deutsches Museum in Munich, across ten environments (sorting screws, plugging in a DC connector, opening a toolbox with a screwdriver…). Two tracks are still “coming soon”: objects placed outside the training area and mid-training on 100 hours of station data.

Evaluated modelOverall successSuccess after approximately 10 demonstrations per environmentSuccess after approximately 300 demonstrations per environment
MolmoAct 228 %4 %44 %
Pi 0.521 %5 %33 %
DiT-Flow19 %2 %29 %
GR00T N1.712 %4 %19 %

🔗 Ai2’s tweet · Reality Check leaderboard

Physis-Lang: captions that explain physics to world models

September 29 — Researchers from NVIDIA, MIT, and the University of Oxford publish Physis-Lang, a framework that adds physical reasoning to video captions to improve world models. The captions make causes, governing laws, and effects explicit; a specialized critic identifies missing or unsupported claims, an agent refines the captioning instructions, and model failures guide the search for videos that will address its gaps. According to NVIDIA, adding this reasoning to the prompt alone, without retraining, improves Cosmos 3’s PhyGenBench score by 5,62 points. With training, Physis-Lang on Cosmos3-Super and Cosmos3-Nano takes the top two spots on the Physics-IQ Verified image-to-video leaderboard (48,2 and 43,3, versus 42,7 for the previous best score). PhyGenBench and VideoPhy-2 are scored by GPT-5.5, and Veo 3.1 retains a slight advantage on the full VideoPhy-2 set.

Video physics benchmarkCosmos3-NanoVeo 3.1Physis-Lang on Cosmos3-Nano
PhyGenBench61,6765,6371,04
Physics-IQ Verified40,2334,9943,41
VideoPhy-2 Hard48,3158,4362,36
VideoPhy-2, full set60,4168,8768,02

🔗 Physis-Lang project page · Announcement by @NVIDIAAI on X


Open models: Hunmin-397B-A17B-CUA and JEV-27B-VL

Hunmin-397B-A17B-CUA grafts Qwen-CUA’s agent capabilities onto a 397-billion-parameter MoE

September 30 — On the Hugging Face community blog, hojun Lee details Hunmin-397B-A17B-CUA, a computer-use (computer use) model released under Apache-2.0 by the mncai organization, with a caveat from the authors themselves: the evaluations are in-house, with a single run for the agent benchmarks, and cannot be compared with published rankings (the base model scores 48,16 on OSWorld here, versus the reported 62,2 on OSWorld-Verified). The model is based on Qwen3.5-397B-A17B, a mixture of experts with 17 billion active parameters, and the entire project ran on a single node with eight B200 GPUs. The team calculated the weight difference between its base model and the already specialized Qwen-CUA, and grafted only a low-rank version of that difference onto selected modules: this transfer alone raises OSWorld-316 from 48,84 to 68,25, without inheriting Qwen-CUA’s weakness in visual grounding. Fine-tuning followed by GRPO reinforcement learning adds 4,3 points.

Evaluation benchmark (in-house protocol)Qwen3.5-397B baseHunmin-CUAQwen-CUA (source)
OSWorld, 360 tasks48,1670,4577,12
WindowsAgentArena, 153 tasks41,7750,8556,68
ScreenSpot-Pro72,7475,6162,20
KMMLU-Pro (Korean)77,9176,1271,41

🔗 Hunmin-CUA post · Weights on Hugging Face

AutoTrust AI releases JEV-27B-VL, an open decision model that also evaluates images

September 30 — AutoTrust AI releases JEV-27B-VL under Apache-2.0, the vision-enabled version of JEV-27B, its open decision model introduced on September 27. Give it images and text, ask a question, and it responds in a single pass with a calibrated probability for each option, in about a hundred milliseconds; when the question calls for it, it reasons step by step while examining the images. Its decision head, however, saw no images during training. Main test: on MicroLens, a public dataset from a short-video app, it ranks 20 candidate videos for 200 users using thumbnails alone, and matches the AUC of collaborative filtering trained on 59 045 users, with a higher rate of the correct video appearing in the top 5. The authors stress that this result comes from a single set of 200 users.

Recommendation method on MicroLensAUCCorrect video in the top 5
JEV-27B-VL, thumbnails only, no interaction data0,72759 %
Collaborative filtering trained on 59 045 users0,72849 %
JEV-27B-VL, titles only0,64946 %
Random order0,48921 %

🔗 JEV-27B-VL post


Leaderboards: Hugging Face’s Open TTS Leaderboard and LILT’s AURORA

Hugging Face launches the Open TTS Leaderboard, an open leaderboard for speech synthesis

September 30 — Hugging Face launches the Open TTS Leaderboard, a speech synthesis (text-to-speech, TTS) leaderboard focused on open and multilingual models. The Hub has more than 8 000 TTS models, but the leading voting arenas struggle to keep pace and underrepresent open models: according to the post, only 16 of the 92 models ranked by Artificial Analysis have open weights. The leaderboard therefore replaces votes with objective measurements: intelligibility (word or character error rate between the requested text and its transcription by Qwen3 ASR), speed (batched inference and time to first audio, on H200 and CPU), and cloned voice fidelity (WavLM similarity). Evaluating a model takes a few hours instead of a few weeks of voting. In English, Kokoro-82M, supertonic-3, and Fish Audio’s s2-pro lead; across languages, OmniVoice, s2-pro, and Fun-CosyVoice3-0.5B lead. The authors caution that neither naturalness nor expressiveness is measured, and will publish their evaluation scripts “soon.”

🔗 Open TTS Leaderboard post · The leaderboard on Hugging Face

LILT launches AURORA, a multilingual leaderboard for agent tasks

September 30 — LILT launches AURORA, a leaderboard that evaluates frontier models on non-English agent tasks, all designed and verified by native-speaking experts, without machine translation. It launches with four benchmarks: a multilingual Terminal-Bench with 324 coding tasks in ten languages, a multilingual version of τ³-bench for customer support, MultiChallenge for long-context instruction following, and GAIA-v2-LILT for agentic reasoning. Claude Opus 5.5 leads the multilingual Terminal-Bench and ranks first on τ³-bench in each of the five languages. On GAIA, models gain an average of 20,7 points on the version audited by LILT compared with a machine translation: for LILT, this gap measures the error introduced by translation. The same dialogue also uses about 1,4 times as many tokens in Korean or Arabic as in English.

Evaluated model (multilingual Terminal-Bench, excerpt)Average success
Claude Opus 5.5 (effort high)76,4 %
Claude Opus 5 (effort high)73,5 %
Gemini 3.8 Flash67,3 %
GPT-6 Sol64,8 %
Command A+4,3 %

🔗 LILT’s AURORA post · The AURORA leaderboard


Public sector and businesses: Claude for Government, Anthropic’s purchasing agent, and OpenAI with America’s SBDC

Claude for Government becomes generally available

September 30 — Claude for Government leaves beta: Anthropic announces its general availability for U.S. federal and state agencies. The platform, in public beta since July, provides Claude’s coding and agentic work capabilities in a FedRAMP High-authorized environment, with features comparable to those available to commercial customers; the Claude Code CLI and Claude for Microsoft 365 are arriving in early access. No per-seat licensing: agencies pay for usage in fixed tiers, with a hard cap that prevents them from exceeding the committed budget, and SCIM group mappings set rate limits, dollar caps, and permitted models by seat tier. On the control side, sensitive operations on Anthropic’s side require two-person approval, usage exports contain only metering data, and conversation history stays on the agency-managed device. Anthropic does not publish any pricing.

🔗 Claude for Government is now generally available

At Anthropic, an agent built on Managed Agents handles sales inquiries

September 30 — Anthropic describes how its sales team rebuilt how it handles inbound inquiries, tens of thousands per month, using a purchasing agent built on Claude Managed Agents, in beta. Available on the Contact Sales and Pricing pages, in claude.ai, and in emails, it asks a few questions, recommends a plan and a seat count, then guides the customer to a purchase, hands them over to a sales representative with the full history, or simply answers questions; the customer can choose a human from the outset. According to the post, written by Carl Johnson, the agent handles thousands of conversations per day; the leads it hands over become opportunities more than twice as often as with the previous form and close about five days faster, and the share of conversations that require a sales representative to close has fallen by roughly half. A single engineer built the first version in a few weeks; the agent often directs small teams toward the Team plan rather than Enterprise, an outcome Anthropic has chosen to accept.

🔗 How Anthropic’s sales team rebuilt inbound with Claude Managed Agents

OpenAI partners with America’s SBDC and publishes a report on small businesses

September 30 — OpenAI partners with America’s SBDC, the national network of U.S. small business development centers, which supports 1 million entrepreneurs per year. In an initial phase, through OpenAI Academy’s community trainer program (Community Trainer Program), launched as a pilot on September 23, OpenAI plans to train around 150 advisers, who will run three-hour workshops across the SBDC networks, aiming to reach at least 1 000 small businesses in person. That same day, the report “Small Businesses, Bigger Capabilities” quantifies usage: during the week of September 9 to 15, around 4 million employees at companies with fewer than 500 people used OpenAI products, including nearly one in five at a company with fewer than 10 employees. In August, agentic output tokens, particularly those from ChatGPT Work and Codex, accounted for two thirds of small businesses’ output tokens, compared with one third in April.

🔗 OpenAI’s post on small businesses


Runway Ads: an autonomous engine for performance advertising

September 30 — Runway launches Runway Ads, which handles the creative side of paid campaigns from end to end. Connect an advertising account and a brand kit: the tool, built on Runway Agent, generates video and image creatives, publishes approved variants on Meta, Google and TikTok, reviews their performance and produces the next wave based on what attracted ad spend. It localizes each creative according to rules set by the team, including on-screen text, resizes it for each format, subjects it to an automatic brand check and respects budget limits. Human approval is enabled by default; automatic publishing can be enabled by campaign or variant type.

Runway internal metric, since JulyResult published in the post
Ads per weekfrom 77 to approximately 900
Return on ad spenddoubled
Conversionapproximately +34 %, stable click-through rate
Cost per subscriber−41 %

Runway Ads is currently being piloted with selected enterprise partners. The announcement tweet promises a broad rollout in the coming weeks and claims the tool has also helped add 100 million dollars in ARR (annual recurring revenue) since July, a figure absent from the post.

🔗 Runway Ads post · Tweet by @runwayml


ElevenLabs valued at 22 billion dollars after a 300 million dollar tender offer

September 30 — ElevenLabs has completed a 300 million dollar share buyback offer (tender offer) for its employees, valuing the company at 22 billion dollars, twice its valuation in February’s Series D. The transaction is led by Wellington and T. Rowe Price; EQT, Goldman Sachs, GIC, OTPP, Sapphire Ventures and BDT & MSD become shareholders alongside existing investors such as Andreessen Horowitz, Lightspeed and ICONIQ. Enterprise customers account for 55 % of revenue: ElevenAgents handles more than 15 million conversations per week, three times as many as in February, and its annual recurring revenue has more than tripled over that period. According to an analysis of its customers, voice agents resolve requests 31 % faster than chat agents. The company has more than 800 employees, four years after an initial funding round at a valuation of 9 million dollars.

🔗 ElevenLabs post


Vera Rubin NVL72 in production at CoreWeave, Cognition measures up to 4.8 times the throughput

September 30 — At CoreWeave Fully Connected in San Francisco, NVIDIA details Vera Rubin NVL72’s move into production at CoreWeave: the platform, with Spectrum-X Ethernet 102.4T networking, is available on CoreWeave Cloud, which received its first production racks earlier in the month. Cognition, the developer of Devin, is the first customer to run production workloads on it: in initial tests on tasks drawn from FrontierCode, it measured up to 4.8 times the total token throughput for SWE-2 inference compared with GB200 NVL72. CoreWeave will also offer the Vera CPU, designed for agents: 128 CPUs and 11,264 cores per rack, enough to run more than 11,000 simultaneous isolated environments, with agent sandboxes that start more than three times faster. Finally, CoreWeave is launching CoreWeave Forge, which brings together Weights & Biases, OpenPipe and marimo to feed production back into training; its serverless reinforcement learning (serverless RL) is advertised as 1.4 times faster, at 40 % lower cost.

🔗 NVIDIA post on CoreWeave and Vera Rubin


Coding agents and development tools: Claude Code 2.1.286, Zed 1.22.0, Gemini CLI, VS Code 1.140, HydraFusion and an MCP server for gcloud

Claude Code 2.1.286: bookmarks in VS Code, capped retries and a more restricted bare mode

September 30 — Claude Code 2.1.286, released at 19:10 UTC, has 88 entries, including 49 fixes. The permission prompt indicates its position when multiple requests stack up (“2 of 5”), and the VS Code extension gains bookmarks (bookmarks) for saving Claude responses in a side panel, a Questions row that shows the questions Claude asked and the answers selected, and option previews in question cards. Two behavior changes matter: a single limit now covers retries for a model call, allowing at most 14 requests with the default settings, and --bare now connects only to MCP servers named on the command line, without system reminders or background tasks. If a skill named verify exists, Claude is instructed to run it just before every commit, except for documentation or tests, and plugins reject npm sources that are git repositories or folders. Finally, when the API rejects the default model, Claude Code retries once using the previous model in the same tier instead of failing on every turn.

🔗 Claude Code 2.1.286 release notes

Zed 1.22.0: one model per subagent

September 30 — Zed releases its stable version 1.22.0, focused on subagents: the spawn_agent tool accepts an optional model parameter to launch a subagent on a model other than the one configured or inherited from its parent, and each subagent’s card displays the model used. Subagents also compact their context automatically, allowing them, according to Zed, to handle longer tasks. On the model side, GPT-6.1 Sol arrives in BYOK with an OpenAI API key, Grok 4.7 becomes stable as the recommended model for SuperGrok and xAI, and the OpenCode Zen and Go model list is fetched on the fly. The release fixes a memory leak of approximately 2 MB per completed terminal command and changes a shortcut: closing the active dock moves to ctrl-alt-w on all platforms. In total, 18 new features and 37 fixes.

🔗 Zed 1.22.0 release notes

Gemini CLI: v0.62.0 becomes stable, and a preview executes plans without confirmation

September 29 — Gemini CLI publishes two releases in the evening (UTC). v0.62.0 becomes stable at 21:17: its 19 entries carry over the preview and nightlies covered from September 16 to 24 (structured titles for MCP tool calls, preservation of the OAuth refresh token (refresh token), Gemini 3.8 Flash and 3.5 Flash Lite). The new development is in v0.63.0-preview.0, released at 20:58: in non-interactive mode, the agent now writes and executes its plan without waiting for confirmation (PR 29539), whereas the Plan Mode instructions previously made it respond with a simple alignment message, without any tool calls. A maxChars limit of 0 or less also disables tool output truncation (PR 29542). The September 30 nightly opens the 0.64.0 series: in ACP mode, the CLI finally reports standard usage, whereas clients such as Zed or OpenHands were overestimating billing by approximately 3 times, according to PR 29549.

🔗 Gemini CLI v0.63.0-preview.0 · Gemini CLI v0.62.0

VS Code 1.140 adopts the Copilot harness

September 30 — Visual Studio Code 1.140 becomes stable with the Copilot harness (Copilot harness): based on the Copilot SDK, it gives the VS Code agent the behavior and capabilities of the GitHub Copilot app and Copilot CLI, and runs in a dedicated host process based on the Agent Host Protocol, allowing the same session to be joined from multiple windows; the notes warn that it may already be selected by default for some users. As experimental features, multi-folder sessions give each chat its own folder or worktree, and the agent can delegate a session to a remote host, selected according to the operating system, memory, CPU count and model. Three controls arrive for enterprises: an explanation in chat of the minimum versions required by the administrator, a default Auto mode level set through a managed setting (autoTier), and the addition of the user’s identity to Copilot’s OpenTelemetry data, an option disabled by default.

🔗 VS Code 1.140 release notes

HydraFusion comes to VS Code and the GitHub Copilot app

September 30 — GitHub extends HydraFusion, the multi-model orchestration launched on September 4 in Copilot CLI, to VS Code (starting with version 1.140, or Insiders) and the GitHub Copilot app, still as a research preview. HydraFusion is selected in the picker like a model, but is not one: it treats workflow selection as an optimization problem and chooses one of three patterns. In Single mode, one model solves the task; in Cascade mode, an efficient model drafts and a quality check accepts the result or escalates to a stronger model; in Critique mode, a read-only critic from another model family reviews the draft, then the writer revises it once. Whereas Auto mode selects a model for each request, HydraFusion coordinates several models within a single turn. It is available in Copilot Pro, Pro+, Business and Enterprise, with an administrator required to enable previews in Business and Enterprise; GitHub publishes no new figures.

🔗 HydraFusion in VS Code and the GitHub Copilot app

Google Cloud previews a remote MCP server that executes gcloud and bq

Google Cloud adds the Google Cloud CLI remote MCP server to its managed remote MCP servers, in public preview. It brings together hundreds of commands from the gcloud and bq (BigQuery) command-line tools behind two MCP tools, run_gcloud_command and run_bq_command. Google gives two reasons for choosing the command line: a command encapsulates multi-step operations that the API would require callers to orchestrate, and models have been extensively trained on these commands’ public documentation. The remote server removes the need to install the CLI in the agent’s environment, making these operations available to agent platforms hosted on the web, such as Gemini Enterprise. Commands execute in an isolated sandbox, behind a proxy with restricted network access and without ambient credentials (ambient credentials), each with the IAM permissions of the calling identity, authenticated through Agent Identity or OAuth 2.0; Model Armor can filter prompts and responses, and each call can be recorded in audit logs. The server itself incurs no charge: only resources created and any data transfers are billed.

🔗 Empower your agents with the Google Cloud CLI remote MCP server (Google Cloud blog)


Briefs

  • Codex CLI 0.159.2 — Released on September 29 at 23:57 UTC, this version contains just one backported fix: on Windows, console windows no longer flash when Codex launches background processes or commands in its sandbox. 🔗 source
  • Plaid in Amp — Amp modes that use GPT-6 Astra gain Plaid, based on OpenAI’s ultrafast tier: requests up to 6 times faster at 6 times the cost per token, available only for inference provided by Amp; sub-agents remain at fast or standard speed, and Amp publishes no pricing. 🔗 source
  • Warp and Factories as Code — Warp presents the configuration of its software factories as “Terraform for agents”: a repository describes agents, automations, access and measurement. The factory.yaml format dates from late August; the interactive guide details federated access to AWS or GCP, webhooks and Slack, Linear or Jira integrations. 🔗 Warp thread · 🔗 configuration guide
  • Devin at Crosby — In a Cognition case study, New York law firm Crosby, with roughly a dozen engineers supporting more than 50 lawyers, assigns Devin the initial response to incidents: 20 to 40 bugs and alerts reviewed per day, most in about 5 minutes, and at least 50% less Sentry noise. 🔗 source
  • claude.dev — Anthropic officially introduces claude.dev as the home for developers building with Claude: technical deep dives, Claude Code and API guides, Agents, Engineering, Playbooks and Skills sections, and a Terminal page as a hidden Easter egg. Posts from the site had already been shared since September 22. 🔗 source
  • Editable ChatGPT Sites — According to the September 29 release notes, a published Site can be modified with Edit site and scheduled through Automations on Plus and Pro; on Enterprise, private Sites can use connected apps, with an Allow use in Sites setting for each plugin, disabled by default. 🔗 ChatGPT release notes · 🔗 Enterprise and Edu release notes
  • Kimi Work 3.2.15 — The September 30 release enables human-AI co-editing of Notion and Feishu documents in the built-in browser, collapses the agent’s activity by default and fixes file corruption when editing a PPT simultaneously. 🔗 source
  • Cue manages finances — The personal agent app launched with Manus 2.0 now tracks subscriptions and spending trends and puts bank accounts on autopilot through Plaid; no countries, plans or terms are specified. 🔗 source
  • GPT-6.1 Sol in Genspark — Genspark makes GPT-6.1 Sol available in AI Chat, Code Agent and Claw the day after its launch, repeating OpenAI’s pitch (intelligence close to Astra at one-fifth of its API price), without specifying plans or credit costs. 🔗 source
  • Qwen3.8-27B-pi — Thomas Kim releases an Apache-2.0 fine-tune of Qwen3.8-27B for the Pi harness that puts effort levels back in order: on Terminal-Bench 2.1, medium matches the base model at xhigh (67 tasks out of 89) with about 41% fewer output tokens. 🔗 source
  • Qwen3.8-27B on Nebius — Qwen shares the arrival of its dense 27-billion-parameter model on Nebius Token Factory, announced on September 28, for coding, research and agentic workflows; no pricing or throughput figures are disclosed. 🔗 source
  • Bekko System One v0 — Yuichi Tateno releases three decision models with 17M, 68M and 400M parameters, the two smallest of which run in a browser, along with the S1MB benchmark: the 400M scores 50.60 against 59.59 for Jev 1.13, but 54.48 against 96.27 on generalization. 🔗 source
  • ZooWork Instinct Tuned 4B — SRP releases an Apache-2.0 decision model with 4 billion parameters, built on Qwen3.5-4B, that scores options in a single pass without decoding: 198 out of 231 (85.71%) on JevBench’s public tasks, which were used during development, the authors specify. 🔗 source
  • Transformers v5.18.0 — Hugging Face adds four architectures, NVIDIA’s Nemotron 3 Diarization and NemotronH Omni, NAVER’s HyperCLOVAX Vision V2 and GTE, and flags six breaking changes. 🔗 source
  • TaskSmith — Adithya S K, who works on reinforcement learning environments at Hugging Face, releases an orchestrator that turns a pull request into a verifiable RL environment: 50 environments drawn from TRL, PEFT, Accelerate, Diffusers and Transformers, delivered as Harbor tasks. 🔗 source
  • TraceML 0.4.1 — The open source tool now measures the entire group of steps making up an optimizer update; on ResNet-50 and a T4 GPU, time spent waiting for data drops from 23% of the step to less than 2 ms after tuning data loading, with a median overhead of 0.35%. 🔗 source
  • Looped transformer — On a toy model with 54,106 parameters and 260 verifiable cases, with the same budget of 80 block passes, one loop learns all 260 cases, two loops learn 125.3 and eight learn 37.0 on average: increasing the number of passes did not make training more economical. 🔗 source
  • An evaluation that can say no — Eric Mey’s tutorial uses movie review sentiment (SST-2) to demonstrate an evaluation that rejects a model despite a 7-point improvement on held-out data, because a critical behavior deteriorates; its scripts run in less than a minute on a CPU. 🔗 source
  • 10,000 synthetic policyholders — Intelligent Actuaries publishes a synthetic study of life insurance lapses and surrenders under CC-BY-4.0: 10,000 fictional policyholders across ten markets, totaling 27,692 policy-year records for 2023–2025 in the SOA-LIMRA survey format. 🔗 source
  • cuObject and SCADA Server SDK — NVIDIA makes its cuObject libraries for accessing object storage through RDMA, bypassing CPU memory, generally available and launches the SCADA Server SDK, prototyped by IBM with Storage Scale, as part of a Storage-Next initiative involving more than 40 participants. 🔗 source
  • NeMo Relay and Hermes Agent — An NVIDIA tutorial coauthored by Teknium of Nous Research traces a Hermes agent: across 108 runs, tool fixes take Qwen3 Coder 30B from 19 to 22 successes out of 27, at the cost of 4.9 model calls instead of 3.8. 🔗 source
  • OpenShell for OpenClaw Enterprise — NVIDIA offers its open source OpenShell environment to govern agents in OpenClaw Enterprise, an open source control plane for persistent agents announced the previous day by the OpenClaw foundation with Red Hat, NVIDIA and OpenAI. 🔗 source
  • GitHub Advanced Security trials — GitHub Team customers can start a trial of GitHub Code Security and GitHub Secret Protection themselves, from the organization’s Overview page, the billing page or a Risk Assessment; the duration is not specified. 🔗 source
  • GHE.com and X25519 — Starting October 7, 2026, GitHub Enterprise Cloud with data residency will reject TLS clients that offer only X25519 for key agreement; P-256 and P-384 remain supported, and SSH is unaffected. 🔗 source
  • Two Runway case studies — French beverage brand Bonjour produced more than 500 ad variants with Runway and reallocated more than 50,000 euros in production spending; the World Juggling Federation assigned one person to create the graphics for a one-hour ESPN broadcast. 🔗 Bonjour · 🔗 World Juggling Federation
  • Ad Variants in Luma — Luma highlights a feature that adapts a finished ad to multiple sizes and markets within a single campaign, with no pricing, plan or launch date, and no blog post. 🔗 source
  • Microduck in production — Pollen Robotics announces that 10,950 Microduck units, its small 399-dollar bipedal robot trained in simulation with an open source reinforcement learning stack, will roll off the production line between December 10 and January 10 in weekly batches. 🔗 source
  • Cognition is hiring — Devin’s developer welcomes Chris Degnan, former CRO of Snowflake, as chief revenue officer, the day after Alex Stamos joined as chief information security officer (CISO). 🔗 Chris Degnan · 🔗 Alex Stamos
  • Claude Founder House — Anthropic is organizing founder days in San Francisco during SF Tech Week (October 6–8 at Terra Gallery) and in Stockholm on October 14, with talks, workshops and sessions with its Applied AI team. 🔗 source
  • AI Engineer Paris — Mistral recaps the event held the previous week at Station F: more than 900 participants, 60 speakers, 57 talks and 22 partners, with no product announcement. 🔗 source
  • NVIDIA doctoral fellowships — Applications for the 2027–2028 Graduate Fellowship Program are open until October 30, 2026, with up to 60,000 dollars per doctoral student; more than 220 fellowships have been awarded since 2002. 🔗 source

What it means

Security now sets the timetable for frontier models. Gemini 4 Argon extends the approach of the Fairwind Program: trusted cyber defenders receive it first, without cyber safeguards, and others will follow “as soon as possible,” with no date, while protections are strengthened. On the same day, OpenAI shows the other side of the problem: what needs protecting is no longer just the model, but also its reasoning. Whether or not it is the work of a single actor, the campaign whose core OpenAI attributes to individuals associated with Moonshot AI makes reasoning traces an asset to defend, with technical countermeasures and information sharing between labs and governments.

A substantial share of today’s figures comes from measurements designed or conducted by those making the announcements. Cohere evaluates Embed 5 using the metric it publishes that same day, and itself warns that its models might look worse under the old measure; HeyGen relies on an internal Elo, Hunmin on an in-house protocol, with a single run for agent benchmarks, and on Reality Check, the organizer fine-tuned the models. By contrast, Perplexity submits its model for blind evaluation on a private turbopuffer benchmark, while the Open TTS Leaderboard and AURORA replace votes or machine translation with verifiable tasks and measurements. Before comparing two scores, look at who produced them.

Image and video generation is increasingly judged on use and cost. Ideogram 4.5 promises an image that can withstand successive edits rather than a prettier image; HeyGen Video is sold by the second, more cheaply than every model in its lineup, and shows that a base model such as MiniMax H3 can yield specialized variants, from HeyGen to Creatify. Runway Ads completes the chain through to ad spending, and ElevenLabs’ valuation, which has doubled since February, rests on enterprise demand for conversational agents.

For developers, the harness is becoming a product in its own right, separate from the model. VS Code adopts Copilot’s harness to align its agent with the app and CLI, HydraFusion coordinates multiple models within a single turn, Zed lets the agent choose a model for each sub-agent, and Manus opens its harness to API keys from third-party providers. Gemini CLI now executes its plans without human intervention in non-interactive mode, Claude Code tightens --bare for scripted use, and Google Cloud exposes gcloud and bq to agents in a sandbox, with the caller’s IAM permissions. Demand is visible even in the infrastructure: the first customer using Vera Rubin NVL72 in production at CoreWeave is Cognition, Devin’s developer.


Sources