Search

Google Cloud launches its Gemini agent for work, HeyGen Voice and dynamic workflows in Claude Managed Agents

ai-powered-markdown-translator

Article translated from fr to en with gpt-6.1-sol.

View project on GitHub ↗

Google Cloud introduced the Gemini agent on October 8 at Gemini at Work 2026: a single agent for all enterprise work, running continuously in the cloud and orchestrating both Gemini and Claude models. Meanwhile, HeyGen is releasing HeyGen Voice, its first in-house voice model, which takes the lead in Artificial Analysis’s Controlled Voice ranking, while Anthropic introduces workflows in Claude Managed Agents capable of launching up to 1 000 agents per run. On the coding side, Codex learns to predict the user’s next message and adopts a new Windows sandbox based on Microsoft Execution Containers.


Google Cloud launches the Gemini agent, a single agent for all work

October 8 — At Gemini at Work 2026, Google Cloud introduced the Gemini agent, which it describes as a single, universal agent for work. From a single input field and a single API, it answers questions, handles knowledge work, creates images and media, and writes and executes code. The idea put forward by Thomas Kurian, CEO of Google Cloud: give it a goal rather than a sequence of instructions.

The agent runs continuously in the cloud, with a single memory, and is available everywhere: web, iOS, Android, Windows, Mac, command line, Google Workspace, Microsoft 365 and Slack, or without an interface (headless) in third-party applications. A task lasting several hours or days keeps running even with the laptop closed. For lengthy tasks, it creates temporary subagents, each with its own identity, and it can become a coworker agent with a permanent role, its own @agents.company.com address and storage. The model becomes a separate choice: according to Google, the agent already orchestrates Gemini models and Anthropic’s Claude models, with other proprietary and open models to follow. Industry-specific versions are entering preview for finance (more than 50 core skills, already used by CME Group and Deutsche Bank) and legal work.

Announced control componentRole described by Google
Each agent’s identityCryptographically attested identity, role-based permissions, audit log in its name
Agent SandboxIsolated execution, with its own network boundary
Agent GatewayNetwork firewall for AI, through which all agent traffic passes
Real-time spending capsIn Cloud Billing, the project’s agent pauses, with one-click resumption

Google Cloud also shares its figures: nearly 500 customers have each processed more than 1 000 billion tokens in a year, and nearly 90 % of the Fortune 100 uses Gemini Enterprise. However, the post provides neither an availability date nor pricing for the agent itself; Google Cloud only names major customers that tested it ahead of release, such as sports brand On for dynamic model selection.

🔗 Gemini at Work 2026 (Google Cloud blog)

Gemini Enterprise makes Claude Opus 5.5 and Claude Sonnet 5.5 available in Antigravity tools

October 8 — On the same day, the Gemini Enterprise release notes state that administrators can enable Claude Opus 5.5 and Claude Sonnet 5.5 in Antigravity 2.0, the Antigravity CLI and the Antigravity extensions for IDEs. The administrator accepts their terms of use directly in the Google Cloud console, without a separate Anthropic account.

Activation settingValue specified by Google
Default stateThird-party models disabled, enabled by administrators
Billing modePay as you go, within the project’s spending cap
Inference locationsglobal, us and eu
Thinking levelsLow, Medium (default), High or Max

Developers keep their Gemini Enterprise license and choose Claude in the model selector. Overages must be enabled before accessing a third-party model, whose cost counts toward the spending cap shared by all models. The same page announces general availability of Gemini 3.8 Flash in Singapore.

🔗 Gemini Enterprise release notes


Voice: HeyGen Voice tops the Controlled Voice ranking, Mistral releases two models for Arabic

October 9 — HeyGen launches HeyGen Voice, which Joshua Xu presents as the first voice model developed in-house by the company, previously known for its avatars. The model is available in HeyGen and through the API, at 30 dollars per million characters; until October 31, API users pay 15 dollars, with the discount applied automatically.

The supporting evidence comes from Artificial Analysis, which published its results a minute and a half before HeyGen. In its Controlled Voice ranking, where all models speak using the same 8 cloned voices in American and British English, HeyGen Voice takes first place with an Elo of 1 201 across 1 468 appearances. It ranks first in the Assistants and Customer Service categories, as well as in British English.

HeyGen Voice takes the #1 spot on the Artificial Analysis Controlled Voice TTS Arena Leaderboard, ahead of Alibaba’s Qwen-Audio-3.1-TTS-Plus and ElevenLabs’ Eleven v4 Turbo — @ArtificialAnlys on X

Evaluated modelControlled Voice Elo (October 9)API price per million characters
HeyGen Voice1 20130 dollars (15 until October 31)
Qwen-Audio-3.1-TTS-Plus1 18219,3 dollars
Eleven v4 Turbo1 16640 dollars
Eleven v41 16280 dollars

This first-place ranking has a specific scope: it covers voice cloning. In the Provider Voice ranking, where each provider uses its own voices, Eleven v4 Turbo and Eleven v4 were still in the lead on the evening of October 9, with HeyGen Voice absent from the top eight. For pronunciation robustness, HeyGen Voice scores 83,1 %, placing it 10th out of 29.

On the API side, cloning is instant: a single recording, of which only the first three minutes are used, produces an active voice in seconds, without training. Speech is synthesized in one batch (a WAV file at 44,1 kHz) or streamed, and Artificial Analysis evaluated the API’s default settings.

🔗 HeyGen announcement

Voxtral Mini 4B Realtime Arabic and LIDstral Arabic: two Mistral models for Arabic

October 8 — Mistral posted two open models dedicated to Arabic on Hugging Face under the Apache 2.0 license, without an announcement. Voxtral Mini 4B Realtime Arabic transcribes Arabic dialects and Modern Standard Arabic in a stream, including when speakers switch languages during a conversation (code-switching). Fine-tuned from Voxtral Mini 4B Realtime, it has around 4,4 billion parameters. According to its model card, it was co-developed with the Moroccan Ministry of Digital Transition and Administrative Reform as part of the partnership signed between Mistral and Morocco in January 2026, which also produced two new benchmarks, ISMA and Darija in the Wild.

Published modelModel functionResult according to the model card
Voxtral Mini 4B Realtime ArabicStreaming Arabic transcriptionAverage character error rate of 8,82 % across 7 benchmarks at 480 ms, compared with 7,91 % for Voxtral Transcribe Arabic
LIDstral ArabicLanguage and dialect identification, 51 classes, running on CPUF1 of 88,67 % on Moroccan darija, compared with 71,28 % for GlotLID v3

🔗 Voxtral Mini 4B Realtime Arabic on Hugging Face · LIDstral Arabic on Hugging Face


Claude Managed Agents: dynamic workflows with up to 1 000 agents per run

October 9 — Anthropic opens dynamic workflows (dynamic workflows) in Claude Managed Agents in public beta, introducing a new type of multiagent orchestration. The agent leading the session writes a program itself, which the server executes in the background by launching many agents in phases before combining their results. Once the multiagent_20261001 type is enabled, the agent decides when to launch a run.

A run launches up to 1 000 agents, including 64 at the same time, lasts 24 hours by default, and a session can keep 10 runs open. Its tokens are billed like those of the session and count toward its budget; Anthropic warns that token usage can be high, and cites an internal test involving 70 bugs seeded in a codebase of 116 000 lines.

Anthropic test configurationBugs found across 3 trials
Single agent14, 15 and 27
Dynamic workflow66 in each trial

🔗 Announcement from @ClaudeDevs · Workflow run documentation


Coding agents: message predictions and MXC sandbox for Codex, Claude Code 2.1.296, Vibe CLI 2.26.1

Codex receives two new features on the same day, while Anthropic and Mistral release new versions of their command-line agents.

Codex predicts the user’s next message, in beta for Pro subscribers

October 9 — OpenAI launches composer predictions (composer predictions) in beta in the Codex desktop application. When Codex has finished responding, a suggested next message may appear in the input field, generated from the current thread without consulting memories or connected apps. The Tab key inserts it; the text remains editable and is never sent automatically. The suggestion covers a complete message, rather than word-by-word completion.

Beta requirementValue specified by OpenAI
Plan and agePersonal ChatGPT Pro, 18 and older
Threads and modelsLocal and SSH threads, with GPT-6 Astra or GPT-6.1 Sol
Default settingEnabled, can be disabled in General > Composer > Show predictions
Cost during betaOutside Codex usage limits, consumes no credits

Sent messages, including accepted predictions, continue to count toward usage as usual. The feature is not available in Chat.

🔗 Announcement from @OpenAIDevs · OpenAI help article

Codex on Windows: a sandbox based on Microsoft Execution Containers

October 9 — OpenAI adds a sandbox mode to Codex on Windows based on Microsoft Execution Containers (MXC), which Microsoft made generally available on October 7. It requires a compatible Windows 11 device (24H2 build 26100.9278 or 25H2 build 26200.9278) and promises faster setup, stricter network control and more granular file access.

According to the documentation, MXC becomes the recommended implementation: it natively isolates processes, without administrator-approved configuration, additional Windows accounts or local firewall rules. The elevated and unelevated modes become legacy fallbacks. The desktop application automatically chooses MXC for consumer accounts; in the CLI or in enterprise environments, it is enabled with prefer_mxc = true in config.toml, and an administrator can block it with allow_mxc = false. Two limitations are documented: managed networking requires allow_local_binding = true, and any remaining child processes stop when the foreground command ends.

🔗 Announcement from @OpenAIDevs · Windows sandbox documentation

Claude Code 2.1.296: subagent-specific compaction and a single model for workflow agents

October 9 — Claude Code 2.1.296 contains 79 entries, including 56 fixes, and was released without an accompanying tweet. It primarily provides more control over subagents.

New feature in the releaseWhat it changes
autoCompactWindow (subagents, --agents)A subagent compacts its context earlier than the main conversation
CLAUDE_CODE_WORKFLOW_SUBAGENT_MODELAll agents in a workflow use the same model; other subagents keep their own
Read’s allow_large optionReads a large text file in a single call, if the context allows it
CLAUDE_CODE_OVERLOADED_RETRY_MAX_DELAY_MSLonger maximum delay between retries after a 529 error
MCP tool descriptions sent upfrontDefault limit increased from 2 048 to 4 096 characters

/cost and the status line now count Sonnet 5.5 cache reads at 0,10 dollar per million tokens. On the security side, the release fixes managed PreToolUse hooks that rejected a call without ending the turn, Bash checks that automatically approved certain commands using BASH_ARGV0, cloud sessions that skipped auto mode checks for Claude in Chrome actions, and Edit and NotebookEdit operations that replaced all non-ASCII characters in non-UTF-8 files; those operations are now rejected.

🔗 Claude Code 2.1.296 on GitHub

Vibe CLI 2.26.1: bundled local Devstral model removed, a more complete non-interactive mode

October 9 — Three days after 2.26.0 and without a tweet, Mistral releases Vibe CLI 2.26.1. Despite its patch version number, the release contains 85 entries (23 additions, 24 changes, 36 fixes, 2 removals), including 37 for the new Rust interface. It removes the bundled local Devstral model: any configuration still pinning it falls back, with a warning, to the default hosted model. That model now offers only two reasoning levels, off and high, and reports a 256k context window, which automatic compaction follows.

Non-interactive mode vibe -p gains a time limit, prompt input from a file or standard input, a export.json report written on every exit, and distinct exit codes: 1 for a usage or configuration error, 2 for an infrastructure failure, 3 for a limit reached or a model refusal. The Rust CLI adds /proxy-setup, a narrator that reads each turn’s summary aloud, /plugins, and /whoami. New settings prohibit the agent from using background processes or sub-agents, and the Document Library connector remains disabled by default.

🔗 Vibe CLI 2.26.1 release notes


Generative images, videos, and games: Qwen-Image-2.1-Turbo, Midjourney’s Thinking mode, Synthesia’s Syren, Google’s Playground

A faster, cheaper image model, a generator that reviews its own work, an agent that edits videos, and a platform for games without code.

Qwen-Image-2.1-Turbo: 8 denoising steps and 0,016 dollars per image

October 9 — Qwen releases Qwen-Image-2.1-Turbo, an accelerated checkpoint (accelerated checkpoint) of Qwen-Image-2.1 that generates and edits images in 8 denoising steps, using the same 7-billion-parameter architecture. According to Qwen, it still produces 2K images and accepts natural-language edits, but no quality or speed figures specific to this version have been published. The weights are on Hugging Face and ModelScope, under Qwen’s research license (Qwen Research License). The checkpoint includes its 8-step sampling schedule and runs with a CFG of 1 by default.

On the same day, Qwen officially opens the Qwen-Image-2.1 Pro and Turbo APIs; QwenCloud allows 120 requests per minute for Turbo.

Model served via API (Model Studio)Price per imageFree images
qwen-image-2.1-turbo0,016 dollars10
qwen-image-2.1-pro0,04 dollars10
qwen-image-2.0-pro0,075 dollars100

🔗 Qwen announcement · Hugging Face model card

Midjourney tests a Thinking mode and introduces shared folders on its alpha site

October 8 — Midjourney’s alpha changelog, published in the evening, introduces two early features, currently available only on alpha.midjourney.com. The first is a test: for an image produced with V8.2, whether generated or edited, the Rerun (Thinking) button asks the model to inspect the result, identify where it deviates from the prompt (objects, layout, anatomy, text), and generate it again. Midjourney says it is seeing better results in prompt adherence, typography, and consistency, and is considering making this a general mode.

The second is folder sharing: users can invite Collaborators, who generate directly inside the folder, or Viewers, and can share via a link or even open the folder to any Midjourney member. Midjourney clarifies that sharing does not change image visibility: without stealth mode, images can still appear on Explore and on profiles.

🔗 Midjourney alpha changelog · Thinking mode test (Midjourney)

Syren: Synthesia hands video creation to an agent, in beta

October 9 — Synthesia launches Syren in beta for all its customers, including free accounts. Users describe the video they want, or provide documents, images, videos, PowerPoint presentations, or YouTube links, and Syren delivers a finished video with animated graphics (motion graphics), avatars, voiceovers, music, and cutaway footage (b-roll), which they can then revise through conversation. According to Peter Hill, Synthesia’s chief technology officer, the agent writes the entire video as code, which an in-house rendering engine, used for Synthesia’s animations since June, plays in real time.

Videos can last up to 3 minutes, in landscape or portrait format, and render in the browser. Billing is credit-based, with no detailed pricing, and Enterprise administrators can disable the tool. The post announces that automatic brand identity learning from existing videos is coming soon. Syren arrives four days after HeyGen’s HyperFrames Studio and the day after Anthropic’s Claude Motion, two other tools in which an agent produces the video.

🔗 Introducing Syren (Synthesia)

Playground: Google opens a platform for creating games without coding

October 7 — Google launches Playground, an experimental gaming platform where users create, play, and share games by describing them in a few prompts, without writing code. In a conversational interface, users start from a blank page, an example to remix, or guided assistance, then adjust the physics, rules, characters, or scenery on request.

The games run in the browser, on phones and computers alike, and can remain private, be shared via a link, or join the Explore gallery, with leaderboards and multiplayer for certain genres; every published game undergoes a safety check. Playground is open in the United States to people aged 18 and over, with creation allowances depending on their Google AI subscription. An integration with Unity Spark, for more ambitious 3D games, is announced as coming soon in closed beta. Google does not specify which model powers the platform.

🔗 Introducing Playground (Google blog)


Health: AMIE’s clinical study published in The Lancet

October 8 — Google and Beth Israel Deaconess Medical Center (BIDMC) publish a prospective feasibility study of AMIE (Articulate Medical Intelligence Explorer), Google’s conversational diagnostic AI, in The Lancet; according to Google, this is its first publication in The Lancet’s main journal. Patients interacted with AMIE up to five days before an urgent care appointment, after which the physician received the transcript and a summary.

Study measurePublished result
Patients enrolled (April to November 2025), then followed up114, then 98
Stoppages for safety reasons0
Hallucinations identified1
AMIE useful for preparing the visit, according to the physician33 cases out of 44 (75 %)
Diagnoses matching the final diagnosis (Google)90 %

Conducted at a single site, without a control group, and funded by Alphabet, the study represents neither authorization nor deployment: the authors call for larger trials.

🔗 Google blog post · The article in The Lancet


Research at OpenAI: the response to the letter from three researchers the company parted ways with

October 9 — In a note published on @OpenAINewsroom on behalf of its research leaders, OpenAI responds to the letter from three researchers it says it parted ways with the previous week. According to OpenAI, an internal investigation concluded that they had violated clear rules on handling sensitive information, with a breach of trust extending beyond what their letter describes; OpenAI stands by its decision. It says these decisions were unrelated to raising safety concerns or speaking out, and that it does not fire any employee for expressing concerns.

The note also announces:

We are actively finalizing contracts with third-party safety assessors and will announce details in the coming weeks. — @OpenAINewsroom on X

OpenAI says it agrees with the letter on one point: preserving the monitorability (monitorability) of frontier models requires a commitment from the entire industry, including OpenAI. This article reports OpenAI’s account; the researchers’ letter was not consulted.


Building open models: six models for about 103 dollars with ML Intern, Ai2’s GPU scheduler

Two accounts of model building: one from the agent side, the other from the compute cluster side.

ML Intern: six open models for about 103 dollars in compute

October 8 — Already covered here on October 1 and 8 from other angles, HuggingChat’s ML Intern mode gets its first numerical assessment on the Hugging Face blog: six projects completed from start to finish, for about 103 dollars in total compute costs. Each starts with a message and ends with a public model on the Hub, including evaluation. The agent plans, requests permission before any paid compute task, runs a short test, then trains, evaluates, and publishes using Hugging Face hardware.

Model produced by ML InternPublished resultCompute cost
Pocket Rewriter (0.8B)99,7 % valid outputs, about a quarter of the tokens used by Qwen-Image 2.1’s 9B rewriterAbout 16 dollars
Citrus Doctor (Qwen3.5-2B)From 14,9 % to 52,8 % correct diagnoses of citrus diseasesAbout 1,90 dollars
Two LoRA for Qwen-Image 2.1Viewpoint changes, scribbles replaced with the requested objectAbout 16 and 24,30 dollars
Huggy LoRA (FLUX.2 klein)Character LoRAAbout 7,60 dollars
Agate 4 stepsDistillation of a small image model, from 50 steps (100 passes with guidance) to 4About 37 dollars

🔗 Hugging Face blog post

Ai2’s GPU scheduler: median wait drops from 5 minutes to 24 seconds

October 9 — Ai2 details its new GPU cluster scheduler in a post by Jeremy Tryba. Thousands of NVIDIA H100, B200, and B300 GPUs, in clusters of 88 to 1 024 GPUs, are shared by about 150 researchers, with demand two to three times greater than supply. The old priority-based system left GPUs “occupied” by empty tasks and saw all workloads eventually reach the highest priority.

The new system relies on GPU time budgets distributed along the research organizational chart, hierarchical fair sharing (fair-share) calculated over a rolling seven-day period, and a “scheduling contract”: each workload declares a protected duration of no more than 8 hours, after which it can be preempted and requeued. Unallocated time remains free, but is always preemptible.

Metric measured by Ai2Old schedulerNew scheduler
Median wait, largest H100 cluster5 minutes24 seconds
p90 wait for debugging workloads2 hours30 seconds
Owed GPU hours delivered over 30 days—98 %

Repairs requiring human intervention decrease by 74 %. Ai2 acknowledges limitations, such as interactive sessions becoming preemptible after 8 hours, whereas they previously could last a week, and publishes no code.

🔗 Ai2 blog post


Decision models: pplx-decider-v1.1-27b tops Decision Bench, statistically tied with nine others

October 9 — Already covered on October 6 and 8, Perplexity’s decision model receives an external evaluation result. On Decision Bench, Atlan Frontier Labs’ open source benchmark (1 071 closed-ended questions, 36 datasets, 15 ranked models), pplx-decider-v1.1-27b achieves the highest raw score, 94,5 %, at the lowest cost, 0,017 dollars per 1 000 decisions. However, the site assigns the same rank to models whose 95 % confidence intervals overlap: ten models share first place.

Model evaluated on Decision BenchMeasured accuracyCost per 1 000 rows
pplx-decider-v1.1-27b94,5 %0,017 dollars
DeepSeek V4.1 Flash94,2 %0,683 dollars
Gemini 3.5 Flash94,2 %3,32 dollars
Jev 1.1393,8 %0,044 dollars

Perplexity also publishes a practical guide (cookbook) for a browsing agent, in which the code, never the model, decides what to click, and v1.1 replaces v1 in the API.

🔗 @perplexitydevs thread · Decision Bench leaderboard


Grok Bot gets its own email address

October 9 — Grok Bot, SpaceXAI’s agent, now has its own email address. The Bot can use it to sign up for services, contact companies on the user’s behalf, or schedule an appointment with someone. The feature rolls out the same day: users simply ask the Bot for an address or mention @bot on X, and within a team, an administrator must enable it first.

The announcement consists of a thread from the @bot account, reposted by @grok, with no post on x.ai. It specifies neither the format nor the domain of the addresses, nor which plans are eligible. Following the dedicated Slack identity for Team Bots in late September, the agent now gains its own identity for acting outside the platform.

🔗 Grok Bot announcement


ElevenLabs establishes a presence in Singapore, its hub for Southeast Asia

October 8 — ElevenLabs makes its arrival in Singapore official, opening an office intended to serve as a hub for Southeast Asia and, more broadly, Asia-Pacific. The post highlights the infrastructure already in place: since July, businesses have been able to host their data in Singapore, with zero-retention options and reduced latency in the region.

ElevenLabs primarily cites its regional customers: Funding Societies, which qualifies leads using conversational AI across five markets, Boost Bank in Malaysia, Salmon Group in the Philippines, MNC Group in Indonesia, and Rezonate, a healthcare platform used by 60 % of Singapore’s public hospitals. No headcount is provided. The expansion follows those in Brussels and Amsterdam in late September and early October.

🔗 ElevenLabs blog post


Briefs

  • Projects in Claude Code — Anthropic has granted access to all Pro and Max subscribers on the waiting list. Projects remains in beta: the list is still open, more access will follow as capacity allows, and the feature is not yet available on Team or Enterprise, according to the documentation. 🔗 source
  • Guide to agent automations — A claude.dev guide describes a reference implementation on Claude Managed Agents that reads Slack channels and GitHub pull requests on a schedule, then posts a summary in Slack. It details the six components and their safeguards against common failures, such as avoiding mistaking an unreadable source for a quiet day or setting the spending cap at 3 to 5 times the cost of a normal run. 🔗 source
  • Block and Claude Fable — In an interview published by Anthropic, Block explains that Claude Fable designs its large code migrations and directs as many as dozens of Opus or Sonnet models, merging up to a thousand pull requests in a single migration; merges and production releases remain reserved for humans, with dual approval for deployments. 🔗 source
  • Claude Enterprise’s Compliance API — Its conversation endpoints (endpoints) also return conversations from the unified Claude experience, in beta and using the existing key: a conversation continued in a cloud session is returned as a single conversation, with Claude’s work in tool_use and tool_result blocks. 🔗 source
  • Codex CLI 0.162.1 — This patch to 0.162.0 fixes a TUI crash on multiline asynchronous questions, along with startup failures when an already running background server had different feature settings from the CLI. 🔗 source
  • Plugins by role in ChatGPT Enterprise and Edu — When role-based access control (RBAC) is enabled, owners and administrators configure plugin access for each role: Available, Install or Disabled for the workspace, and Default for a custom role to inherit that setting. 🔗 source
  • Sophos and OpenAI Daybreak — According to the cybersecurity vendor, its agents built with Daybreak reduce the average response time for the cases they handle from around 38 minutes to around 89 seconds, and 52% of its MDR cases are resolved end to end by AI; destructive actions remain under human supervision. 🔗 source
  • Asana and StackAI’s browser agent — According to Asana, GPT-6 Astra in Codex optimized this agent in about a week: on GPT-6.1 Sol, a test run costs $0.47 and takes around four minutes, making it 76 times cheaper and 5 times faster than the original production configuration. 🔗 source
  • LegalOn and Codex — LegalOn Technologies says it reduced its estimated daily Codex costs by around 65% by moving from GPT-5.5 in unlimited fast mode to distributing tasks across GPT-6 Luna, GPT-6.1 Sol and GPT-6 Astra, with fast mode on demand and caps per department and per person. 🔗 source
  • Gemini CLI v0.64.0-preview.1 — This preview suppresses false positives from the Untrusted Command Flags Detected warning on harmless shell commands: read-only options are exempted, shell variables are preserved, and loops are allowed when each subcommand has its own ALLOW rule. No nightly was released on October 9, and the stable version remains v0.63.0. 🔗 source
  • Antigravity 2.22.0 — Plugins can display their own interface in the side panel (UI Extensions), conversations can be organized by dragging and dropping them between sections of the sidebar, and timestamps display the month name. 🔗 source
  • Gemini 3.7 Flash deprecated — In the Gemini API, calls to gemini-3.7-flash are redirected to gemini-3.8-flash, and the old deep-research-pro-preview-12-2025 agent will be shut down on October 23 in favor of deep-research-preview-04-2026 or deep-research-max-preview-04-2026. 🔗 source
  • VST3 and AU plugins for Google Flow Music — Catching up on October 6: Spaces, the instruments and effects created using natural language, can be exported as plugins for digital audio workstations (Digital Audio Workstations), such as producer Khris Riddick-Tynes’s No Chaser plugin. 🔗 source
  • Google Earth AI and public health — Catching up on October 6: during the Ebola outbreak in the DRC, a prototype geospatial reasoning agent enabled the WHO Regional Office for Africa to identify 48 exposed localities and more than 45,500 people at risk within minutes; the PDFM model also identifies areas at risk of cholera up to 8 weeks in advance. 🔗 source
  • Copilot code review billing — Organization owners can charge reviews requested by their licensed members to the organization that owns the repository, instead of using those members’ quotas, if paid AI Credits usage is enabled. They can also restrict review requests to licenses provided by the organization or enterprise. 🔗 source
  • Copilot CLI 1.0.95 — Now stable, it authenticates on macOS through Microsoft Entra’s native broker, with a browser fallback; preview 1.0.96-0 shows in the timeline who made each permission decision: the user, Assisted Permissions, a policy or the unattended fallback. 🔗 source
  • GitHub MCP Server 2.0 — Catching up on October 6: GitHub’s official MCP server returns typed outputs, described by output schemas (output schemas), for agents using Code Mode or Programmatic Tool Calling; these are advertised only to clients using the MCP 2026-07-28 specification or later. 🔗 source
  • Claude Haiku 5.5 in Genspark — Genspark is rolling out the model in AI Chat, Code Agent and Claw, echoing Anthropic’s claim that it is around 75% cheaper to run than Haiku 4.5, with no dedicated plan or pricing. 🔗 source
  • Meta VR template for v0 — Meta is releasing a v0 template with Vercel based on the Immersive Web SDK: describe an experience, and v0 codes and previews it; one click deploys it on Vercel so it can be opened in the Quest browser. Gaze and hand input for Meta VR Glasses, expected in spring 2027, are already supported. 🔗 source
  • v0 for teams — v0 presents this way of viewing teammates’ work, joining their conversations and building together in a video; the tweet highlights features introduced in stages since September, with no new pricing. 🔗 source
  • DeepSeek Harness 0.2.1-alpha.2 — The 26th preview, still without a stable release: the plugin catalog installs Claude Code and Codex packages on demand, two experimental plugins appear (Git Worktrees, reasoning translation), and global instructions can come from a shared AGENTS.md. 🔗 source
  • OGX 1.5.0 — Formerly Llama Stack, renamed in April and moved out of the meta-llama organization, it adds native passthrough for the /v1/messages API for DeepSeek and Fireworks AI, a Text-Embeddings-Inference provider, Exa and Serply web search, and the MCP 2.x client, with four breaking changes. 🔗 source
  • GenIA — Without an announcement, Meta has published the code for this work conducted with the Tübingen AI Center under the noncommercial CC BY-NC 4.0 license: it reconstructs 3D objects from an image, a few views or a video by aligning the frozen SAM 3D Objects prior at inference time, without retraining. 🔗 source
  • decision-model tag on the Hub — Hugging Face gives decision models their own tag, with an icon and a filter on the Models page; decision GGUFs are tagged automatically using their llama.cpp metadata. 🔗 source
  • Darwin-27B-ZTC-v2 — VIDRAFT announces that the new version of its judge, under Apache-2.0, leads the System One Mosaic Benchmark (137 benchmarks) with a Borda score of 89.58, ahead of OpenJev-27B (87.50) and Jev 1.13 (85.05); a ranking reported by its authors, who present it as a snapshot of a public leaderboard. 🔗 source
  • Decision model confidence — In a community post, Stephen Solka shows that GPT-6 Luna Decisions, queried through OpenRouter, answers “does not match” to 600 SHA-256 checks, yielding 50% accuracy and 99.8% average confidence; in causal reasoning, a 99% threshold isolates 47 correct answers out of 49, but covers only 8.2% of cases. 🔗 source
  • openai-math as reinforcement learning environments — FineEnvs ports 369 of the 405 results formalized in Lean from the openai/math publication into Harbor environments: the agent must prove a theorem in a Lean 4 sandbox, and Comparator awards a reward only for a complete proof of the exact statement; an experimental dataset under Apache-2.0. 🔗 source
  • JOSIE-2 — Gökdeniz Gülmez publishes the technical report for this family of 2B, 4B and 9B models post-trained from Qwen3.5 on two M4 Macs, using around 3,000 selected examples: in reasoning mode, JOSIE-2-4B reaches 95.5 on ARC-Challenge (+12.1) and 69.2 on TruthfulQA (+20.3), on just two benchmarks. 🔗 source
  • Gradio 6.30 — gr.Workflow gains an application view, and the Run this node command reuses upstream results that are still current instead of rerunning the entire subgraph; gr.ImageEditor preserves the alpha channel of loaded images. 🔗 source
  • tokenizers 0.23.3 — This patch to the Python package alone supports huggingface_hub v2, unblocking Transformers installation according to the notes, and reverts a breaking change to truncation introduced in 0.23.1. 🔗 source
  • AI scientific peer review — Sakana AI presents a paper accepted at TMLR: a benchmark inserts contradictory claims into papers to check whether AI reviewers detect them, and its Multi-Layered Review system detects more than the other systems tested, according to Sakana AI’s announcement, which provides no figures. 🔗 source
  • Suno Studio — Clips align better with the downbeat and stretch to follow the project tempo, fades become exponential or logarithmic, and the synthesizer, keyboard and plugin views can be detached and moved to a second screen; Studio is included in the Premier subscription. 🔗 source
  • World Models’ Last Exam in Physics — On this physics exam from Einsia for video models (40 tasks, 9 categories, eight models), Seedance 2.5 leads with 57.76 out of 100, ahead of MiniMax H3 (54.89), the best open model according to MiniMax, and NVIDIA’s Cosmos 3 Super (42.46); a third-party benchmark. 🔗 source
  • Omniverse simulations built by agents — The NVIDIA blog shows company teams having GPT-6 Astra, and in one case Claude Fable 5, assemble simulations using the ovphysx, ovstage, ovrtx and ovui libraries, including digital twins created or improved in around three days; no new product. 🔗 source
  • Managing a Shopify store — Once the store is connected, Grok Bot can check orders, track inventory and keep product listings up to date; no plan or limits are specified. 🔗 source
  • Searching X — Catching up on October 7: Grok Bot can search, read and monitor X for all users without configuring the X connector, for example to track product feedback or receive a weekly industry summary; at the end of August, connecting an account was still required. 🔗 source
  • Grokipedia v0.3 — According to its official account, the encyclopedia written by Grok published more than 4,400 reader-requested articles in a week and makes more than 2,200 edits per day, with an average of 78 sources per new article; a page now streams its edits live. 🔗 source
  • mistral CLI 0.8.0 — mistral apps deploy also deploys workflow workers and all modules at once, including in continuous integration using just an API key, and mistral apps dev starts a local Postgres database in Docker; in return, the .env in the current directory is no longer read. 🔗 source
  • Mistral Python SDK 3.2.0 — Chat and agent calls support log probabilities (logprobs) for both the response and the prompt, as well as top_k, a repetition penalty and a minimum token count; an automatically generated release, without an announcement. 🔗 source
  • Qwen Code v0.25.1-preview.1 — The second preview of v0.25.1 in three days, with 16 new features and 27 new fixes, including session-centered multi-agent collaboration and PreToolUse hooks that modify a tool’s input with full revalidation; the stable version has not yet been released. 🔗 source
  • Claude Haiku 5.5 in Perplexity’s Agent API — Catching up on October 7: the Agent API supports anthropic/claude-haiku-5-5 at Anthropic’s pricing, $0.10 per million input tokens and $0.50 per million output tokens up to 100,000 input tokens, then $0.50 and $2.50 beyond that. 🔗 source
  • Shared or dedicated inference — A Cohere guide for Embed and Rerank argues that the shape of requests matters more than their number, and calculates illustrative break-even thresholds against an instance dedicated to an NVIDIA A10 GPU: around 4 requests per minute for long documents, 20 for a catalog and 29 for reranking (reranking). 🔗 source
  • Agent or MCP — Perplexity publishes a guide that distinguishes the agent, which makes decisions, from MCP, which connects it to tools, with a comparison table, an online store example and four risks with their safeguards; according to the guide, Claude, ChatGPT or Cursor can delegate a task to Computer through Perplexity’s MCP server. 🔗 source

What this means

One agent now orchestrates many others, and each is given an identity. The Gemini agent launches sub-agents, each with its own identity, and can become a coworker agent with its own address; Claude Managed Agents lets an agent write a program that launches up to 1,000 of them; Grok Bot gets an email address to act externally, and Block puts Claude Fable in charge of dozens of models working on the same migration. The emerging question is no longer whether the agent can do the job, but how to keep it under control. Google responds with Agent Gateway, Agent Sandbox, an audit log for each agent, and real-time spending caps; Anthropic with a session budget that suspends all runs once reached; OpenAI with a Windows sandbox that applies stricter network and file controls.

Platforms themselves are embracing multiple models. Google, which develops Gemini, has its own agent orchestrate Claude models, explicitly separates model selection from agent selection, and makes Claude Opus 5.5 and Sonnet 5.5 available in Antigravity, billed under the same spending cap as its own models. Genspark and Perplexity’s Agent API add Claude Haiku 5.5 in the days following its launch, LegalOn divides its coding work among three OpenAI models according to the nature of the tasks, and Block is building an automatic model selector.

Unit costs continue to fall, and every announcement puts a number on it: $0.016 per image for Qwen-Image-2.1-Turbo, 2.5 times cheaper than the Pro version; $15 per million characters for HeyGen Voice until October 31, then $30, less than Eleven v4 Turbo; $0.017 per 1,000 decisions for pplx-decider-v1.1-27b; six models published on the Hub for around $103 in compute with ML Intern. Top rankings, however, require careful reading: HeyGen Voice only leads the ranking with controlled voices, and pplx-decider shares its rank with nine other models whose confidence intervals overlap.

In healthcare, caution remains warranted. The AMIE study published in The Lancet is a feasibility study, conducted at a single site, without a control group, and funded by Alphabet: 98 patients, no conversations stopped for safety reasons, and one hallucination identified. It is a step toward larger trials, which the authors themselves call for, not an approval or a rollout to patients.


Sources