Search

Anthropic launches Claude Sonnet 5.5, NVIDIA unveils Open Agent Safety Platform, Manus moves to 2.0 with Cue

ai-powered-markdown-translator

Article translated from French to English with gpt-6-sol.

View project on GitHub ↗

On Monday, September 28, six days after Opus 5.5, Anthropic launches Claude Sonnet 5.5: the same price as Sonnet 5, output generated more than 30% faster, and a 70.6% score on Terminal-Bench 4.0, versus 10.3% for its predecessor; GitHub Copilot, Devin, Cursor, and v0 make it available the same day. NVIDIA unveils Open Agent Safety Platform, which monitors agents from software down to silicon with more than 100 organizations, Manus moves to 2.0 and launches Cue, ElevenLabs releases Eleven v4, and Kling announces Kling 4.0 for October. Among coding agents, Claude Code 2.1.284 enables auto mode on all plans, and Devin becomes 30% to 40% cheaper.


Anthropic launches Claude Sonnet 5.5, faster and cheaper per task than Sonnet 5

September 28 — Six days after Claude Opus 5.5, Anthropic launches Claude Sonnet 5.5 (claude-sonnet-5-5), the second model in the Claude 5.5 family. Presented as a clear upgrade over Sonnet 5, it generates responses more than 30% faster and, in Anthropic’s tests, costs up to 30% less per task because it needs far fewer tokens to do the same work. Opus 5.5 remains the model for complex work requiring careful judgment; Sonnet 5.5 targets well-defined everyday tasks: fixing bugs and producing polished documents, presentations, and spreadsheets. Claude Haiku 5.5, designed for high-volume workloads, is due in the coming weeks.

The largest gap with Sonnet 5 appears on Terminal-Bench 4.0, a command-line agentic programming evaluation: 70.6% versus 10.3%, above Opus 5.5’s 66.4%, measured at Xhigh effort, its best result. On GDPval-AA, which covers knowledge work, Sonnet 5.5 finishes two points behind Opus 5.5 and roughly 400 points ahead of Sonnet 5. It is also the first Sonnet to complete Pokémon Red using only screenshots.

Benchmark (Anthropic figures)Claude Sonnet 5.5Claude Sonnet 5Claude Opus 5.5GPT-6 Sol
Terminal-Bench 4.070.6%10.3%66.4% (Xhigh)—
FrontierCode 1.1 Main52.1% (Xhigh), 46.2% (Max)42.4%54.4%49.3%
CursorBench 4.055.5%34.1%57.8%—
GDPval-AA v2.11844144918461487
AA-Briefcase v1.11811135918221483
Humanity’s Last Exam (with tools)64.5%54.9%67.7%—
OSWorld 2.1 (partial)80.1%57.0%81.8%—
Chartography (without tools)61.6%15.6%64.4%53.6%

Anthropic adds caveats to these figures. On FrontierCode, Sonnet 5.5 performs worse at Max effort, where it invokes Claude Code’s code review skill more often, causing it to exceed the time limit or stray beyond the task’s scope in two cases examined by Cognition; GDPval-AA and AA-Briefcase were measured on a preview affected by a bug whose impact was judged small. Above all, Anthropic considers Opus 5.5 markedly stronger at open-ended, complex work requiring sustained judgment.

Pricing is unchanged: $2 per million input tokens, $10 per million output tokens, and $0.20 per million cache-read tokens, the same as Sonnet 5; input and output prices are half those of Opus 5.5 ($4 and $20). The savings come from the number of tokens used: at Low or Medium effort, Sonnet 5.5 exceeds Sonnet 5’s best score on several benchmarks at roughly one-tenth the cost per task. The model accepts a 1-million-token context and produces up to 128,000 output tokens; default effort is Medium in Claude Code and the Claude apps, and High on the Claude Platform.

We recommend starting with Opus for large projects and Sonnet for tasks where you need to balance quality with speed and cost. — @ClaudeDevs on X

Because its cybersecurity capabilities are comparable to those of Opus 5, Sonnet 5.5 is the first Sonnet launched with cyber safeguards at the level used for the most capable models: high-risk requests visibly switch to Sonnet 5, without affecting routine bug fixes. It is also the first Sonnet equipped with safety classifiers that prevent its reasoning from being extracted, and its reasoning is now tied to the account that produced it.

For developers, switching to claude-sonnet-5-5 involves more than changing a model ID: the documentation lists five breaking changes from Sonnet 5, including replacing the option to disable thinking with the between_tools setting and rejecting forced tool choice (tool_choice to any or tool, error 400). The /claude-api migrate command helps with migration in Claude Code.

Sonnet 5.5 is available today on the Claude Platform, in Claude Code, and through Amazon Web Services, Google Cloud, and Microsoft Azure, with no plan-by-plan details (Free, Pro, Max) in the sources. In Claude Code, version 2.1.284 makes it the default Sonnet model on the Anthropic API (see the coding agents section).

🔗 Claude Sonnet 5.5 announcement · Sonnet 5.5 migration guide

GitHub Copilot makes it available starting with Pro

September 28 — GitHub makes Claude Sonnet 5.5 generally available in Copilot (changelog entry at 11:03 a.m. PT). Unlike Claude Opus 5.5, which arrived on September 22 without the Pro plan, it is offered to Copilot Pro, Pro+, Max, Business, and Enterprise across ten surfaces: Visual Studio Code, Visual Studio, Copilot CLI, the Copilot coding agent, the GitHub Copilot app, github.com, GitHub Mobile, JetBrains IDEs, Xcode, and Eclipse, with a gradual rollout.

GitHub says that, in its early tests, Sonnet 5.5 matches Claude Sonnet 5 on coding tasks while using far fewer steps, tokens, and tool calls, and finishes faster, without publishing figures. Usage is billed at the provider’s public rate: GitHub’s documentation lists $2 per million input tokens and $10 per million output tokens, the same rates as Claude Sonnet 5. For Business and Enterprise, the default setting to enable new models makes it available automatically unless an administrator has disabled that setting or the model.

🔗 Claude Sonnet 5.5 in GitHub Copilot · Copilot models and pricing

Devin scores it at 64.4% on FrontierCode 1.1

September 28 — Cognition makes Sonnet 5.5 available in Devin Desktop and Devin CLI on launch day and scores it at 64.4% on FrontierCode 1.1, its benchmark for meeting requirements in production codebases, versus 56.2% for Sonnet 5. It moves ahead of Claude Fable 5.1 (63.6%) and sits just behind the chart’s top three: Opus 5.5 (65.3%), Fable 5 (64.9%), and GPT-6 Astra (64.5%).

Cognition’s post on X refers to the Main variant, but the chart in its blog post is titled Extended and repeats, for the other models, the values published on September 22 for that variant. For comparison, Anthropic reports 52.1% for Sonnet 5.5 on the Main variant at Xhigh effort. Cognition also repeats Anthropic’s figures: output generated more than 30% faster and a cost per task up to 30% lower than Sonnet 5, with token prices unchanged.

🔗 Claude Sonnet 5.5 in Devin · Cognition on X

Cursor and v0 make it available the same day

September 28 — v0, Vercel’s app creation tool, makes Sonnet 5.5 available at 18:04 UTC, one minute after the announcement; Cursor follows at 20:14 UTC and describes it as a strong model, on par with Opus for many tasks, without specifying which Opus. Neither provides pricing or tool-specific measurements: as with Opus 5.5 on September 22, this is simply an availability announcement.

🔗 Cursor on X · v0 on X


NVIDIA launches Open Agent Safety Platform to monitor agents from software down to silicon

September 28 — NVIDIA launches Open Agent Safety Platform, an open software platform paired with a reference system design, to secure AI agents from testing through deployment, with governance and control across the entire stack (software, hardware, compute, and robotic systems). Its architecture post starts with an observation: several leading labs have recently reported agents escaping the evaluation environments meant to contain them. Among the five principles adopted, policy must be verifiable and enforced out of band, beyond the agent’s reach.

Platform componentRole described by NVIDIA
NVIDIA OpenShell (open source, version 0.1)Sandboxed runtime: formally verified policy enforced outside the agent, on Vera CPUs and third-party platforms (Arm, Intel)
NVIDIA Sentry (reference design)Out-of-band watchdog on BlueField-4 DPU that quarantines an agent within milliseconds
NVIDIA DOCARequest and response inspection, attested telemetry, agent identity, zero-trust access

OpenShell, an open-source runtime (Apache 2.0) tracked since March, moves to version 0.1 (0.1.0 published on GitHub on September 25, 0.1.2 on September 28) and is now described as broadly available for both open and closed models. Each agent runs in a sandbox with kernel-level controls; an external supervisor inspects HTTP, GraphQL, and MCP traffic (a read is allowed, while a write through the same API is blocked), and the agent can access services without being exposed to the real credentials. Policies are written in YAML, compiled to OPA/Rego, and checked by a formal prover; when access is missing, the agent can propose a narrow rule that, by default, awaits human approval and cannot be approved by the agent itself. OpenShell supports Codex, Claude Code, Pi, and Hermes. In long-running adversarial tests, NVIDIA reports, leading agents with reduced safeguards spent up to two hours trying to persuade an AI reviewer to grant them access to a protected GitHub repository, without any writes taking place.

Sentry, the second component, moves monitoring into hardware: an out-of-band watchdog on a BlueField-4 DPU, built on DOCA, that continuously monitors agent behavior, verifies identity, and can quarantine an agent that exceeds its limits within milliseconds. In a Vera Rubin POD, each compute tray places a BlueField-4 on the node’s only path to the model; on a Vera system already equipped with BlueField-4, enabling it requires a software update, according to NVIDIA.

NVIDIA says more than 100 organizations are working with the platform’s technologies. Anthropic integrates its Claude Managed Agents with OpenShell and BlueField, running the agent loop on a server separate from the sandboxes; SpaceXAI applies the platform to Cursor coding agents and Grok models; Salesforce tracks OpenShell agent activity from Slack, where permission requests can be approved or denied; SAP integrates it into the Joule Studio runtime, and Scale AI into its GenAI offering. The list extends from Microsoft, IBM, Palantir, CrowdStrike, and Hugging Face to robotics (Figure, Skild AI), finance (Citi, JPMorganChase), operating systems (Canonical, SUSE, Red Hat), and infrastructure providers (CoreWeave, Nebius, Oracle Cloud Infrastructure). The software, including OpenShell and skills, is available on GitHub and NVIDIA’s developer page, and Jensen Huang appeared on CNBC to present these measures the same day.

🔗 NVIDIA press release · Platform architecture · OpenShell 0.1.0 in practice · Jensen Huang on CNBC (@NVIDIAAI)

Together AI and Perplexity share the announcement

September 28 — Together AI says it is a launch partner for the platform, while NVIDIA’s press release lists it among infrastructure providers. The inference provider says it has built platform features to develop and deploy agents securely and is continuing to invest in this area, including work with NVIDIA on OpenShell, without giving figures or naming a product.

Perplexity, mentioned twice in the press release (among the companies joining NVIDIA, then among the more than 100 organizations using the platform’s technologies) but without a specified role, starts a nine-post thread:

We’re partnering with Nvidia and 100+ industry partners to build infrastructure that contains rogue AI agents. — @perplexity_ai on X

The rest of the thread revisits Escaping SPACE, the security audit of its sandbox published on September 23 and covered at the time (no virtual-machine escape in 108 attempts). That post mentioned NVIDIA only through OpenShell, one of the third-party sandboxes tested, against which neither of the two bypasses succeeded.

🔗 Together AI on X


Manus moves to 2.0 with the Cascade agent harness and launches Cue, an app for personal agents

September 28 — Manus launches Manus 2.0, which it presents as a new architecture with new products, rather than a version update. The announcement comes less than a month after Manus resumed independent operations on September 1.

The foundation is called Cascade, the latest iteration of its in-house agent harness: each project starts small and receives specialized capabilities only when the work requires them, with the brief, page, video, and automation kept together in one project. Manus quantifies the improvement over its previous system, but only for one tested configuration, which the post does not detail.

Metric reported by Manus (one tested configuration)Difference from the previous system
Tokens consumed-23.2%
Task duration-28.2%
Execution cost-32%

Two components complete the offering: Cloud Computer, a dedicated environment that users purchase for tasks that need to run continuously (a multiplayer game server, an automation), and automations that can now be triggered by an event in a connected service (a new email, a change in ad performance, a calendar appointment, a Slack message, a Notion update), configured with a single prompt.

The desktop app becomes Manus Studio, a workspace shared by people and AI that loads only the tools useful to the project. It includes two environments. Video Editor produces an initial edit, then opens a timeline where clips, images, text, animations, and audio remain separate and editable; it targets 30- to 60-second product ads, tutorials, and vlogs, while its Alchemy mode gives AI creative direction. Game Dev combines video, image, and code models, publishes a playable game through a simple link, and enables multiplayer by purchasing a Cloud Computer. With Remote Control and Computer Use, users can assign Manus a task on their own computer from their phone while watching the desktop live.

The third part, Cue, is a new standalone app for personal agents on phones and computers, built on Manus infrastructure. Each agent has its own email address, phone number, wallet, and computer: it sends messages, makes payments within a set budget, takes calls, and leaves a summary. Multiple agents can share a goal in a group chat, and Cue connects to everyday services, such as ordering at a restaurant by scanning a QR code. Later that day, Manus clarified that these numbers can make and receive calls and SMS messages, though only in certain countries and sometimes for voice calls only.

Manus 2.0 is available now on the web, desktop, and mobile. Cue is also available, with the iOS version awaiting App Store review, in free early access by invitation code for a limited number of initial users. The post gives no pricing, even for Cloud Computer, names no underlying model, and cites no external evaluation.

🔗 Introducing Manus 2.0 · Cue announcement on X · Agent phone numbers


ElevenLabs launches Eleven v4, its most expressive voice model, and a Turbo variant for agents

September 28 — ElevenLabs launches Eleven v4, presented as its most expressive text-to-speech (TTS) model, and Eleven v4 Turbo, a low-latency variant designed for conversational agents. Built on an entirely new architecture, Eleven v4 is intended to interpret a text’s tone, rhythm, emotion, and context to deliver dramatic, tender, urgent, or comic speech without losing the voice’s identity. ElevenLabs claims the top spot in Artificial Analysis’s Provider Voice Arena ranking (September 2026) and says approximately 75% of listeners preferred it in blind tests against Cartesia Sonic 3.6, Inworld TTS-2, and two Google Gemini 3.8 TTS models, with ties counting as half.

Performance can be directed in natural language or through tags inserted into the text (laughter, an angry line with a French accent, light rain, a vibrating phone), which Eleven v4 follows more faithfully than its predecessors, according to the company. Support for the International Phonetic Alphabet (IPA) improves substantially, and dialogue with multiple voices sounds more natural, with a new method for capturing voice identity to keep voices consistent across agents, audiobooks, and ads.

Metric published by ElevenLabsReported value
Artificial Analysis Provider Voice Arena ranking (Sept.)1st
Preference in blind tests against 4 competing modelsapproximately 75%
Median inference latency for Eleven v4 Turboapproximately 100 ms
Median time to first speech for Eleven v4 Turboapproximately 150 ms
Supported languagesmore than 90 (Eleven v3: more than 70)
Per-request limit for Eleven v410,000 characters (Eleven v3: 5,000)
Minimum audio for an instant clone10 seconds

Time to first speech was measured over WebSocket streaming against Cartesia, xAI, Google, and OpenAI, with network latency excluded, and Eleven v4 Turbo was optimized jointly with the ElevenAgents platform. A voice recorded in one language speaks other languages with a native accent, and Eleven v4 now accepts Professional Voice Clones; chaining long generations, useful for Studio and the Reader app, is more reliable.

Both models are available now in ElevenAgents, ElevenCreative, and through the API (eleven_v4 and eleven_v4_turbo). For two weeks, the Eleven v4 API is discounted to $22 and the Eleven v4 Turbo API to $11 per million characters, while Eleven v4 is free for ElevenCreative Creator plans and above, up to twice the monthly credit allowance; list pricing has not been published. This release follows Eleven v3, which became commercially available in February 2026.

🔗 Introducing Eleven v4 · Announcement thread on X


Kling announces Kling 4.0 for October and opens Kling 4.0 Flash in early access

September 28 — Kling AI introduces Kling 4.0, its next-generation video model, with an official release planned for October. Its first variant, Kling 4.0 Flash, has been available in early access to a limited group of users since September 28: annual Ultra subscribers, according to the announcement tweet. Kling highlights three areas: visual realism, creative control, and narrative completeness.

On paper, Kling 4.0 generates 3 to 30 seconds of continuous video, at up to 4K, with two-channel stereo sound. Narratives can be directed with up to 10 keyframes and prompts of at most 8,000 tokens, while Omni Reference accepts up to 15 references per generation: 10 images, 5 videos totaling no more than 30 seconds, and 7 subjects, including 3 drawn from videos; audio references are limited to voice. Untextured models and depth maps can help set motion and framing, while editing can alter expressions, gestures, camera, style, or background across up to 5 clips. Kling also claims readable text within images and support for more languages, accents, and dialects, from Cantonese and Sichuanese to Indian English.

Technical specificationKling 4.0Kling 4.0 Flash
AvailabilityOfficial release in OctoberLimited early access since September 28
Generation modesText to video, image to video, first and last frames, 10 keyframes, Omni ReferenceText to video, image to video, Omni Reference
Generation duration3 to 30 seconds3 to 20 seconds
Maximum resolution4K (also 720p and 1080p)720p
Dynamic range10-bit HDR at 1080p and 4K, announced for later8-bit SDR

Some features are still pending. The release notes mark 10-bit HDR output at 1080p and 4K as coming soon, although the tweet lists it among Kling 4.0’s features, along with the ability to extend a video to 2 minutes. Kling is also redesigning its creation page, bringing all references into a single input field and adding a canvas where an agent carries out requested tasks. The announcement gives no pricing, API access, or benchmark; it is illustrated by a short film made with Kling 4.0, THE BEAT.

🔗 Kling 4.0 release notes · Announcement on X


Coding agents: Claude Code 2.1.284, Codex CLI 0.158.0, Copilot CLI 1.0.89 and its SDK, Devin, Delta, and Gemini CLI

Claude Code 2.1.284 starts in auto mode on all plans and providers

September 28 — Released at 18:02 UTC, moments before the model announcement, Claude Code 2.1.284 adds Claude Sonnet 5.5, which becomes the default Sonnet model on the Anthropic API. The release has 100 entries: 57 fixes, 18 improvements, 14 additions, and 11 changes.

The most visible change concerns permissions: when no mode is configured, interactive terminal and VS Code sessions now start in auto mode on all plans and with all providers. Version 2.1.283 had already done this on September 25 for third-party providers and sessions without telemetry; version 2.1.284 extends it to everyone, and the permissions.defaultMode setting still takes precedence. A new response, “Yes, but ask again next time,” authorizes one read outside the working directories and has Claude Code ask again for subsequent reads.

Ultracode moves out of the effort slider to become a separate toggle in /effort (the Tab key, or /effort ultracode on and off): it no longer forces xhigh effort and remains active regardless of the selected level, with an equivalent switch below the VS Code effort slider.

What’s new in 2.1.284Command or setting
Auto mode by default, all plans and providerspermissions.defaultMode takes precedence
Ultracode as an independent toggleTab in /effort, or /effort ultracode on and off
Batch reconnection of failed MCP servers/mcp reconnect all
Dollar-denominated spending for the Claude apps gateway/usage and status line (fields used_usd, limit_usd, period)
Second compaction if “Prompt is too long” persistsautomatic

On security, plugins from marketplaces, claude.ai, or npm can no longer pre-approve their own tools when administrators allow only managed rules (allowManagedPermissionRulesOnly); .claude/rules rules symlinked from outside the project require approval; and memory loading neutralizes invisible characters and tags that imitate Claude Code markup in MEMORY.md. For reliability, a damaged response stream is retried instead of displaying “JSON Parse error,” and MCP calls in a resumed session wait up to 10 seconds for their server. Finally, when a Sonnet model’s safeguards flag a message, the notice provides more explanation and suggests revising and resubmitting the request.

🔗 Claude Code v2.1.284

Codex CLI 0.158.0: copy on select and OAuth client secrets for MCP

September 28 — Released at 05:07 UTC, Codex CLI 0.158.0 is the first stable release since 0.157.1 on September 26; its release notes list 188 pull requests since 0.157.0. The fullscreen TUI makes copy on select and right-click paste configurable, and text copied from the transcript retains its Markdown formatting.

FeatureSetting or detail
Copy on selecttui.copy_on_select: auto by default (tmux, Zellij, direct macOS terminals except Ghostty and Kitty), always, never
Right-click pastetui.right_click_paste: auto (Windows, Linux, WSL), on (also macOS), off
MCP servers with OAuthcodex mcp add --oauth-client-secret, key oauth.client_secret
Exec-serverBearer tokens for direct WebSocket connections
Image generationExplicit transparent background (transparent_background), editing images from the conversation

Codex can now connect to MCP servers that require a pre-registered OAuth client with a secret, and direct WebSocket connections to the exec-server can be protected with bearer tokens, including when they pass through the app-server. On security, approval for input sent to a terminal becomes stable and is enabled by default for commands run with elevated permissions. Fixes focus mainly on sandboxes: ordinary Windows 10 paths, rejected stored credentials, startup on Linux with nested writable roots, and protection of Git metadata on Linux and macOS.

🔗 Codex CLI 0.158.0

Copilot CLI 1.0.89 reaches stable release, Copilot SDK 1.0.15 adds typed outputs

September 28 — GitHub releases Copilot CLI 1.0.89 as a stable version (19:20 UTC). Preview 1.0.89-5, described on September 27, already added support for .claude/rules files; several features among the 35 release note entries are new. Auto routing now suggests a level that can be changed with a shortcut or a click, and drops the unsupported Fast profile: a saved Fast preference falls back to Balance. Pull request creation follows repository templates, including required sections and checklists. In a local session, pressing Escape twice in an empty input restores a prompt whose turn has not started, and directly installed plugins can be enabled and disabled (copilot plugin enable). In the sandbox, wrapped gh and git commands (timeout 60 gh …) work, and localhost is accessible on Windows when local network access is enabled.

Shortly afterward (19:38 UTC), the GitHub Copilot SDK, which embeds Copilot’s agent runtime in an application, reaches version 1.0.15 after five previews. The main addition is typed structured output in all six SDKs (TypeScript, Python, Go, Java, C# and Rust): the developer provides a JSON schema or an idiomatic type, and the model returns validated, typed output instead of free-form text. An experimental handler also lets an application require human review before installing an MCP server or skill, with an explicit decision to confirm, deny or cancel. In Rust, the memory retained by installation of this runtime drops by about 99%.

🔗 Copilot CLI v1.0.89 · Copilot SDK v1.0.15

Devin becomes 30 to 40% cheaper, with Devin Fusion leading FrontierCode 1.1 Extended

September 28 — Cognition announces a substantial reduction in Devin’s usage costs: 30 to 40% less in Fusion and Normal modes, 15 to 20% less in Ultra mode, and up to 70% less for Devin Review, its code review tool. The company attributes the gains to integrating the latest models, including its own SWE-2, and improving Devin’s cloud harness: grouping more actions into one tool call (formatting, analyzing and testing at once), running independent calls in parallel, and making better use of the prompt cache. The figures given for the harness come from illustrative examples, not measurements.

Cognition says it has maintained or improved the intelligence of each mode. Fusion pairs a capable primary model with a less expensive sidekick, and the chart in the post puts it first on FrontierCode 1.1 Extended.

Configuration evaluated by CognitionFrontierCode 1.1 Extended scoreAverage cost per task
Devin Fusion68.8$0.60
Opus 5.5 (high)65.2$0.90
Fable 5.1 (medium)63.6$2.68
GPT-6 Astra (high)63.1$2.62
SWE-2 (high)60.1$0.64
GPT-6 Sol (high)59.6$0.87

The Opus 5.5 (65.2) and GPT-6 Astra (63.1) values, reported at high effort, differ from those in the chart published the same day for Sonnet 5.5 (65.3 and 64.5), which does not specify effort. The more economical Devin is available today, with no new pricing schedule published.

🔗 Devin is now up to 40% more cost-efficient

Devin’s September 25 release notes and Devin Mobile beta

September 25 — Previously overlooked, Devin’s September 25 release notes add 57 hosted MCP integrations to the Marketplace, including Buildkite, Brex, Expo, Vanta, WorkOS and Wix. Devin also produces an interactive HTML report when asked for a report that would benefit from one (charts, tables, diagrams) without a specified format. A child session starts with a summary of its parent session, shown as a removable chip in the input area, and a session whose machine has gone to sleep wakes as soon as someone opens a preview URL or connects to it via SSH. Automations can run on a shared Outpost, and Devin reads Azure Pipelines logs for failed checks on Azure DevOps pull requests.

On September 28, Cognition also opens a waitlist for the Devin Mobile beta. The signup page sums it up in one sentence: Devin runs in the cloud or on your machine, tests in its own browser, and does not stop until the pull request is ready to merge. Neither a general release date nor pricing has been published.

🔗 Devin release notes · Devin Mobile on X

Zed announces Delta 0.18, Gemini CLI publishes an empty nightly

September 28 — Zed announces, without a release date, that version 0.18 of Delta, its multiplayer environment for coding with agents, will run Delta threads in existing Git worktrees; version 0.17, released on September 23, already isolated each parallel task in its own workspace. On Google’s side, the Gemini CLI nightly published on September 28 contains no changes: it uses the same commit as the September 26 nightly, while stable v0.61.0 and preview v0.62.0-preview.0 remain in place.

🔗 Zed on X · Gemini CLI v0.63.0-nightly.20260928


Perplexity’s Agent API gains reusable Profiles, versioned Skills and shared connectors

September 28 — Perplexity adds a reusable, versioned configuration layer to its Agent API for teams. The centerpiece, the Profile, stores everything that defines an agent under a single versioned identifier: model, instructions, tools, Skills, connectors and runtime settings. A project administrator configures the agent once and makes it available to authorized developers; each application then supplies only its own data.

Custom Skills package instructions, procedures and supporting files in a versioned form: a platform team can put its deployment checks, rollback criteria and report format in one, then publish a new version instead of editing every application. Managed connectors, launched in late August, become a resource the administrator shares across the project: applications reference them without each storing their own credentials or tool definitions.

Added resourceRole in the Agent APIAnnounced availability
ProfilesModel, instructions, tools, Skills, connectors and runtime settings under a versioned identifierAvailable
Custom SkillsVersioned instructions, procedures and supporting files invoked by project membersAvailable
Managed connectorsApproved integrations shared by the administrator (GitHub, Slack, Google Drive, Datadog, Linear, Notion)Preview

Everything is configured in the API Console, and members call these resources with an API key from the same project; the post announces no pricing. On the same day, an API changelog entry switches the Agent API’s xhigh preset from GPT-5.6 Sol (openai/gpt-5.6-sol) to Claude Opus 5.5 (anthropic/claude-opus-5-5), while prompts, reasoning effort, tools, token budgets and step limits remain unchanged. Three days earlier, the low, medium and high presets had moved to GPT-6.

🔗 Agent API now supports reusable agents · Perplexity API changelog


SpaceXAI launches Team Bots, Grok Bots shared across an entire team

September 28 — SpaceXAI launches Team Bots, Grok Bots built around a team role or workflow and shared with all its members, who can also work with the Bot individually. The Bot is shared, but conversations remain private: context and memories are separate for each user, while skills are shared across the team. Each Team Bot has its own identity in Slack and can be invited into a channel where the whole team can ask it questions.

Bot componentRole described by SpaceXAI
ContextFiles, instructions and skills
PluginsWork in Salesforce, Notion or GitHub, connected individually or for the team
CredentialsAccess to third-party APIs without a plugin
MemoriesWhat the Bot retains to improve in its role, kept separate for each user

SpaceXAI describes its own uses. Each major sales account has a Team Bot that analyzes customer news, Gong calls, Notion documents and Slack threads overnight, then posts a morning update in Slack. The engineering Bot, connected to Notion, Linear, Hex, Datadog and Cursor, has directed hundreds of Cloud Agents: a five-person team delivered more than 100 pull requests per day and launched Team Bots in a few weeks. Among customers, insurer Harper says it built a Team Bot in 24 hours that helps its customers reinstate expired policies, saving them more than $120,000 according to its CEO.

Team Bots is available today in public beta on the Teams and Enterprise plans, with preconfigured Team Bots for sales, product management, marketing and data analysis. SpaceXAI has announced neither pricing nor quotas.

🔗 Team Bots (SpaceXAI) · @bot announcement on X


H Company releases Holo4, open-weight agent models of 27B and 35B-A3B

September 28 — H Company releases Holo4, its new series of agent models, in two sizes: a dense 27-billion-parameter model and a 35B-A3B Mixture of Experts, both served through the H Models API and published on Hugging Face in BF16, FP8, NVFP4 and 4-bit GGUF. H also adds Holotron4 Nano, created by applying its post-training recipe to NVIDIA’s Nemotron 3 Nano Omni. A single model acts through a graphical interface, code, MCP or APIs, on a computer, on the web, on Android or in a code sandbox; H trained it with supervised learning followed by reinforcement learning, notably on about 10,000 verifiable tasks generated by its Agentic Task Factory from plain documentation.

Evaluated modelBase and licensePublished score
Holo4 27BQwen3.8-27B, CC-BY-NC-4.0 (non-commercial)61.7% on OSWorld 2.0; 85.2% on OSWorld at $0.08 per task
Holo4 35B-A3BQwen3.6-35B-A3B, Apache-2.030.9% on OSWorld 2.0
Holotron4 Nano (Holotron4-30B-A3B)Nemotron 3 Nano Omni, NVIDIA Open Model Agreementgains over the base shown only in a chart
Claude Opus 5.5 (reference cited by H)closed model81.8% on OSWorld 2.0

H specifies its evaluation conditions: Holo4 is measured in its own harness, with a single run on OSWorld 2.0, and compared with public scores obtained using other harnesses and effort levels; on AutomationBench, 480 of the 600 public tasks belong to the split used to collect training data. H publishes all the trajectories behind its public scores, available for step-by-step inspection or download on Hugging Face, and says DSpark drafts to speed up inference are coming in the next few days.

🔗 Holo4 post on Hugging Face · H Company announcement


Robotics: Sakana AI’s SAIL raises success from 25 to 73%, ProjectSim questions simulation measurement

September 28 — Sakana AI presents SAIL (Scaling In-Context Imitation Learning), work carried out with the University of Tokyo and accepted at the IROS 2026 conference; the paper has an arXiv identifier from March 2026, and today’s news is its public presentation. The question it asks: can an existing vision-language model (VLM) produce reliable robot movements without retraining if given more computation at inference time?

SAIL generates a complete trajectory from a few successful demonstrations placed in context and executes it in simulation; a second VLM then estimates task progress from the video, and that feedback guides a Monte Carlo tree search (MCTS). Only the selected trajectory goes to the real robot. Gemini Robotics-ER 1.5 fills both roles, with no changes to its weights.

Evaluated methodSearch nodesAverage success across six simulated tasks
Single generation125%
Depth-first search1537%
Breadth-first search1551%
SAIL655%
SAIL1565%
SAIL4573%

These rates count the configurations (20 per task, in the ALOHA simulator) for which a successful trajectory is found within the budget; success is verified by the simulator, not by the VLM. On a LeRobot SO-101 arm, SAIL succeeds in 5 of 6 attempts to place a block in a bowl with a budget of 15 candidates, and an imitation policy trained on the trajectories found also succeeds in 5 of 6 attempts while running faster. Sakana acknowledges the limitations: open-loop execution, additional computation for the search, and a real-world evaluation limited to one task and six attempts per method.

🔗 Sakana AI post · SAIL project page

ProjectSim surveys robotics benchmarks

September 28 — The gap between simulated and real-world scores is the focus of ProjectSim’s first experiment. In a survey post with no results of its own, Luca Cilio reviews manipulation benchmarks, from MetaWorld and LIBERO to shared physical tests such as RoboChallenge. Citing RoboCasa365 figures, he notes that GR00T N1.5, fine-tuned on 30,000 demonstrations, succeeds on 43% of individual skills but drops to 4.4% on combinations absent from its fine-tuning data. ProjectSim’s experiment aims to determine whether more faithful simulated scenes—with reconstructed geometry, appearance, and physical properties—bring simulation scores closer to those measured on a real robot; no results have been published yet.

🔗 ProjectSim post


Mistral opens a Munich hub for industrial AI and Physics AI

September 28 — Mistral is opening a hub in Munich. It will host research teams working on Physics AI and Industrial AI, alongside applied engineers who work directly with partner companies. Mistral presents itself there as a long-term technology partner rather than a software vendor.

Partner named by MistralAnnounced work
BMWCrash simulation and engineering AI
Siemens EnergyIndustrial AI applications
Technical University of Munich (TUM)Wind tunnel digital twins for automotive aerodynamics

The hub follows Mistral’s acquisition of Emmi AI in May 2026, which brought it more than 30 physicists, researchers, and engineers specializing in computational fluid dynamics (CFD), structural mechanics, and multiphysics simulations. With TUM and Professor Nikolaus A. Adams, Mistral will develop digital twins for automotive aerodynamics by combining real-time wind tunnel sensor measurements with CFD simulations computed offline to produce accurate aerodynamic predictions in real time.

The post reiterates the sovereignty argument: customers can access the model weights and run the models on their own infrastructure under European law, without data leaving their organization. It also says Mistral will build one gigawatt of European computing capacity by 2030. It quotes Germany’s Federal Minister for Digital Transformation, Karsten Wildberger, and the head of the Bavarian State Chancellery, Florian Herrmann. Mistral is hiring in the Munich area but gives no headcount or investment figure.

🔗 Hallo, Deutschland! (Mistral AI)


NVIDIA and Nscale measure DSX MaxLPS: 37% more GPUs and 49% more throughput within the same power budget

September 27 — NVIDIA has published a quantitative evaluation of DSX MaxLPS, its software for distributing power among AI factory resources according to operator policies, conducted with Nscale. According to NVIDIA, it enables deployment of up to 40% more GPUs within the same approved power budget. The post emphasizes that this is coordinated allocation, not an increase in the site’s power supply.

The test ran on GB300 NVL72 systems at Nscale’s data center on the Verne campus in Keflavík, Iceland, which is powered entirely by renewable energy. It used Kimi K2.5 in FP4, NVIDIA Dynamo, and TensorRT LLM. Within the same provisioned budget of 264.4 kW, the fleet grew from 140 to 192 GPUs without reducing the throughput of existing instances.

Measured metricStatic baselineWith DSX MaxLPSRecorded difference
GPUs managed140192+37.1%
Normalized aggregate throughput1,084,503 tokens/s1,618,443 tokens/s+49.2%
Total measured power166.2 kW198.9 kW+19.7%
Power budget utilization62.9%75.2%+12.3 points
Throughput per provisioned watt4.10 tokens/s/W6.12 tokens/s/W+49.2%

The post also reports the trade-off: while median and P75 latencies remain within 5% of the baseline, P99 time to first token rises by 17% from a baseline of 15.7 seconds. NVIDIA recommends a five-step validation process for operators, from defining the power boundary to setting production limits. This evaluation follows Lambda’s validation on HGX B200, presented on September 15 (19 nodes within the budget for 16); projections for Vera Rubin, with liquid coolant entering at 45 °C, are presented separately.

🔗 NVIDIA technical post


Briefs

  • DeepSeek Harness 0.2.0-rc.1 — The first release candidate in the 0.2.0 series and the 23rd prerelease of DeepSeek’s agent harness, still without a stable version: conversations using models on a DeepSeek account can search the web without an additional API key, and automation tasks have moved to an optional plugin package. 🔗 source
  • Pixel Canary on Vercel AI Gateway — Since September 25, Vercel has offered this coding model from an unnamed developer for free during its stealth period: 28 of 31 Next.js tasks at pass@4, tied with GPT-6 Astra (high), and 30 of 31 with documentation supplied through AGENTS.md; zero data retention is unavailable. 🔗 source
  • GitHub Actions self-hosted runners — On GitHub Enterprise Cloud, full enforcement of minimum versions begins September 29: a runner below version 2.329.0 can no longer register, and a runner below the minimum execution version stops accepting jobs. A new REST API provides end-of-support dates. GitHub Enterprise Server is unaffected. 🔗 source
  • NexteraBERT — Rikka Botan has released a bidirectional encoder with 212.6 million parameters, pretrained on 130 billion tokens (about 15 times fewer than ModernBERT): 87.90 on GLUE versus 87.97 for ModernBERT-base, 54.70 versus 53.63 on MTEB v2, and 5.22 times the throughput at 65,536 tokens, according to its own measurements; code is under the MIT license. 🔗 source
  • Real-time LTX-2.5 — A community post takes this 22-billion-parameter audio and video model from 90 seconds per clip to generation faster than playback: 1.83 seconds for a 5-second clip on a 96 GB card, using 4 steps instead of 8, NVFP4 computation, and CUDA Graph; code is under Apache-2.0. 🔗 source
  • SCRIBE — Adalat AI, which develops speech recognition for Indian courts, has published this open-source evaluation tool on PyPI (with a paper at Interspeech 2026). It scores words, numbers, punctuation, and domain terms separately, whereas word error rate assigns 100% to a Malayalam sentence that no proofreader would change. 🔗 source
  • 228 JEV projects — Eric Kang, who maintains the Awesome JEV list, selected 228 open-source projects (10 types, at least 50 stars, each pinned to a commit) from the 13,473 repositories returned by a GitHub search for “jev,” a total that includes unrelated projects. It is a manual selection, not an independent census. 🔗 source
  • HeyGen and Jev — HeyGen has published a guide on X that places Jev, TypeSafe AI’s decision model, ahead of its MCP server: Jev sorts roughly 40 prospects by whether to use video, angle, language, and fit; Claude writes the scripts; then HeyGen MCP renders the batch, the only step that consumes credits and requires approval. 🔗 source
  • Four creations with Gemini 3.8 Flash — Google’s blog showcases four projects built with the model launched September 2: a live satellite tracker built with Antigravity, animated Seigaiha waves, a 3D tyrannosaur skeleton, and an automatic gearbox with 10 camera views; it offers no figures or new product feature. 🔗 source
  • A Brooklyn caterer and Gemini — Google’s blog describes how Edy Massih, owner of Edy’s Grocer in Greenpoint, uses Gemini to scale recipes for 200 guests, generate shopping lists, and adapt menus to dietary needs; no new feature is announced. 🔗 source
  • Lenfest Institute — OpenAI is committing $5 million, plus up to $5 million in software credits and engineering support, to the next phase of the Lenfest Institute’s AI fellows program, launched in 2024 in 11 US newsrooms: twice its previous support. 🔗 source
  • ChatGPT Health tab — Selecting a chart, metric, or record in the Health tab now produces personalized explanations, sometimes with an interactive chart, without writing a prompt. Health remains available only in the US (ages 18 and over, Free, Go, Plus, and Pro plans, web and iOS). 🔗 source
  • WebMCP Challenge winners — OpenAI Developers presents the ten winning projects from the hackathon launched in late August (MASIL, Alza, ArchMorph, Aisle, Roque Nights, Mandate, Observatory, Bouquet Studio, Faraday, JupyterLite WebMCP), without rankings or amounts by project. 🔗 source
  • macOS security fix — On September 25, version 26.924.20706 of the macOS desktop app, listed under the Codex app section of the ChatGPT and Codex changelog, fixed vulnerability CVE-2026-100754 reported by Patrick Wardle (Objective-See Foundation), without describing the flaw. 🔗 source

What it means

The price per token is no longer moving; the number of tokens is falling. Sonnet 5.5 keeps Sonnet 5’s pricing but, according to Anthropic, costs up to 30% less per task because it uses fewer tokens. Devin becomes 30% to 40% cheaper by integrating its latest models, including SWE-2, and batching tool calls. Manus reports a 32% reduction in execution cost with its new harness in a tested configuration. The lever is now the amount of work consumed, by both model and harness, rather than the listed price, and the midrange is catching up with the top tier: on Terminal-Bench 4.0, Sonnet 5.5 exceeds Opus 5.5’s score measured at Xhigh effort.

Agents are gaining autonomy by default, while safeguards are moving beyond their reach. Claude Code now starts in auto mode on every plan, while NVIDIA puts monitoring in a DPU the agent cannot access and, by default, requires human approval for every access rule it proposes. Tools are following the same pattern: Codex CLI submits terminal input for commands requiring elevated permissions for approval, and the Copilot SDK can require human review before installing an MCP server or skill. The more an agent acts on its own, the more control needs to reside outside the agent itself.

The agent is also becoming a team member with its own identity. Perplexity Profiles turn an agent into a versioned configuration reusable across a project; SpaceXAI Team Bots have their own Slack identifiers and keep separate memories for each user; and Cue agents receive an email address, a phone number, and a wallet. An agent that can pay, call, or write on someone’s behalf makes the control questions NVIDIA raised that same day very concrete.

Finally, the day’s figures need careful reading. Devin measures Sonnet 5.5 on a FrontierCode variant that its tweet and chart name differently, and two Cognition charts published the same day give different scores for Opus 5.5. Holo4 is measured in H’s harness in a single run on OSWorld 2.0. SAIL’s 73% is achieved in simulation, and the gap between simulation and reality is precisely the subject of ProjectSim’s first experiment. In media, Eleven v4 draws on an external ranking and blind tests, while Kling 4.0 arrives without a benchmark or pricing, with some features deferred.


Sources