ai-powered-markdown-translatorArticle translated from fr to en with gpt-6-sol.
At Meta Connect 2026, Meta gives Muse, its personal agent, a real-time voice and a video avatar, promises to bring it to its glasses, and makes its Meta Model API generally available. Claude Code cloud sessions leave preview with a one-time credit for subscribers, while Perplexity launches Fast Search, powered by Photon, an in-house search engine written in Rust. In a sign of the times, Meta and Google launch video avatars that speak in real time just hours apart.
Meta Connect 2026: Muse speaks, Meta refreshes its glasses
September 23 — Meta opened Meta Connect 2026 with a keynote at 4 p.m. PT (1 a.m. on September 24 in Paris), focused on Muse, the personal agent launched on September 8. Muse gains a real-time voice mode: users can have long conversations with it while it works in the background, and customize its voice by describing it (faster, slower, with a specific accent). It also gains an animated face, Muse Realtime Avatar, described below. On Mac, computer use is available: with the user’s permission, Muse can operate any application and continue working while the user is away.
| New Muse feature | Announced availability |
|---|---|
| Real-time voice mode, customizable voice | Presented at Connect, no date |
| Computer use in Muse for Mac | Available |
| Muse on AI glasses | “In the coming months” |
| Muse Charm, pocket-sized device | December, according to Mark Zuckerberg |
| Muse’s own email address | Announced, no date |
On the glasses, users will call Muse by name to act on what they are looking at, such as a product on a shelf or a poster. According to Zuckerberg, Muse Charm, equipped with a real-time voice model, will fit on a keychain; the official recap only promises more details by the end of the year.
A day after @Muse announced Shopify, PayPal, Expedia, and Instacart, Meta lists the full set of connectors: Walmart, Best Buy, American Eagle Outfitters, DICK’S Sporting Goods, Fanatics, Gap, Michael Kors, Sephora, Ulta, and Wayfair for shopping; Shop Pay and PayPal for payments; Instacart and Expedia, still “coming soon”; and Notion, Granola, GitHub, and Box for work. GitHub detailed its connector that same evening: review pull requests, track issues and notifications, and leave comments without leaving the agent. It did not specify subscription plans or countries. Meta gives no pricing or countries for these new features.
🔗 Meta’s recap · Mark Zuckerberg’s thread · Muse’s GitHub connector
Glasses: Ray-Ban Meta Audio, Gen 3, Meta VR Glasses, and Ray-Ban Display in France
September 23 — Meta is also refreshing its glasses lineup and promises more than 100 models by the end of the year. A new category, Ray-Ban Meta Audio has no camera: 43 g, up to 12 hours of battery life, plus 48 hours with the case. Third-generation Ray-Ban Meta glasses gain an hour of battery life (9 hours), an action button, and six microphones. Meta Ray-Ban Display, glasses with a screen, are coming to France: following the UK and Canada, preorders are opening there gradually, as in Italy and Germany, ahead of availability on October 13.
| Announced device | Starting price | Announced availability |
|---|---|---|
| Ray-Ban Meta Audio | 349 dollars | Preorder, ships October 13 |
| Ray-Ban Meta (Gen 3) | 449 dollars | Available September 23 |
| Meta Ray-Ban Display | Not specified | France, Italy, Germany: October 13 |
| Meta VR Glasses (about 100 g) | 1,299.99 dollars | Spring 2027 |
Meta Model API generally available, Muse Spark 1.3 at Oracle and Google Cloud, Replit for Meta devices
September 24 — The second day of Meta Connect was aimed at developers. The main announcement: the Meta Model API, which provides access to Muse models, reaches general availability worldwide, with support and compliance guarantees for businesses. Muse Spark 1.3 is now available on Oracle Cloud AI Platform and enters private preview on Google Cloud, two new preferred cloud partners according to Meta.
The Google Cloud listing (Gemini Enterprise Agent Platform) sets out the terms: model meta/muse-spark-1.3 in preview, available on request, a context window of 1 million tokens, text-only input and output, MCP-compatible tool calls, function calling, structured output, and reasoning support, but no batch predictions or provisioned throughput, and only a global endpoint. Pricing had not yet appeared on the pricing page when we checked.
| Developer announcement | Announced status |
|---|---|
| Meta Model API | General availability worldwide |
| Muse Spark 1.3 on Oracle Cloud | Available |
| Muse Spark 1.3 on Google Cloud | Private preview |
| Muse Code | Coming to Windows |
| Wearables Device Access Toolkit 1.0 | Launched after a year in preview |
| WebMCP and Meta AI Connectors | Developer preview |
| Meta Global AI Developer Hackathon | 1 million dollars, dates unannounced |
Muse Code, Meta’s terminal coding agent, which left beta on August 31 on macOS and Linux, is coming to Windows. Ash Jhaveri (“VP AI & Hardware Ecosystems,” according to his X bio) reiterates that Muse Spark 1.2 weights are coming “soon,” a promise repeated since August 10, still without a date. For glasses, WebMCP lets Meta AI call only the functions a web application chooses to expose.
🔗 Meta Connect 2026 developer recap · Muse Spark 1.3 on Google Cloud
Replit builds for Meta headsets and glasses
September 24 — Announced at Meta Connect, Replit opens a path to Meta devices: users describe an application in everyday language, Replit Agent writes the code and shows a preview in the browser, then the application or web experience opens on a Meta Quest headset, Meta Ray-Ban Display, or Meta VR Glasses, with no tool to install. Three paths are described: Expo (building, packaging, and deployment) or the browser for Meta Quest and VR Glasses, and a simulation in Replit followed by web deployment for Ray-Ban Display. The descriptions of Muse differ: Replit’s page presents a Replit connector in Meta’s agent as “coming soon,” with no date, while a Replit post published the same day says Muse can already create applications in Replit. No pricing is mentioned.
🔗 Replit × Meta Connect 2026 · Muse in Replit, according to @Replit
Claude Code: cloud sessions leave preview, Projects goes local
September 23 — Claude Code cloud sessions leave research preview, @ClaudeDevs announced in a thread published at 21:23 UTC. A cloud session runs on Anthropic’s infrastructure: work continues even when the laptop is closed.
Cloud sessions are officially available and out of research preview! They let you keep Claude Code working, even when your laptop is closed. — @ClaudeDevs on X
To let subscribers try them, Anthropic is offering existing subscribers a one-time credit, separate from usage limits:
| Credit condition | Published details |
|---|---|
| Amount on Pro | 100 dollars |
| Amount on Max | 250 dollars |
| Deadline | Claim before October 7 |
| How to claim | Link in the announcement or CLI command /claim-credit |
| Prerequisites | Connected GitHub account, terms apply |
| Scope | One credit per account, cloud sessions only |
| Entry points | claude.ai/code, Code tab in the mobile app, desktop app, claude --cloud |
In response to questions, @ClaudeDevs clarified on September 24 at 01:57 UTC that cloud sessions count against the Pro or Max plan like the rest of Claude Code: the optional credit is simply used first. The documentation confirms this model: rate limits are shared with all Claude usage on the account, parallel tasks consume more of that allowance, and the virtual machine is not billed separately. Cloud sessions are available on Pro, Max, and Team, and to Enterprise users with premium or Chat + Claude Code seats. Cloning and pull requests go through GitHub; a GitLab or Bitbucket repository can be sent as a local bundle with CCR_FORCE_BUNDLE=1.
🔗 Cloud sessions documentation · @ClaudeDevs’ clarification on the plan
Projects runs its threads locally
September 23 — An hour and a half later (22:49 UTC), @ClaudeDevs announced local support for Projects in Claude Code, launched on September 17 in beta for Pro and Max: a project’s threads can now run on the user’s machine. Until now, each was a cloud session on its own branch, and the documentation recommended a local session outside a project for work requiring tools accessible only from the user’s machine (a local database, an emulator, or an API behind a VPN). Anthropic also says it is opening Projects to more Pro and Max subscribers on the waitlist; migration of users from older Claude projects is still underway. The announcement is ahead of the documentation: on the evening of September 24, the Projects page still said local sessions could not be part of a project, and the post does not specify where users launch these threads.
🔗 @ClaudeDevs’ announcement · Projects documentation
Perplexity: Fast Search and Photon
September 24 — Perplexity adds a fast mode, Fast Search, to its Search API, enabled by the search_type: "fast" parameter, and introduces the engine behind it: Photon, a retrieval and ranking service written in Rust. It replaces the open-source engine the company had adapted and now handles all Perplexity searches.
Fast Search runs on Photon, our new Rust-based retrieval and ranking service that we built with a small team of engineers and hundreds of agents. — @perplexity_ai on X
| Comparison criterion | Standard Search | Fast Search |
|---|---|---|
| Search API, 1,000 requests | 5 dollars | 1 dollar |
Agent API web_search tool, 1,000 calls (excluding tokens) | 2.50 dollars | 1 dollar |
| Score across 3,554 tasks from six agentic benchmarks | 64.0 % | 64.3 % |
| Estimated model and search cost for those tasks | 187.60 dollars | 59.73 dollars |
| Call latency, median and 95th percentile (self-reported) | Not published | 160 ms and 230 ms |
The trade-off is explicit: in internal tests of rare queries, relevance falls by 0.24 points and answer availability by about 3 points. Perplexity recommends Fast Search for routine agentic work and standard search for difficult or ambiguous questions. The post notes that its latency comparison with other APIs relies on figures reported by each provider, without a controlled test; for the Search API, the Python and TypeScript SDKs currently require a workaround for their validation. In production, Photon reduced the engine’s internal p99 latency from about 800 to 65 ms, using about 20 % fewer machines.
🔗 Photon: Building a Retrieval and Ranking Engine From Scratch · Fast Search documentation
Real-time video avatars: Meta and Google launch hours apart
Meta and Google launched the same type of product just hours apart: a video avatar that listens, speaks, and reacts in real time, powered by a voice model.
Muse Realtime Avatar
September 23 — Meta Superintelligence Labs details Muse Realtime Avatar, Muse’s face. From a reference image (a portrait, full-body illustration, animal, or everyday object), the model animates a character live, with facial expressions, hand gestures, and body movements. A single stream of speech tokens, produced by the Muse Realtime Voice model, drives both sound and image, keeping the voice and lips synchronized. Distillation reduces computation from 120 to 2 evaluations per segment.
| Metric reported by Meta | Reported value |
|---|---|
| Portrait video | 448x768, 25 frames/s |
| Voice and video response latency | About 870 ms |
| Concurrent sessions per GB200 | 12 |
| Preference over Runway Characters | 78 % versus 22 % |
| Preference over HeyGen LiveAvatar | 88 % versus 12 % |
These preferences come from a human evaluation conducted by Meta, which notes that the difference in mannerisms compared with Runway is not statistically significant. Each video carries the invisible Meta Video Seal watermark.
Gemini 3.8 Live with Live Avatar
September 24 — Nine days after Gemini 3.8 Live, Google gives its voice dialogue model a face: the avatar listens, looks at what the user shows it (through a camera or screen share), and responds with synchronized voice and video. It is generally available in Gemini Enterprise and accessible to developers through the Live API on its agent platform, with endpoints in the United States and the European Union; Gemini 3.8 Live Extended Thinking remains in private preview. The avatar speaks 97 languages and calls tools without interrupting the conversation. A custom avatar requires allowlist registration. According to the model card, output is capped at 24K tokens, and a continuous interaction lasts a few minutes. Avatar video costs $1 per million tokens, at 6,192 tokens per second of speech; listening is not counted. Everything carries a SynthID watermark.
What Claude costs: Opus 5.5 caching and refusals billed again
Two September 24 publications address what Claude users pay: where Opus 5.5 saves money, and the return of billing for certain refusals.
Longer sessions, and why Opus 5.5 costs less for them
September 24 — A claude.com post by Michael Segner aggregates Claude Code usage from March to September 2026. The number of prompts per session has stayed the same, but Claude works 3.3 times longer on each one, with more than 40% more model calls and 68% fewer interruptions. Context per request has grown 2.6-fold, and the input-to-output token ratio has risen from 189:1 to 324:1. Cache reads therefore account for the largest share, making the 60% cut in their price with Opus 5.5 especially valuable; according to Anthropic, a cached token on Opus 5.5 costs one-fifth as much as on competing models. On Opus 5.5 and Fable 5.1, changing effort during a session no longer resets the cache, contrary to what the guide published on September 22 said.
Blocked requests are billed again in three categories
September 24 — Anthropic is resuming billing for requests that its safeguards block before Claude responds, @ClaudeDevs announced. According to the tweet, only categories with low false-positive rates are affected: biology, distillation attacks, and frontier LLM development. The stated reason is “coordinated attacks” against its systems in recent weeks. Anthropic says 99.7% of Claude Code, Claude.ai, or Cowork accounts triggered none of these billable blocks in recent tests, and it is aiming for a false-positive rate below 0.1%; an unjustified block can be reported with /feedback. According to the documentation, a billed refusal is charged at the model’s rate on the Claude API as well as on Bedrock, Google Cloud, and Microsoft Foundry; every refused request counts toward rate limits.
| Refusal category (documentation) | Billed before any output |
|---|---|
bio | Yes |
frontier_llm | Yes |
reasoning_extraction | Yes |
cyber | No |
general_harms | No |
🔗 @ClaudeDevs announcement · How refusals are billed
AI and health: Ebola in the DRC, OpenEvidence, AlphaFold Database, and NV-Reason-CT
Two September 22 announcements covered here bring Claude to public health, while NVIDIA has published two open projects, one in virology with Google DeepMind and EMBL-EBI, the other in radiology.
Ebola in the DRC
September 22 — In “The Situation Report,” a post dated September 22 and shared by @AnthropicAI on the evening of September 23, Anthropic describes the use of Claude against the Bundibugyo Ebola outbreak in eastern Democratic Republic of the Congo. According to the September 19 bulletin, the outbreak has 7,672 confirmed cases and 3,699 deaths (a 48.2% fatality rate), and no vaccine is approved against this strain. A partnership convened by CEPI, with the WHO Regional Office for Africa and the National Institute for Biomedical Research (INRB), is working with two Anthropic teams. WHO Africa wrote a skill that extracts figures for each health zone, compares them with the previous day, and flags changes in trends: the daily situation report now takes less than an hour instead of a full day. CEPI uses Claude to compare vaccine candidates, while scientific decisions remain with experts, and INRB can have Claude Science assemble the virus’s genomes and phylogenetic tree.
🔗 The Situation Report (Anthropic)
Free OpenEvidence for around 100 countries
September 22 — During the United Nations General Assembly, OpenEvidence announced a partnership with Anthropic to offer a specialized version of its platform free to clinicians in around 100 low- and middle-income countries, including Uganda, Angola, Sudan, Haiti, and Mongolia (a list provided by OpenEvidence, according to Reuters). OpenEvidence answers clinical questions using peer-reviewed medical research and care guidelines; the service is already free in the United States and Canada. According to Reuters’ exclusive report, Anthropic’s technology powers the service behind the scenes, and OpenEvidence adapts the system to each region; financial terms were not disclosed. The announcement comes from OpenEvidence and Reuters: Anthropic has not shared it on its official channels.
🔗 OpenEvidence announcement · Reuters exclusive
Protein complexes from more than 2,800 viruses in the AlphaFold Database
September 24 — NVIDIA has joined a coalition of eight organizations, including Google DeepMind, EMBL-EBI, and CEPI, that is publishing predicted 3D structures of protein complexes from more than 2,800 viruses in the AlphaFold Database, from common cold viruses to emerging threats such as mpox. These predictions, labeled by confidence level, were inferred with AlphaFold2, optimized by NVIDIA BioNeMo Inference Runtime; around 30% of the added interactions are absent from the Protein Data Bank, the main database of experimentally determined structures. NVIDIA is also releasing the BioNeMo Structure Prediction Pipeline that produced the data so researchers can apply it to their own targets. The publication coincides with a meeting on pandemic prevention held in New York during the UN General Assembly; the database now contains more than 260 million predictions.
🔗 How Open Science Can Help Researchers Prepare for the Next Pandemic
NV-Reason-CT
September 23 — NVIDIA has published NV-Reason-CT, an open vision-language model for chest and abdominal CT scans. Instead of treating a volume as a stack of 2D slices, a 3D ViT encoder divides it into 13,824 visual tokens passed to the Qwen3.5-4B language model; the model writes structured reports, works through reasoning inspired by a radiologist’s, and answers follow-up questions. On CT-RATE, the leading public benchmark for 3D CT understanding, NVIDIA reports a macro-F1 of 0.614 and a macro-AUROC of 0.871, ahead of the published models it compares against (VoxelFM: 0.581 and 0.870). The weights are on Hugging Face. NVIDIA stresses that this is a research foundation to be refined for a specific use, not an autonomous diagnostic tool or an approved clinical product.
OpenAI calls for international standards for frontier AI
September 23 — Sam Altman, OpenAI’s CEO, addressed the United Nations Security Council, and OpenAI has published the speech as delivered. He described two dangers to avoid: losing control of the future to AI, and a concentration of power that would allow one person, company, or country to impose its worldview using the most powerful models. As systems capable of improving their own successors approach, this moment “calls for extreme caution,” he said; according to Altman, OpenAI has already slowed down unilaterally and will do so again.
It doesn’t matter whether people put the risk of catastrophe at 10%, or 1%, or 12%, or .1%. None of these levels are remotely acceptable. — Sam Altman, address to the United Nations Security Council
He made four requests of governments: a mechanism connecting national and international standards for frontier AI, common standards for comparing evidence and verifying compliance, incident reporting protocols, and secure channels among governments, critical infrastructure operators, and experts. This echoes a post OpenAI published on September 21, also covered here. OpenAI says fully autonomous recursive self-improvement does not yet exist and should not be pursued until it can be done safely. The company calls on the United States to lead global technical standards with the network of AI safety institutes (ten countries, including France), coordinated by the US CAISI (Center for AI Standards and Innovation). For OpenAI, these standards would be neither licenses nor a mandatory review before release; each country would decide how to incorporate them into its laws. Dialogue between the United States and China on these issues would be “a positive step.”
🔗 Building standards for the next phase of AI
AI in cyberdefense: Daybreak for Ukraine and Fuzzing Taskflow
Two announcements put models to work for defenders: one opens a program to a government, while the other assigns fuzzing to an agent.
Daybreak opens to Ukraine
September 23 — OpenAI is opening its Daybreak cyberdefense program to the Ukrainian government, in partnership with the Ministry of Digital Transformation: tools to identify software vulnerabilities, then develop and test fixes faster to protect civilian infrastructure. The announcement was made on the sidelines of the UN General Assembly by Dmytro Kushneruk, Ukraine’s consul general in San Francisco, and Sasha Baker, OpenAI’s head of national security policy. OpenAI notes that CERT-UA, the national incident response team, handled nearly 6,000 cyber incidents in 2025, targeting hospitals, energy, and telecommunications among others. The post cites two European results: ENISA found vulnerabilities in software used by EU institutions, all since fixed, and CERT Polska helped discover six vulnerabilities in third-party router software. Neither a duration nor an amount is specified.
🔗 OpenAI extends cyber access to Ukraine for civilian defense
GitHub Security Lab’s Fuzzing Taskflow
September 24 — GitHub Security Lab has released Fuzzing Taskflow, an autonomous fuzzing pipeline for C/C++ projects, licensed under MIT and built on its Taskflow Agent framework. Give it a GitHub repository, and the agent identifies entry points, writes harnesses, runs AFL++, reads coverage, classifies each crash into one of seven verdicts, and writes a report for each unique bug, with a suggested fix marked “needs review.” The agent makes decisions while MCP tools execute them; each iteration doubles its time budget, from 30 to 960 seconds (around 32 minutes per target), and the loop stops when two iterations each gain less than 1% coverage. The default model is Claude Sonnet 5, which passed all the team’s internal tests. Security Lab cautions that the pipeline runs LLM-selected commands without a container and should be used in a disposable environment. The post gives no figures for bugs found.
Project Suncatcher: TPUs in orbit
September 24 — Announced last year, Project Suncatcher, Google’s project to study placing AI compute in space, is moving to its first test in orbit. A prototype satellite built with Planet is due to launch on SpaceX’s Transporter-18 mission, and the September 24 post says Google’s first TPUs will be in orbit “next week.” The goal is to verify that the chips withstand launch, radiation, and temperature swings. On the ground, Trillium TPUs exposed to a proton beam at UC Davis while running at full AI load withstood a radiation dose greater than that of a five-year mission; heat-pipe and radiator cooling was tested in a thermal vacuum chamber. The next step in 2027 is two satellites to test very high bandwidth laser links. In low Earth orbit, a satellite receives up to eight times more solar energy than it would on the ground, according to Google.
Generative images and videos: Midjourney, Runway, Google Research, and Nunchux
On the creative side, Midjourney is refining its editing tools, Runway is putting numbers to its model router, Google Research is chaining together videos several minutes long, and Nunchux is speeding up MiniMax-H3 on AMD GPUs.
Midjourney: inpainting, —tile, and fast models in alpha
September 24 — Midjourney has released a new version of the --tile parameter, which produces repeatable patterns, for V8.1 and V8.2: according to Midjourney, visible seams between tiles disappear. The editing model’s inpainting and outpainting now change only selected pixels, allowing successive edits without degrading the rest of the image. On alpha.midjourney.com, Midjourney is experimenting with fast models (fast models), presented on X as an early glimpse of real-time models: with the “Live previews” option, the Styles bar previews the latest prompt in each style, and a click starts generation. No names or schedule have been given. The September 23 alpha changelog also adds default parameters (Settings → Advanced → Your defaults) and makes Korean available to everyone on midjourney.com.
🔗 @midjourney announcement · Midjourney updates
Runway puts numbers to its Model Router
September 24 — Runway has published its first quantitative evaluation of Model Router, the Runway Dev router that chooses a generation model for each request based on a preference for cost, quality, or latency. The protocol compares Seedance 2.5, used as the baseline, with three router configurations across 250 image-to-video prompts judged double-blind.
| Configuration tested | Usable videos | Cost per clip | Spending difference |
|---|---|---|---|
| Baseline (Seedance 2.5) | 78 % | $1.80 | — |
| Quality router | 77 % | $1.28 | -29 % |
| Quality capped at $1 | 74 % | $0.61 | -66 % |
| Cost router | 38 % | $0.30 | -83 % |
According to Runway, the capped configuration retains 95% of the baseline quality; it performs better on dialogue and lip sync (96% versus 92%) and worse on physics (52% versus 60%). Runway recommends it for production and reserves cost mode for exploration; 64% of developers already using the router have set it to cost.
🔗 Evaluating Cost vs. Quality Using Runway Model Router
Google Research’s AI video co-director
September 24 — Google Research introduces AI video co-director, a multi-agent orchestration layer built on Gemini and Veo to produce coherent narrative videos several minutes long, where diffusion models struggle to string shots together without costumes and settings drifting. It comprises four research frameworks: a bandit algorithm that chooses the creative strategy, narrative mode, and aesthetic before a multimodal judge scores the result; CANVAS, which maintains a visual memory of characters, locations, and objects; A²RD, which generates one segment at a time and has produced a continuous ten-minute film; and VQQA, which rewrites the prompt based on critiques from a vision-language model. On GenAD-Bench, one of three benchmarks created by the team (400 advertising scenarios), the system achieves a maximum quality score of 81.4. This is research: there is no product or availability date.
🔗 Automating coherent long-form video generation
MiniMax-H3 on AMD MI355X by Nunchux
September 24 — Shared by MiniMax, results published on September 23 by Nunchux AI show MiniMax-H3, MiniMax’s open-weight video model, running on a server with eight AMD MI355X GPUs: a 5-second video generated in 1.33 seconds and a 15-second video in 5.39 seconds, at 1344×768. Compared with SGLang on the same hardware, running in BF16 with 50 Euler steps, the speedup ranges from 21.8 to 26.7 times. The timings include text encoding, denoising, and video and audio decoding, but exclude final MP4 encoding. Nunchux attributes the gain to its model optimizer and inference engine, which uses a custom MXFP6 kernel. The speed makes it possible to change the prompt during playback. Free access is promised “very soon” via a waitlist; the measurements are Nunchux’s own.
🔗 Video generation on AMD MI355X · Post shared by @MiniMax_AI
Coding agents: Claude Code, Delta, Devin, Qwen Code, and Kimi Code
Five coding agents have new releases: Claude Code from Anthropic, Delta from Zed, Devin from Cognition, Qwen Code from Alibaba, and Kimi Code from Moonshot AI.
Claude Code 2.1.282
September 24 — Claude Code 2.1.282, released at 18:38 UTC, has 86 entries, including 59 fixes. The maxProseWidth setting limits the width of Claude’s text in wide terminals, while tables and code blocks retain the full width. Several changes limit what a project can impose: project and local settings ignore OpenTelemetry variables that enable exporting or capture content (CLAUDE_CODE_ENABLE_TELEMETRY, OTEL_LOG_*), as indicated by a startup notice, /status, and claude doctor. The anthropic-skills and claude-ai namespaces are reserved for skills synced from claude.ai, and the managed allowClaudeInChromeWithManagedMcp setting allows claude --chrome to run alongside an exclusive managed-mcp.json. Auto mode now uses the server-side classifier by default on a direct API connection, even when telemetry is off, and fixes prevent extended thinking from being lost when sessions resume.
🔗 Claude Code 2.1.282 release notes
Delta 0.17 from Zed
September 23 — Zed releases version 0.17.0 of Delta, its multiplayer environment for coding with agents, which has been in public beta since September 16: 20 new features, 26 improvements, and 59 fixes. From a thread, users can now ask the agent to launch multiple tasks, each in its own top-level thread with an isolated workspace, so they can run in parallel; agents can also write to one another across threads. The /approve and /request-changes commands deliver a review verdict in any thread. The release also includes the two features announced on September 22: thread colors, inherited by subthreads and subagents, and a progress indicator for automatic compaction in the context panel, which breaks down the share taken by instructions, messages, thinking, and tools.
🔗 Delta release notes · Announcement by @zeddotdev
Devin Review: composable triggers and Perforce
September 23 — Devin’s September 23 release notes, spanning 26 sections, give considerable attention to Devin Review, its pull request review tool. Automatic review now uses composable triggers: users target an account, organization, repository, repository prefix, or member, then filter by author, label, branches, changed files, author type (human, bot, or Devin), or diff size, combining conditions as needed. Each trigger sets when the review runs, and exclusions always take precedence. Devin Review now supports Perforce: a Helix Swarm review can trigger it through a test definition or, starting with Swarm 2026.3, an outgoing webhook. The capability selector now names its models (Ultra: Fable 5.1 and GPT-6 Astra), and Devin’s GitHub app requests six new permissions that administrators must approve.
🔗 Devin release notes for September 23
Qwen Code v0.24.5
September 24 — Qwen Code v0.24.5, a stable release with no announced breaking changes, has 87 entries, including 28 features and 38 fixes. Qwen Live, the agent’s real-time voice conversation feature, now runs through the browser’s Web Shell on all platforms; on macOS, the native host becomes optional (QWEN_SERVE_LIVE_NATIVE_HOST=1). During a call, users can share their screen, transmitted at no more than one frame per second, without screen sharing alone triggering a response. Web Shell gains a timeline of an execution across three lanes (main session, tool calls, subagents), and messaging channels route messages to separate sessions according to their prefix (messageRoutes). For privacy, --bare mode now respects the decision to opt out of usage statistics. Qwen Code Desktop v0.24.5 and the TypeScript SDK v0.1.15 followed on the same day.
🔗 Qwen Code v0.24.5 release notes
Kimi Code 2.1.1 reverses the hardening in 2.1.0
September 24 — Less than 19 hours after 2.1.0, presented on September 23 as a security hardening release, Kimi Code 2.1.1 reverses it. The release note says it removes “some overly defensive changes”; pull request #4013 is more explicit: it reverses the hardening of the workspace trust boundary on the grounds that, once a user approves a repository, its contents are their responsibility. Four protections are removed: file tools again follow symbolic links outside the working directory, local configuration (local.toml) applies without waiting for approval, the repository’s git configuration is no longer restricted during background operations, and the home directory or root are no longer rejected as additional directories. A failed read of trust information still means “unapproved.” File watchers, disabled in 2.1.0, are reenabled by default.
Desktop agents: Perplexity’s Portable Computer and Kimi Work 3.2.14
Two desktop agents that work on the user’s computer changed on September 24: Perplexity opened Portable Computer to PCs with AMD processors, and Moonshot AI released Kimi Work 3.2.14.
Portable Computer on AMD Ryzen AI Max
September 24 — Portable Computer, the fully local version of Perplexity’s Computer agent, now runs on Windows PCs with AMD Ryzen AI Max Series processors, including the Ryzen AI Halo development platform. Until now, the published requirements listed only NVIDIA hardware, from DGX Spark to PCs with RTX GPUs. It requires at least 24 GB of GPU-accessible memory, Windows 10 or 11, and around 20 GB of disk space. Two local models are offered on these machines, PPLX 27B, post-trained by Perplexity, and Qwen 27B, while the product page offers only PPLX 27B on RTX PCs. A local classifier detects names, account numbers, and identifiers before anything leaves the PC, and local inference does not use Computer credits. The offering is available to Pro and Max subscribers, both individuals and businesses.
🔗 Portable Computer comes to AMD-Powered Agentic PCs
Kimi Work 3.2.14: side chats and version history
September 24 — Two days after 3.2.12, Moonshot AI releases Kimi Work 3.2.14, the next version of its desktop agent, without an accompanying post. The main addition is Side Chat: from the “+” panel, users can open an independent conversation within a session that inherits the context of the original conversation. They can also quote a selected passage from the conversation. Office files gain version history, and the browser lets users annotate web pages and then cite those annotations in the input box. Remote Control now allows users to preview files and send attachments from their phone, and Keep Awake is now activated only on request.
Open models and labs: Tev1, LFM2.5-VL-DSpark, LeRobot, and Sakana AI
In the open ecosystem: a small decision model, a draft model that speeds up a vision-language model, a new data format for robotics, and Jürgen Schmidhuber’s arrival at Sakana AI.
Tev1-4B-experimental from Together AI
September 23 — Together AI releases Tev1-4B-experimental, a small decision model: give it a situation, a question, and a list of 2 to 24 options, and it responds with the letter of its chosen option. It is a fine-tune of Qwen3.5-4B inspired by Jev, the decision model often cited in Hugging Face community posts; the model card specifies that it is not a Jev engine, as the model retains Qwen’s standard language head. Together highlights the cost: 0.042 per million input tokens and $0 for output tokens. On its main development set, it makes 880 correct decisions out of 1,000, a result the model card distinguishes from an independent benchmark. The license for the fine-tuned weights is “being finalized.”
🔗 Together AI announcement · Model card
LFM2.5-VL-DSpark from Liquid AI
September 24 — Liquid AI extends its DSpark speculative decoding technique to the LFM2.5-VL-3B vision-language model, after introducing draft models for text on August 20. A 279.5-million-parameter draft model (+8.9%) proposes several tokens ahead for the target model to verify, without changing the final answer under greedy decoding. On a Mac M5 Max with MLX, decoding is 2.30 to 3.13 times faster, and end-to-end latency improves by a factor of 1.56 to 2.62; on H100, decoding is up to 2.66 times faster. The actual gain remains limited by image encoding, which the technique does not accelerate. The draft model is supported from day one by llama.cpp, MLX-VLM, and SGLang.
🔗 Accelerating vision-language models with LFM2.5-VL-DSpark
LeRobot and LanceDB
September 24 — LeRobot, Hugging Face’s robotics learning library, now reads datasets in Lance format natively, using the same training APIs. This makes it possible to train directly from the Hub or object storage without downloading hundreds of gigabytes, while shuffling data across the entire dataset. On DROID (27.6 million frames, 369 GB of video), 10,000 steps of SmolVLA on 8 H100s took 1 h 27 from remote storage, versus 2 h 00 with the standard reader on a local NVMe copy, which first required a 384 GB download. The final loss is identical: the difference comes from time spent waiting for data, 1.7% versus 37.4%. On a small dataset, the two approaches perform equally well.
🔗 How to Train Your Robot: The LanceDB Edition
Jürgen Schmidhuber joins Sakana AI
September 24 — Sakana AI announces that Jürgen Schmidhuber has joined as Chief Scientific Advisor, a role he will hold alongside his current positions. Known for his work on world models, meta-learning, and the Gödel machine, he will be involved in the Tokyo startup’s RSI Lab, which focuses on recursive self-improvement, and will travel regularly to Tokyo. According to Sakana, his work, including his 1987 thesis on recursive self-improvement, directly inspired the Darwin Gödel Machine and The AI Scientist. The announcement also reveals a direction for the company: Sakana is working on its first “agent-native” world models, capable of simulating the physical consequences of an action before carrying it out, with Japan’s manufacturing industry in its sights.
🔗 Jürgen Schmidhuber joins Sakana AI
Briefs
- Claude Tag and personal connectors — In a Slack channel, Claude Tag can now use the personal connectors of the person who invokes it (calendar, drive, CRM), and only that person’s, with review of each response or an auto mode; they are never used for scheduled routines. Rollout is underway for Team, with Enterprise to follow. 🔗 source
- Sora 2 leaves the API — As announced on March 24, the Sora 2 models and the Videos API were shut down on September 24, with no equivalent replacement API; five identifiers have been removed, from
sora-2andsora-2-proto their dated snapshots. Only the documentation records the shutdown. 🔗 source - ChatGPT Ads in Southeast Asia — ChatGPT’s ad platform is coming to Indonesia, Malaysia, the Philippines, Singapore, Thailand, Vietnam, and Taiwan, and now covers more than 60 countries, up from more than 40 at the end of August; only Free and Go plans display ads. 🔗 source
- Grab and OpenAI — The GO Forward with AI program aims to train 30,000 Grab partner drivers, delivery workers, and merchants over two years, starting in Singapore, through workshops built on OpenAI Academy and ChatGPT Work. 🔗 source
- Three case studies at OpenAI — Ringg handles more than 7 million completed calls per month and has cut model costs by about 90% for workloads moved from GPT-4.1 to GPT-5.6 Luna; invideo says GPT-6 Astra has roughly tripled its success rate in color grading; Harvey offers a testimonial without figures. 🔗 source
- ChatGPT for iOS 1.2026.258 — The September 23 entry in the ChatGPT & Codex changelog redesigns the home screen, adds an iPad split view that keeps the task list beside the open task, and fixes SSH pairing failures on Mac and Linux. 🔗 source
- Gemini CLI v0.61.0 — Gemini 3.8 Flash and Gemini 3.5 Flash Lite, already available in nightly builds, move to stable through a cherry-pick that the official release note does not mention; v0.62.0-preview.0 opens the next preview channel. 🔗 source
- Gemini in Chrome — On desktop, Gemini in Chrome generates interactive quizzes from the open tab (initially in English in the United States and India) and can now analyze almost any podcast or video file played in full, beyond YouTube. 🔗 source
- Google Photos — The virtual wardrobe, which catalogs clothes in photos so users can try on outfits, is opening to eligible users in the United States, India, and Brazil, alongside retro Moods styles and 15 new Remix templates. 🔗 source
- API Gateway as an MCP server — In public preview, Google Cloud API Gateway turns REST API operations into MCP tools from an annotated OpenAPI 3.x specification, reusing authentication, quotas, and logs, with up to 1,000 tools per gateway. 🔗 source
- PostgreSQL for agents in AlloyDB — In preview, AlloyDB creates isolated, read-only instances within seconds, with data no more than a second behind, to absorb bursts of agent queries; Google reports more than 3 million queries per second. 🔗 source
- OLMo 3 7B reproduced on TPU — Google Cloud’s TPU team reproduced the pretraining of Ai2’s OLMo 3 7B in MaxText (about 5.9 trillion tokens), producing a model indistinguishable from the original on eight benchmarks and reaching 44.5% FLOP utilization on Ironwood after uncovering a data loader bug. 🔗 source
- The Talking Museum — Google Arts & Culture Lab is opening a virtual study room with more than 190 cultural objects in 3D, which visitors can rotate while Gemini explains their history and answers questions. 🔗 source
- DeepSeek Harness 0.1.7-rc.2 — The second release candidate in two days and the 22nd prerelease without a stable version: tasks keep running in the background after the Desktop window closes, and scheduled tasks survive a restart and can repeat as often as once a minute. 🔗 source
- A request router on Jetson — On the Hugging Face blog, Eric Mey replaces a 9B model with a fine-tuned ModernBERT-large to choose the response type: 180 correct routes out of 180 versus 175, in 42 ms versus 771 ms, on different hardware; the model, data, and code are under Apache-2.0. 🔗 source
- Enriched Carbon genomic corpus — Parag Ekbote has published a CPU-based pipeline, without the Carbon model, that adds length, GC content, coding status, and taxonomy to 32.4 million records in HuggingFaceBio’s pretraining corpus. 🔗 source
- ZCode v3.14.3 — The only GitHub release of ZCode, Z.ai’s harness, since its code was opened on September 21: concurrency can be changed during a running workflow without stopping it, and workflow scripts use fewer tokens. 🔗 source
- xAI Python SDK 1.20.0 — For
grok-imagine-video-1.5,last_frame_urlsets the clip’s final frame andkeyframesplaces up to four keyframes within it; these options were already described in the API documentation. 🔗 source - Copilot code review — The personal code review settings page is opening to all Copilot plans, including Business and Enterprise, with automatic review extended to new pushes and drafts; enterprises can set a default review effort. As announced in late August, GitHub’s default will change from Lite to Balanced on September 28. 🔗 source
- Proof of presence on GitHub Enterprise Cloud — In public preview, high-impact actions (token creation, webhooks, security settings) can require reauthentication or MFA with the identity provider, valid for two hours; available to EMU and GHEC-DR enterprises using Microsoft Entra ID. 🔗 source
- End of Node 20 in GitHub Actions — Runners now execute JavaScript actions with Node 24, the fallback option has been removed, and self-hosted runners on macOS 13.4 or earlier and ARM32 are no longer supported. 🔗 source
- Copilot app canvases — Burke Holland makes the case for these small full-stack applications that interact with the agent, instead of an all-chat interface, citing token savings; there is no new feature, as canvases have existed since June. 🔗 source
- Amp shares its runners — A runner launched with
--sharecan be made available to the whole workspace from ampcode.com; Amp warns that other members can then execute code under the host’s identity, and administrators can disable sharing. 🔗 source - v0 model selector — Catch-up from September 22: the new selector lists AI Gateway models with their per-token prices and is coming to Enterprise teams, where the owner must enable Allow custom models. 🔗 source
- BioNeMo MoE recipe — NVIDIA details training biological mixture-of-experts models with BioNeMo and Transformer Engine: up to 2.21 times the throughput of the Hugging Face baseline on Mixtral-8x7B and eight B200s, with the fused kernel limited to Blackwell. 🔗 source
- Seedance 2.5 Draft mode in Runway — Faster drafts that cost fewer credits for exploration before producing the best ones at full quality; Runway does not specify the credit costs, and the same mode arrived on Pika on the 22nd. 🔗 source
- Runway MCP in the Cursor Marketplace — Runway’s MCP server has been installable from the Cursor Marketplace since September 23, and agents have been able to view Brand Kits in read-only mode since the 22nd. 🔗 source
- Pika API Club — Pika puts numbers on its members’ savings compared with Fal.ai: 2,000 to $990 at AI-Bridge. 🔗 source
- HeyGen study on avatars — Among 1,000 small businesses, 71.6% have recorded a video they never published; among avatar users, 71.1% say they save time and 55.3% publish more often. 🔗 source
- Luma survey — Catch-up from September 22: of 760 creative professionals, 81% have published AI-made content, but no task has an adoption rate above 42%, and 62% worry about becoming dependent on AI at the expense of human creativity. 🔗 source
- Perplexity’s market research guide — An eight-step guide with example prompts, verification steps, and a five-point checklist; it introduces no new feature and points readers to Deep Research and Spaces. 🔗 source
What it means
The agent is moving beyond the chat window. Muse speaks in a voice users can describe, appears as an avatar, controls a Mac, and is due to arrive on glasses and then, according to Mark Zuckerberg, in a pocket-sized device in December. Google is also giving Gemini 3.8 Live an animated face for reception and customer service, Perplexity is putting its local agent on AMD-powered PCs, and Claude Code keeps working in the cloud after the laptop is closed. The timelines vary: computer use on Mac and Live Avatar are available, while Muse on glasses is promised “in the coming months” and its email address has no date.
The cost of agentic work is becoming a quantified selling point. Anthropic is offering 250 in credit for cloud sessions and shows, using usage data, that cache reads now account for the largest share of Claude Code usage, where the input-to-output ratio has risen from 189:1 to 324:1. Perplexity charges 5 and puts savings on agentic tasks at about 68% for a comparable score; Runway retains 74% usable videos while spending 66% less; Together AI trains a decision model for $17. Each measures savings using its own method: useful figures, but seldom directly comparable.
Trust is managed case by case, sometimes in opposing directions. Anthropic is once again charging for some refusals to curb “coordinated attacks,” while targeting fewer than 0.1% false positives. Claude Code 2.1.282 prevents a project from imposing its telemetry, while Kimi Code 2.1.1 removed the protections introduced in 2.1.0 less than 19 hours earlier, arguing that an approved repository is the user’s responsibility. GitHub is testing proof of presence before sensitive actions. On the defensive side, OpenAI is opening Daybreak to Ukraine, and GitHub Security Lab is assigning fuzzing to an agent while warning that it makes mistakes and must run in a disposable environment.
Finally, rules and infrastructure are being discussed at scale. At the UN, Sam Altman is calling for international standards for frontier AI, which OpenAI wants led by the United States through the network of AI safety institutes, without mandatory licensing or review before publication. Google is preparing to send its first TPUs into orbit to test space-based computing, and Perplexity has rewritten in Rust the engine that serves all its searches, reducing p99 latency from about 800 to 65 ms. Standards negotiated among states, chips in orbit, a rewritten engine: the layer that runs these agents is being built alongside them.
Sources
- Everything Meta Announced at Connect 2026 (Meta)
- @finkd: Mark Zuckerberg’s Connect 2026 thread
- @github: Muse’s GitHub connector
- Ray-Ban Meta Audio and the AI glasses lineup (Meta Newsroom)
- What’s New with Meta Ray-Ban Display (Meta Newsroom)
- Meta Connect 2026: Developer Recap (Meta)
- Muse Spark 1.3 on Gemini Enterprise Agent Platform (Google Cloud)
- Replit × Meta Connect 2026 (Replit)
- @Replit: Muse in Replit
- @ClaudeDevs: cloud sessions leave preview
- @ClaudeDevs: cloud sessions and plans
- Use Claude Code in the cloud (documentation)
- @ClaudeDevs: local Projects
- Projects in Claude Code (documentation)
- Photon: Building a Retrieval and Ranking Engine From Scratch (Perplexity)
- @perplexity_ai: Fast Search
- Fast Search (Perplexity documentation)
- Portable Computer comes to AMD-Powered Agentic PCs (Perplexity)
- Bringing Your Muse to Life (Meta AI Research)
- Introducing Gemini 3.8 Live with Live Avatar (Google)
- Claude Code 2.1.282 (GitHub)
- Coding sessions are longer and use more context (claude.com)
- @ClaudeDevs: billing for blocked requests
- How refusals are billed (Claude documentation)
- The Situation Report (Anthropic)
- @OpenEvidence: the partnership with Anthropic
- Anthropic and OpenEvidence (Reuters)
- Sam Altman at the United Nations Security Council (OpenAI)
- Building standards for the next phase of AI (OpenAI)
- OpenAI extends cyber access to Ukraine for civilian defense (OpenAI)
- AI-powered fuzzing with the GitHub Security Lab Taskflow Agent (GitHub Blog)
- Behind Project Suncatcher (Google)
- @midjourney: inpainting, —tile, and fast models
- Updates (Midjourney)
- Evaluating Cost vs. Quality Using Runway Model Router (Runway)
- Automating coherent long-form video generation (Google Research)
- Video generation on AMD MI355X (Nunchux)
- @MiniMax_AI: the MiniMax recap
- How Open Science Can Help Researchers Prepare for the Next Pandemic (NVIDIA)
- Introducing NV-Reason-CT (NVIDIA Developer)
- Delta release notes (Zed)
- @zeddotdev: Delta 0.17
- Devin release notes for September 23 (Cognition)
- Qwen Code v0.24.5 (GitHub)
- Kimi Code 2.1.1 (GitHub)
- Kimi Code PR #4013 (GitHub)
- Kimi Work release notes (Moonshot AI)
- @togethercompute: Tev1-4B-experimental
- Tev1-4B-experimental (Hugging Face)
- Accelerating vision-language models with LFM2.5-VL-DSpark (Liquid AI)
- How to Train Your Robot: The LanceDB Edition (Hugging Face)
- Jürgen Schmidhuber joins Sakana AI (Sakana AI)
- Claude Tag and personal connectors (claude.com)
- Sora 2 model card (OpenAI)
- ChatGPT Ads expands to Southeast Asia and Taiwan (OpenAI)
- Grab and OpenAI bring practical AI skills to Southeast Asia (OpenAI)
- Ringg and OpenAI (OpenAI)
- ChatGPT & Codex changelog, September 23 (OpenAI)
- Gemini CLI v0.61.0 (GitHub)
- 5 ways to upgrade your study habits with Chrome (Google)
- 5 Google Photos updates (Google)
- Turn your REST APIs into MCP tools with Google Cloud API Gateway (Google)
- Announcing PostgreSQL for agents in AlloyDB (Google Cloud)
- Reproducing OLMo 3 7B Pre-training in MaxText (Google)
- The Talking Museum (Google Arts & Culture)
- DeepSeek Harness v0.1.7-rc.2 (GitHub)
- Train Your Own Request Router (Hugging Face)
- From Genomic Sequences to Taxonomic IDs (Hugging Face)
- ZCode v3.14.3 (GitHub)
- xAI Python SDK v1.20.0 (GitHub)
- Copilot code review, September 23 (GitHub Changelog)
- Require proof of presence for high-impact actions (GitHub Changelog)
- Node 20 is no longer available in GitHub Actions (GitHub Changelog)
- When chat is the wrong UI (GitHub Blog)
- Shared Runners (Amp)
- v0 changelog for September 22 (v0)
- Efficient MoE training for biological foundation models (NVIDIA Developer)
- @runwayml: Seedance 2.5 Draft mode
- Runway changelog (Runway)
- Pika API Club Member Spotlights (Pika)
- The State of AI Avatars 2026 (HeyGen)
- How Creative Teams Are Integrating AI in 2026 (Luma)
- AI for market research: a step-by-step guide (Perplexity)