ai-powered-markdown-translatorArticle translated from fr to en with gpt-6.1-sol.
Anthropic is adding two features to Claude, Claude Dashboards for dashboards kept up to date and Claude Motion for animations written in code, and bringing Docs, Slides and Design out of beta on all plans, including Free. On the same day, the company is launching its Cyber Mission, with a program to defend critical infrastructure and OSS Scanner, which offers open source projects free scans by its most powerful models. OpenAI, meanwhile, is rolling out the Ultrafast tier for GPT-6.1 Sol, up to 8 times faster than Standard mode, while Anthropic and NVIDIA are committing 150 million and 1 billion dollars, respectively, to the Genesis Mission.
Claude Dashboards and Claude Motion in beta, Docs, Slides and Design on all plans
October 8 — Anthropic is adding two features to Claude. With Claude Dashboards, in beta on paid plans, users connect a company data platform (Amazon Redshift, BigQuery, ClickHouse, Databricks or Snowflake) or another connector, such as Salesforce opportunities, then ask a question in everyday language: Claude writes the query, builds the dashboard and keeps it up to date as the data changes. Clicking a figure displays the query that produces it, and each chart indicates when its data was refreshed. Anthropic intends it for quick, exploratory questions; for deeper analysis, the dashboard can be sent to Amplitude, Grafana, Hex, Mixpanel, Omni, Perplexity, PostHog or Sigma, with Looker, monday.com and Tableau to follow.
Claude Motion, in beta on Team and Enterprise, is aimed at short animations: a 30-second explainer video drawn from a quarterly report, an animated chart in a presentation or a product demonstration. Since no video generation model is involved, the animation contains no AI-generated footage or people.
Claude Motion turns reports, charts, or product walkthroughs into short animations. Claude writes each one as code, with no video model involved, so you can change any word, number, or timing, then export an MP4. In beta on Team and Enterprise plans. — @claudeai on X
On the same day, Claude Docs, Slides and Design are losing their beta label and becoming available on all plans, including Free. According to Anthropic, more than 45 million documents, presentations and designs have been created since they became available in conversations on September 16. The release from beta brings CMEK support for artifacts, administrator control over artifact templates, collaborative editing with Claude, sharing outside the organization if the administrator allows it, PowerPoint and PDF exports that match the editor, direct export to Google Slides and editing from the mobile app.
| Feature | Status as of October 8 | Applicable plans |
|---|---|---|
| Claude Dashboards | Beta | Paid plans |
| Claude Motion | Beta | Team and Enterprise |
| Claude Docs, Slides and Design | Out of beta | All, including Free |
| claude.ai/design (standalone site) | Open until December 14 | — |
The practical consequence: the standalone claude.ai/design site, now integrated into Claude, remains open until December 14. An organization’s design systems (design systems) can be migrated all at once from the Artifacts page, but conversations and comments on the standalone site, along with its public links, will disappear when it closes. On Enterprise, Dashboards and Motion are disabled by default, while Docs, Slides and Design will be enabled by default on October 15.
🔗 Claude post: Dashboards and Motion · @claudeai thread
Runway and Luma, launch partners for Claude Motion
October 8 — Two generative video companies are presenting themselves as launch partners for Claude Motion. In Runway, any animation can be exported to add generated footage, edit shots by describing the desired change and have Runway Agent assemble a longer edit around the animation; the same 16:9 animation can also be adapted into vertical and square formats. In Luma, animations open directly through the Luma connector (MCP) added to Claude: the Ray and Uni video models restyle them while preserving motion, reframe them to 9:16, 1:1, 4:3 or 21:9, and the Luma agent produces additional versions. Neither Runway nor Luma specifies pricing or credit costs. Anthropic also lists Adobe, Descript, HeyGen, Higgsfield and invideo among the tools for taking an animation further, with Canva and Captions announced as coming soon.
🔗 Runway post · Luma post
Anthropic Cyber Mission: a program for critical infrastructure and OSS Scanner for open source
October 8 — Two days after expanding its Cyber Verification Program, Anthropic is launching the Anthropic Cyber Mission, presented as a long-term commitment to securing the systems everyone depends on. The starting point comes from Project Glasswing: finding vulnerabilities has never been easier, but verifying, prioritizing and fixing them remains difficult. The mission starts in two areas, critical infrastructure and open source software; other areas have been announced, with no date.
For critical infrastructure, the Critical Infrastructure Defense Program (CIDP) provides frontier Claude models, on-site engineers and Anthropic’s threat research to service providers working with operators of operational technology: programmable controllers, control software and industrial networks in electricity, water or transportation, which often cannot be shut down to apply a patch. Eleven founding partners are participating: Accenture, Booz Allen, CrowdStrike, Deloitte, Dragos, Hitachi, Insane Cyber, Nozomi Networks, Palo Alto Networks, PwC and Rockwell Automation. The program starts with a small cohort and is set to expand to other partners and sectors in the coming months.
For open source, OSS Scanner, inspired by Google’s OSS-Fuzz, offers registered projects free periodic scans by Anthropic’s most powerful models, including Claude Mythos. Each report includes a reproducer, an explanation and, when available, a candidate patch, but is sent without human review: Anthropic expects a true positive rate above 90 % and warns that some reports will contain errors, such as incorrectly assessed severity. The service addresses the bottleneck described by the Frontier Red Team: the vulnerabilities found far exceed what humans can triage.
| Metric published by Anthropic | Published value |
|---|---|
| Candidate vulnerabilities found in six months | More than 29 000 |
| Vulnerabilities reviewed and triaged manually | Around 6 000 |
| Verified critical or high-severity vulnerabilities (48 projects) | 97 |
| Meeting the required standard for coordinated vulnerability disclosure (CVD) | 85, or 88 % |
| Real but duplicate, or invalid | 11 and 1 |
| Share of vulnerabilities found by LLMs on the CyberGym benchmark | Less than 20 % at the start of last year, more than 85 % this year |
wolfSSL reports that, of 74 reports received, all but two were valid and five became CVEs. Registration requires a pull request to the anthropics/oss-scanner GitHub repository and is restricted to lead maintainers of projects deemed critical, assessed case by case, while the Defender Advantage Fund launched in August keeps the service free. Anthropic’s forecast is cautious: within two years, AI will favor defense, but not necessarily in the short term, since exploitation costs less while verification, disclosure and remediation remain slow.
🔗 Anthropic post: Anthropic Cyber Mission · Frontier Red Team post on OSS Scanner · anthropics/oss-scanner repository
Ultrafast for GPT-6.1 Sol: up to 8 times faster, at 12 and 60 dollars per million tokens
October 8 — Announced as “coming soon” on September 29, the Ultrafast tier is arriving for GPT-6.1 Sol. OpenAI Developers says it is rolling out that same day in the API, Codex and ChatGPT Work, with intelligence described as close to GPT-6 Astra. OpenAI intends it for work where every second counts: debugging an outage, having agents navigate applications or providing live experiences.
Ultrafast is rolling out today for GPT-6.1 Sol in the API, Codex, and ChatGPT Work. Near-Astra intelligence at up to 8x faster speeds than Sol Standard, so you can build as fast as the ideas come. — @OpenAIDevs on X
In the API, users simply call gpt-6.1-sol with service_tier: "ultrafast" in the Responses API. The mode is available to all API users, subject to rate limits, with global processing and data residency in the United States and the European Union, while GPT-6 Astra in Ultrafast remains limited to US residency. Pricing follows the rule already applied to Astra: six times the Standard price, or 12 dollars per million input tokens and 60 dollars for output, five times cheaper than GPT-6 Astra in Ultrafast (60 and 300 dollars).
| Processing mode for gpt-6.1-sol | Input price | Cached input | Cache write | Output price |
|---|---|---|---|---|
| Standard, short context | 2,00 dollars | 0,10 dollar | 2,50 dollars | 10,00 dollars |
| Fast, short context | 4,00 dollars | 0,20 dollar | 5,00 dollars | 20,00 dollars |
| Ultrafast, short context | 12,00 dollars | 0,60 dollar | 15,00 dollars | 60,00 dollars |
| Ultrafast, long context | 24,00 dollars | 1,20 dollar | 30,00 dollars | 90,00 dollars |
Prices per million tokens, API Pricing page as of October 8.
In Codex and ChatGPT Work, access is restricted to the Pro 500 plan, eligible Enterprise workspaces billed by usage and credit-based Edu plans. On Enterprise, Ultrafast is disabled by default and workspace owners must enable it for selected users or everyone. The Codex documentation details the cost: the allowances included in the subscription are consumed 8 times faster than in Standard mode, and purchased credits and on-demand Enterprise usage are billed at six times the price. Other self-service plans do not have access at launch, even with purchased credits.
🔗 OpenAI API changelog · OpenAI API pricing · Ultrafast in Codex (documentation)
Science: 150 million dollars from Anthropic and 1 billion from NVIDIA for the Genesis Mission, an ultraviolet map of the sky
Two commitments to the Genesis Mission, the US federal program seeking to accelerate scientific discovery through AI, were announced on October 8 at the same summit, “Science: A New Golden Age,” organized in Washington by the White House Office of Science and Technology Policy (OSTP). On the same day, Anthropic is showing what Claude Science can produce for an astrophysicist.
Anthropic: 150 million dollars over three years and Claude for more than 15 federal agencies
October 8 — Anthropic is committing 150 million dollars over three years to the Genesis Mission. The funding is intended to make Claude available to more than 15 agencies participating in the mission, including NASA, the National Institutes of Health and the National Science Foundation. Specifically, Anthropic will provide Claude, Claude Code and API credits to several hundred research projects, work with agencies and national laboratories on the administration’s priorities, including fusion energy and quantum computing, and provide training and technical support. The commitment extends the partnership established in December 2025 with the Department of Energy; in July, Google DeepMind (40 million dollars) and OpenAI (17 million dollars) announced their own commitments to the mission.
🔗 Anthropic post: commitment to the Genesis Mission · @AnthropicAI tweet
NVIDIA: 1 billion dollars over five years for American science
October 8 — At the same event, NVIDIA is announcing commitments worth 1 billion dollars over five years to develop US capacity for superintelligence research and development in areas such as quantum computing, healthcare and energy security. The press release lists three areas without breaking down the amount: support for academic research institutions, investments to accelerate US leadership in quantum computing and support for cloud providers serving government missions. NVIDIA also says it is collaborating on several projects funded during phase 2 of the Genesis Mission, announced that same day, in quantum computing, fusion, accelerator design and microelectronics, and highlights its partnership with the DOE, which includes the department’s largest scientific supercomputer at Argonne National Laboratory.
The first complete ultraviolet map of the sky, produced with Claude Science
October 8 — The Anthropic Science blog explains how Brice Ménard, an astrophysicist at Johns Hopkins University and a researcher at Anthropic, used Claude Science to produce what Anthropic presents as the first complete ultraviolet map of the sky. The space-based GALEX survey (NASA, 2003-2013) covered only about two-thirds of it, in some 38 000 observations. Claude orchestrated agents that gathered and calibrated public surveys against one another, then filled in the missing third through reconstruction (inpainting), based on the relationship between ultraviolet and other wavelengths. In deliberately masked areas, the estimates fall within about 10 % of the actual measurements.
The work would have taken humans weeks of work, but, with Claude, just took a few days while Ménard worked on other projects. — @AnthropicAI on X
Prohibited uses and abuse: Anthropic’s 2026 usage policy and two influence operations disrupted by OpenAI
Two publications from 8 October show what the labs prohibit and how they track abuse: Anthropic rewrites its usage policy, while OpenAI details two influence operations it has banned.
Anthropic publishes the 2026 version of its usage policy, effective 12 November
8 October — Anthropic publishes the annual update to its usage policy (Usage Policy), which will take effect on 12 November. It consists mainly of clarifications informed by observed abuse. A new section brings together previously scattered rules on deceptive campaigns and artificial activity, targeting all deceptive activity, political or commercial, including building tools for influence operations. The election section, renamed “Do Not Undermine Democratic Processes” (Do Not Undermine Democratic Processes), drops the blanket ban on personalized voter targeting, which had blocked civic uses such as informing voters in other languages. The text also clarifies that the weapons ban covers their software and components as well as arming drones, prohibits tracking people without their consent, and bans sustained, gratuitous abusive behavior toward models in extreme cases only.
🔗 Anthropic post: 2026 usage policy update
OpenAI disrupts Dark Clark and Bogus Bylines, including its first category 5 operation
8 October — OpenAI publishes a report on two covert influence operations it has banned, which maintained fake entities (false front) with help from its models. “Dark Clark,” originating in Russia, targeted Latin America: its operators mainly wrote their internal reports with ChatGPT and ran a purported research center, the Social Research Center, led by a fictional character, “Mia Clark.” “Bogus Bylines,” originating in Iran, placed nearly 100 articles under seven fabricated Western journalist bylines in a dozen online media outlets, one of which had nearly 2 million Facebook followers. OpenAI says it has exposed 30 such operations in two and a half years.
| Disrupted operation | Account origin | Main target | Category on the IO Breakout Scale (1 to 6) |
|---|---|---|---|
| Dark Clark | Russia | Latin America | 5, a first according to OpenAI |
| Bogus Bylines, articles | Iran | International online media | 4 |
| Bogus Bylines, comments | Iran | Social media | 2 |
🔗 OpenAI report on false-front operations
Claude Haiku 5.5 in third-party tools: GitHub Copilot offers it starting with Pro, Devin measures it at 58,4 %
Launched on 7 October, Claude Haiku 5.5 arrived the same day in two coding tools, each of which compares it with Claude Sonnet 5.
GitHub Copilot: Haiku 5.5 on all paid plans, including Pro
7 October — Just over two hours after its launch by Anthropic, Claude Haiku 5.5 becomes generally available in GitHub Copilot for the Pro, Pro+, Max, Business and Enterprise plans: including Pro, like Claude Sonnet 5.5 and unlike GPT-6.1 Sol in late September. It can be selected across ten interfaces, from VS Code to Xcode. GitHub positions it for sub-agents, quick edits and terminal tasks; according to its initial tests, it matched Claude Sonnet 5 on many coding tasks with substantially fewer tokens and steps, though no figures were published. It is billed at the public rate: 0,10 dollar per million input tokens and 0,50 dollar for output up to 100 000 input tokens, then 0,50 and 2,50 dollars, or one-tenth the price of Claude Haiku 4.5 in the first tier.
🔗 Claude Haiku 5.5 in GitHub Copilot · GitHub Copilot models and pricing
Devin: 58,4 % on FrontierCode 1.1, ahead of Claude Sonnet 5
7 October — Cognition makes Claude Haiku 5.5 available in Devin. On FrontierCode 1.1, its proprietary benchmark that scores models on real engineering tasks based on code quality and mergeability, Haiku 5.5 scores 58,4 %, ahead of Claude Sonnet 5, at roughly one-eighth of its cost per task according to Cognition, which provides no precise amount. It also joins the sidekicks (sidekicks) in Devin Fusion, which pairs a primary model with a supporting model: with Claude Opus 5.5 as the primary model and Haiku 5.5 as support, Fusion maintains a score of 66,2 while reducing cost and latency.
| Model evaluated by Cognition | FrontierCode 1.1 Extended score |
|---|---|
| Claude Sonnet 5 | 56,2 |
| Kimi K3 | 58,2 |
| Claude Haiku 5.5 | 58,4 |
| GPT-6.1 Sol | 60,4 |
| Claude Sonnet 5.5 | 64,4 |
| Claude Opus 5.5 | 65,3 |
Coding agents: Claude Code 2.1.295, Codex CLI 0.162.0, Antigravity CLI 1.3.1, VS Code 1.141 and Delta 0.19
Five coding agent tools receive updates: Claude Code tightens its hooks, Codex CLI manages worktrees, Antigravity CLI redesigns change review, VS Code arranges agent sessions in a grid, and Delta expands teamwork capabilities.
Claude Code 2.1.295: hooks that block when they fail
8 October — Claude Code had two releases during the day. Version 2.1.294 fixes prompt and agent hooks written as natural-language instructions, which could allow through what they were supposed to block. Version 2.1.295 has 143 entries, including 97 fixes. With the new onFailure: "block" option for command and HTTP hooks, a hook that fails to start, times out or returns an unexpected code blocks the action instead of allowing it through; several fixes follow the same approach for mod guards and the --tools and --restricted options. Claude Code also supports the Program Status Protocol (OSC 7501), which lets compatible terminals display whether the agent is working, waiting or finished, and MCP tool descriptions loaded through tool search increase from 2 048 to 16 384 characters.
🔗 Claude Code 2.1.295 · Claude Code 2.1.294
Codex CLI 0.162.0: managed worktrees and pinned tasks
8 October — Codex CLI 0.162.0, the first stable release since the previous day’s 0.161.0, lists 225 PRs. When the worktrees feature is enabled, Codex has tools to create and list managed Git worktrees from trusted local projects, and the agent command center allows tasks to be pinned with the p key, in a shared “Pinned” group when the server supports it. The terminal gains /copy to browse and copy transcript blocks, Ctrl+Insert to copy a selection, and a scroll-speed setting, while URLs become clickable even in approvals and MCP prompts. Custom model providers compatible with the Responses API can declare live web access and remote compaction. On the fixes side, apply_patch preserves CRLF line endings, the Linux sandbox is hardened, and Windows releases receive a signed PowerShell installer.
🔗 Codex CLI 0.162.0 release notes
Antigravity CLI 1.3.1: redesigned review in /diff, medium verbosity by default
7 October — Antigravity’s CLI released five versions between 2 and 7 October. Version 1.3.1 improves code review in the terminal: in the file view of /diff, the left and right arrows open the adjacent file at its first change, each file remembers its position, n and N move on to the next or previous file, and the header indicates the current location. Its ten fixes address, among other things, /rewind in very long conversations, plugins on Windows, and sessions using GEMINI_API_KEY keys, which retry when the connection to the Gemini API drops. Version 1.3.0, released on 6 October, changes the default verbosity from “high” to “medium,” which groups tool calls and reasoning into concise summaries; /config lets users switch back. Earlier, version 1.2.16 assigned image generation to a built-in sub-agent, and version 1.2.15 introduced Android binaries for Termux.
🔗 Antigravity changelog, CLI tab
VS Code 1.141: agent sessions in a grid and worktree cleanup
7 October — VS Code 1.141 focuses on Copilot agent sessions. The Agents window can arrange multiple sessions in a grid, and the Chat: Open Worktree Cleanup command displays the disk space used by worktrees from inactive sessions so they can be deleted, on demand or automatically once the pull request is merged. Conversations started in Copilot CLI or the GitHub Copilot app appear from their first request, without reloading the editor, and a Codex conversation opened in the ChatGPT app or Codex CLI can be continued there with its history: selecting a Copilot model lets users continue with their GitHub Copilot subscription, for example after reaching their ChatGPT usage limit. For enterprises, the Copilot harness sandbox can be enforced through managed settings, in preview, and this harness is gradually becoming the default for users with experiments enabled.
Delta 0.19: teammate mentions and an inbox
7 October — Zed Industries releases Delta 0.19.0, its environment for coding together with agents, followed on 8 October by a corrective 0.19.1 release. Teammates can now be mentioned in a message or comment by typing @ followed by their username, and the notification arrives in a new inbox (Inbox). The @delta agent reads and modifies a large portion of the settings from the conversation. A setting, Agent Thread Creation, prevents agents from creating top-level conversations without disabling sub-agents, the agent can merge worktrees from multiple threads, and CLI commands run without opening a window. In total, the notes list 11 new features, 22 improvements and 45 fixes. On 8 October, a post shared by Zed describes how the team replaced Google Docs with Delta for its internal documents.
🔗 Delta release notes · Delta post on moving away from Google Docs
Generative media: Nano Banana 2.1, Brand Kit and ElevenReader in Brazil from ElevenLabs, SemanTok from Stability AI
Four announcements involving images, video and voice: an image model at half the price from Google, reusable brand guidelines and an audio reading app from ElevenLabs, and research on video generation from Stability AI.
Nano Banana 2.1 generally available in the Gemini API, at half the price
6 October — Google made Nano Banana 2.1 (gemini-nano-banana-2.1) generally available in the Gemini API and Google AI Studio, without a blog post. This image generation and editing model updates Nano Banana 2, with better visual quality, improved text rendering, more consistent characters across turns, and panoramic formats from 1:4 to 8:1, without tiling artifacts in 2K and 4K. It accepts up to 14 reference images and uses Google Search for grounding. The price per image is halved in 1K and 2K, and Google is deprecating Nano Banana 2 (gemini-3.1-flash-image), with no shutdown date.
| Output pricing | Nano Banana 2.1 | Nano Banana 2 |
|---|---|---|
| 1K image | 0,0336 dollar | 0,067 dollar |
| 2K image | 0,0504 dollar | 0,101 dollar |
| 4K image | 0,113 dollar | 0,151 dollar |
| Million image tokens | 30 dollars | 60 dollars |
🔗 Gemini API release notes · Nano Banana 2.1 model page
ElevenLabs Brand Kit: brand guidelines stored in the workspace
8 October — ElevenLabs adds Brand Kit to ElevenCreative, its creation suite. Colors, fonts, logos with their usage rules, reference images and brand rules, including what the brand never does, form a single workspace object that can be referenced with @ in Image & Video, in Flows or through the ElevenCreative agent. According to ElevenLabs’ thread, a kit can be created in a few minutes by pasting a website address or uploading the organization’s assets; the whole team then generates content from the same instructions, and an agency can create a kit for each client. ElevenLabs plans to extend the kits to the brand’s voice and sound, with no date announced. Brand Kit is available now, with no pricing specified; HeyGen (August) and Runway (September) already offer comparable tools.
🔗 ElevenLabs post: Brand Kit · @ElevenLabs thread
Stability AI’s SemanTok: a 201M model that matches a model 3,4 times larger
7 October — Stability AI’s Interactive Research team presents SemanTok, a video tokenizer for world models (world models) that generate video token by token, announced on X on 8 October. In “coarse-to-fine” tokenizers (coarse-to-fine), the first tokens describe the scene as a whole and subsequent tokens add details; SemanTok makes these initial tokens semantically richer and therefore easier to predict. To do so, it injects frozen DINO features into its encoder and adds lightweight heads that reconstruct them from each token prefix. According to Stability AI, a 201M-parameter autoregressive model built on SemanTok matches or exceeds a VideoFlexTok model 3,4 times larger, and short prefixes cost less to predict. The paper is on arXiv; neither code nor weights are mentioned.
🔗 Stability AI research post · Paper on arXiv
ElevenReader arrives in Brazil with Fábio Porchat’s voice and 70 000 books
October 8 — ElevenLabs launches ElevenReader in Brazil, its app that turns books, documents and articles into audio, free on iOS and Android, with the voice of actor and comedian Fábio Porchat. The catalog is based on a partnership with Bookwire Brasil: 70 000 titles in Brazilian Portuguese, licensed from the country’s leading publishers, most of which had no audio edition. ElevenLabs has also produced Dom Casmurro, by Machado de Assis, narrated by Fábio Porchat with original music and sound design. According to Bookwire Brasil’s managing director, this is the largest selection of Brazilian books ever offered in audio.
| Offering in Brazil | Access conditions |
|---|---|
| ElevenReader app (iOS, Android) | Free |
| Dom Casmurro narrated by Fábio Porchat | Free for everyone |
| 70 000 titles licensed through Bookwire | With ElevenReader Ultra, price unspecified |
🔗 ElevenLabs post: ElevenReader in Brazil
Specialized models: LightOnOCR-3 for documents, Falcon-ASR for Arabic, Carbon-A for genomes
Three specialized models, two with open weights, for reading documents, transcribing Arabic and annotating genomes.
LightOnOCR-3: three open OCR models that also identify page layout
October 8 — LightOn releases LightOnOCR-3 under the Apache 2.0 license, in three sizes (0.8B, 1B and 4B). In addition to transcribing a page, the models can identify each region: with the visual grounding prompt (grounding), each block is preceded by a labeled bounding box, images receive a short description and charts become HTML data tables. On an H100, the 4B processes 21 % more pages per second than Chandra-OCR-2, which is also built on Qwen3.5. LightOn notes that its scores are calculated after output normalization, published with the repository, which substantially changes the scores of all the leading models.
| Benchmark | LightOnOCR-3-4B | LightOnOCR-3-0.8B | LightOnOCR-3-1B | Cited reference |
|---|---|---|---|---|
| olmOCR-Bench, overall score | 86,3 | 85,5 | 84,5 | Infinity Parser Pro (35,1B) 87,6 ; Chandra 2 85,8 |
| ParseBench, five categories | 75,1 | 74,6 | 71,4 | Infinity Parser Pro 74,3 |
| fr-bench-pdf2md, French documents | 74,1 | 70,5 | 69,6 | Chandra 2 69,0 ; MistralOCR4.1 54,1 |
Falcon-ASR: 1,6 billion parameters for Emirati Arabic, with no weights released
October 7 — Abu Dhabi’s Technology Innovation Institute (TII) introduces Falcon-ASR, a speech recognition model with 1,6 billion parameters focused on Arabic, particularly the Emirati dialect. The same set of weights also transcribes English, French, Spanish and Portuguese, without requiring the language to be specified, with timestamps for every word. Like Falcon-OCR-Arabic and Falcon-Emirati, introduced the previous day, it is available only as a demo: TII announces API access and native apps without releasing the weights.
| Word error rate (WER) evaluation | Falcon-ASR score | Cited reference |
|---|---|---|
| Average across six Arabic datasets (Open Universal Arabic ASR) | 20,92 % | 23,17 %, best published result as of September 30 |
| Emirati, TII internal evaluation | 22,73 % | Qwen3-Omni, second place, 4,07 points higher |
| Average across seven English datasets (Open ASR Leaderboard) | 5,74 % | No reference cited |
Carbon-A: 566 million candidate genes across 22 617 species
October 8 — Hugging Face’s biology team (HuggingFaceBio) releases Carbon-A, a model with 1,2 billion parameters under the MIT license that identifies protein-coding regions directly in DNA, using a single model for mammals, plants, fungi and protists. Across 42 reference genomes, it achieves an average nucleotide-level F1 of 0,944 and, according to the authors, outperforms all seven tools compared, including AUGUSTUS, Helixer and Tiberius. The team ran it on GenBank’s eukaryotic genomes to build the Carbon Annotation Database: 48 167 assemblies from 22 617 taxa and 566 million predicted coding loci, covering 11 times more taxa than the training corpus. Laboratory validation with ActiveSite and UCSD, using full-length RNA sequencing, confirms some of the genes; the authors note that this demonstrates transcription, but not yet protein production. A new batch is planned in three weeks.
Reinforcement learning: TermGrade’s 1 004 environments and TRL v1.15.0
Two releases for reinforcement learning: graded terminal environments and a library that supports training on longer sequences.
TermGrade: 1 004 terminal environments graded by six models
October 8 — The company ai& releases TermGrade, 1 004 terminal environments for training and evaluating agents through reinforcement learning. Each task runs in a Linux container with an instruction and tests, and was retained only if its own reference solution passed its tests: of 66 165 candidate tasks written by DeepSeek-V4-Pro, 1,5 % passed this filter. Six configurations of open models, from Kimi-K3 to gemma-4-31B-it, attempted every task, and all 36 144 attempts are published, including failures. A model learns most from tasks it succeeds at roughly half the time: trained on 151 tasks in this range, gemma-4-31B-it reaches 46,1 on Terminal-Bench 2.1, a gain of 3,1 points over the base model, with an average gain of 2,1 points across five training runs. Everything is released under Apache-2.0.
🔗 TermGrade post · TermGrade collection on Hugging Face
TRL v1.15.0: a fused LM head for sequences up to 6,9 times longer
October 8 — TRL v1.15.0, Hugging Face’s training library, introduces a fused LM head (fused LM head) that directly calculates each token’s log-probabilities and entropy with a Triton kernel, without ever constructing the full logits tensor; it is enabled by default for SFT, DPO, KTO, GRPO, RLOO and distillation. At 8 192 tokens, peak memory falls by 52 to 82 % depending on the method. These measurements were taken on a B300 capped at 79 Gio with random gemma-3-1b weights, whose vocabulary contains 262 000 tokens; the gains are greatest for small models with large vocabularies. The release also introduces breaking changes: Python 3.10 support ends, use_liger_kernel is deprecated and the experimental MiniLLM trainer is removed.
| Training method | Maximum length in v1.14.2 (tokens) | Maximum length in v1.15.0 (tokens) | Gain factor |
|---|---|---|---|
| DPO | 10 240 | 59 392 | ×5,80 |
| KTO (batch of 2) | 9 216 | 63 488 | ×6,89 |
| GRPO | 28 672 | 114 688 | ×4,00 |
| RLOO | 23 552 | 100 352 | ×4,26 |
| SFT (nll loss) | 20 480 | 107 520 | ×5,25 |
Inference: llama.cpp in Windows ML, open-source ML Drift, Together AI on IBM Cloud’s B300 GPUs
Three announcements for running models, from Windows PCs to GPU clusters.
llama.cpp comes to Windows ML as an experimental feature
October 7 — At its Windows event, Microsoft adds experimental support for llama.cpp to Windows ML, its local inference framework: users can download a GGUF model from Hugging Face and run it locally through the same Windows ML stack, using new APIs dedicated to specific tasks. A lower-level native Windows runtime API also enters experimental preview. Microsoft says it is contributing directly to the llama.cpp project with NVIDIA and the community: optimized CUDA kernels, better scheduling between CPU and GPU, CUDA graphs, speculative decoding (Eagle-3, MTP, D-Flash2), execution across multiple GPUs and the NVFP4 format. The post places these new features on PCs equipped with NVIDIA RTX Spark chips, such as the Surface Laptop Ultra. On X, Georgi Gerganov of llama.cpp welcomes the project’s presence at the Windows event and sees it as a sign that software and hardware stacks are finally coming together.
🔗 Microsoft Foundry on Windows post · Tweet @ggerganov
ML Drift: LiteRT’s GPU engine released as open source
October 8 — The Google AI Edge team releases ML Drift, its GPU compute engine for on-device inference, as open source under the Apache 2.0 license. It abstracts away differences between OpenGL ES, OpenCL, Metal and WebGPU, serves as LiteRT’s GPU acceleration engine, is available as a standalone library and succeeds the TensorFlow Lite GPU delegate, which will no longer receive new features. The engine switches kernels depending on whether the LLM is reading the prompt or generating tokens, and provides a SKILL.md guide for coding agents to write and verify shaders. It is already running in production: up to 40 % lower per-frame latency for YouTube Shorts effects, and up to 30 % gains for Lightroom and Photoshop on mobile, according to the post; on Gemma, Google measures up to 12 % lower memory usage than other software frameworks (frameworks).
🔗 ML Drift announcement on the Google Developers Blog · google-ai-edge/ml-drift repository
Together AI becomes the first customer of a B300 GPU inference cluster on IBM Cloud
October 6 — Together AI announces that it is working with IBM and NVIDIA to expand its enterprise inference capacity, starting with a large NVIDIA B300 GPU cluster hosted on IBM Cloud and connected through NVIDIA’s Spectrum-X Ethernet network. According to Together AI, this is the first dedicated inference cluster of this scale on IBM Cloud, and the company is its first customer: it operates the inference layer, IBM provides the cloud, and NVIDIA provides the chips and network. Together AI attributes this expansion to demand for open models: it says it serves hundreds of trillions of tokens every month to more than a million developers. The post provides neither the cluster’s size nor a timeline.
🔗 Together AI post · Tweet @togethercompute
Briefs
- Faster Codex steering — In the ChatGPT desktop app, Codex responds sooner to follow-up messages sent during a task to correct an approach or change direction; the Follow-up behavior setting determines whether these messages steer the current run or wait for the next one. 🔗 source
- ChatGPT Go in Warp — ChatGPT Go subscribers can now use their plan in Warp for any terminal question or coding task; according to a member of OpenAI’s Sign in with ChatGPT team, quoted by Warp, included usage is available to Go customers, alongside Plus and Pro, across all partner apps. 🔗 source
- Oracle — According to an OpenAI case study, Oracle has 130,000 active ChatGPT users and more than 95,000 Codex users; talent market research that used to take 2 to 4 days can be prepared in 15 to 20 minutes, and Codex translates business questions into SQL queries. 🔗 source
- Pollo AI — This video creation platform, which claims more than 26 million users, routes its agent’s requests with GPT-5.6, assigns difficult tasks to GPT-6 Astra and all its images to GPT-Image-2.5, and says it reduces the time spent choosing a model by more than 50%. 🔗 source
- Devin release notes for October 7 — A notification inbox brings together session alerts, with configurable delivery to phones, merge queue status appears throughout the app, and private keys stored as secrets are masked line by line in command output. 🔗 source
- Copilot CLI 1.0.94 — Now stable after six prereleases, it adds Claude Haiku 5.5 to the model selector, passes the displayed shell code to the Assisted Permissions judge to avoid unnecessary approvals, and allows an MCP server to be enabled before discovery. 🔗 source
- Gemini CLI v0.65.0-nightly.20261008 — Telemetry supports custom OTLP headers (
telemetry.otlpHeaders), a request open since October 2025; the CLI no longer loops between verification and OAuth, and the history always ends with a user message, preventing a 400 error after/rewind. 🔗 source - Antigravity SDK 0.1.21 isolates experimental APIs and configures sub-agent skills — The first version listed in the changelog since 0.1.18 on September 21: an
@betadecorator and angoogle.antigravity.betamodule separate preview features from the stable API, andskills_configlets a sub-agent inherit its parent agent’s skills, replace them, or do without them. 🔗 source - Developer Knowledge API — Google presents the ecosystem around this API, which serves its developer documentation (Google Cloud, Firebase, Android) as up-to-date Markdown: an
gcloud developer-knowledgecommand preinstalled in Cloud Shell, a skill for coding agents connected to an MCP server, with Antigravity, Claude Code, Cursor and GitHub Copilot mentioned, and libraries in seven languages. 🔗 source - AQuA — Google Cloud releases this quality agent (Ambient Quality Agent) as a reference implementation. It samples up to 1,000 sessions from an ADK agent in production, groups failures, tracks them in BigQuery and launches a root cause diagnosis with Gemini 3.8 Flash. 🔗 source
- Drafts count toward pull request limits — In response to low-quality contributions, GitHub allows draft pull requests to be included in per-user limits, whereas contributors could previously open an unlimited number of them. 🔗 source
- Archiving pull requests — It is no longer restricted to administrators: the triage, write, maintain and admin roles can archive a pull request, closing it, hiding it from the public and blocking all further activity, including administrator comments. 🔗 source
- GitHub timelines and screen readers — Screen readers navigate issue and pull request timelines as lists, with the item count and current position, and announce loaded events, on github.com and GitHub Enterprise Server 3.23. 🔗 source
- Paper Reproductions with ML Intern — At Hugging Face, Abubakar Abid announces a feature that tasks agents with independently reproducing experiments from a published paper, starting from Paper Pages, to produce open artifacts that help identify solid research; no figures or documentation. 🔗 source
- Hugging Face chat over SSH — Adrien Carreira, Hugging Face’s head of infrastructure, presents a chat with open models accessible through the
ssh chat.hf.cocommand, with no setup or key required; no documentation or model list accompanies it. 🔗 source - huggingface_hub 2.2.0 — A small patch release:
hf jobs statsnow returns a snapshot and then returns control (-ffor live monitoring), all parallel downloads go through hf_xet, and two changes break compatibility. 🔗 source - swarmlab — Hugging Face has opened this testbed without an announcement. It studies how swarms of LLM agents from different providers find the truth or split, depending on communication topology, delays and roles; a single line of YAML is enough to test a hypothesis. 🔗 source
- Darwin-27B-ZTC — VIDRAFT releases a judge under Apache-2.0 that returns a probability distribution over possible answers in a single pass, without generating a token: 0.743 zero-shot accuracy on 2,000 typed-decisions judgments and a Brier score of 0.097, according to its authors. 🔗 source
- Superfluid — basecompute releases this local LLM server under Apache-2.0, compatible with the OpenAI, Anthropic and Ollama APIs, which prioritizes interactive questions over agent work: a 0.36-second response with eight active agents, compared with more than 10 seconds for llama-server, according to its authors. 🔗 source
- funes 1.8.0 — The tool that indexes coding agent sessions adds a multilingual embeddings model: on 30 questions about 2,000 conversations in Chinese, the correct conversation appears in the top eight results 30 times out of 30, compared with 23, and the mean reciprocal rank rises from 0.601 to 0.983. 🔗 source
- GLiNER and Indian languages — A community post shows that keeping combining marks within words, without retraining, raises GLiNER2’s average zero-shot F1 from 13.2 to 68.0 across nine Indian languages, and that of a model fine-tuned on Hindi from 59.9 to 82.6. 🔗 source
- Multilingual Terminal-Bench — LILT details its benchmark of 324 coding tasks in ten languages: Claude Opus 5.5 outperforms Opus 5 by about 3 points with 30% fewer tokens, resulting in a cost per task roughly 45% lower; Gemini 3.8 Flash takes a median of 41 turns per task, compared with 5 for GPT-6 Sol. 🔗 source
- GPU collective primitives — Milos Kotlar predicts the communication time of five primitives on four RTX 4090 GPUs with a mean error of 10.9%, compared with 22.9% for a common model, but with a worst case of 61.6%. 🔗 source
- Best-of-N in video reasoning — facebookresearch posts the code for the study Selecting or Solving? without an announcement, under the noncommercial CC BY-NC 4.0 license, covering best-of-N selection and verifier juries in video reasoning. 🔗 source
- KDD Cup 2026 Data Agents — NVIDIA details the harness that placed its KGMON team second with a required small LLM: data converted into SQLite tables, the schema provided upfront, a small toolset and targeted document reading; the post gives neither a score nor a model name. 🔗 source
- Microduck — Matthieu Lapeyre of Pollen Robotics announces more than three hours of battery life on a 2,600 mAh battery for this small bipedal robot equipped with the new in-house electronics board, with a processor that tops out at 60 °C instead of 114 °C. 🔗 source
- Albums on Suno — Introduced on October 5, the feature is live: users bring their tracks together into a complete release, choose the cover, arrange the tracks and publish when they want; neither plans nor track limits are specified. 🔗 source
- HeyGen and Ryan Serhant — HeyGen launches a daily video series hosted by the official avatar of real estate agent Ryan Serhant, featured in the Netflix series Owning Manhattan, on Instagram, TikTok and X; no financial details. 🔗 source
- Runway and the Tech Coalition — Runway joins this alliance against online child sexual exploitation and its Lantern program for sharing signals across platforms, and reiterates its safeguards: data filtering, adversarial testing, hashing, reporting to NCMEC and C2PA provenance. 🔗 source
- SpaceXAI and Omarchy — SpaceXAI becomes a founding sponsor of the Omacom Foundation with $1.5 million in Grok tokens; Grok 4.7 and its successors will power the project’s agents, including Omabot, which reviews pull requests. 🔗 source
- Perplexity’s guardrails guide — Perplexity publishes a guide that places AI guardrails at three stages (input, agent actions, output), compares five types of protection and illustrates seven cases, from prompt injection to hallucination; no new functionality. 🔗 source
- pplx-decider-v1.1-27b on OpenRouter — Perplexity’s open-weight decision model arrives on OpenRouter at the Decisions API price: $0.02 per million input tokens, free output and a 262,000-token context. 🔗 source
- Cohere in Ottawa — Cohere opens its first office in Ottawa, near Lansdowne Park, where nearly 30 employees already work (commercialization, engineering, infrastructure, security, sovereign AI); neither floor area nor a staffing target is provided. 🔗 source
What this means
The assistant now produces work artifacts, more than just answers. A Claude Dashboards dashboard updates when the data changes, a Claude Motion animation remains editable word by word because it is written in code, and Docs, Slides and Design, now out of beta across all plans, already claim more than 45 million creations. These artifacts then move through existing tools: the dashboard goes to Grafana or PostHog, the animation to Runway or Luma, which add generated shots to it. ElevenLabs’ Brand Kit follows the same logic: brand guidelines become a workspace object, referenced with @, rather than an instruction repeated in every prompt.
Security is changing scale, both in defense and in rule enforcement. Anthropic finds that discovering vulnerabilities has never been easier: more than 29,000 candidates in six months, including around 6,000 manually triaged. OSS Scanner commits to sending reports without human review, with an expected true-positive rate above 90%, to ease this bottleneck. Usage rules are becoming more specific in parallel: Anthropic’s 2026 policy groups deceptive campaigns into a dedicated section, and OpenAI presents Dark Clark as the first category 5 operation it has disrupted since it began publishing reports. In the tools, Claude Code 2.1.295 now lets a failing hook block the action rather than allow it through.
Speed and price are becoming settings users can choose. Ultrafast sells GPT-6.1 Sol at six times the price for up to 8 times the speed; according to GitHub and Cognition, Claude Haiku 5.5 matches or outperforms Sonnet 5 on coding, at roughly one-eighth of the cost per task according to Cognition; Nano Banana 2.1 halves the price of a 1K image. Inference spans devices, with llama.cpp in Windows ML and ML Drift on phone and computer GPUs, through to the B300 cluster for which Together AI is the first customer on IBM Cloud, and TRL reduces peak memory use during training at 8,192 tokens by 52 to 82%.
Finally, many of today’s figures are measured by those who publish them. FrontierCode is Cognition’s proprietary benchmark, GitHub says Haiku 5.5 matches Sonnet 5 without publishing figures, LightOn calculates its scores after a normalization that changes those of all the leading models, TII evaluates Falcon-ASR in Emirati Arabic only on an internal dataset and does not publish its weights, and Carbon-A’s authors compare it themselves with seven tools. Open science provides a counterpoint: TermGrade publishes its 36,144 attempts, including failures, Carbon-A its database of 566 million candidate genes, and Claude Science’s ultraviolet map distinguishes what is measured from what is predicted, pixel by pixel. Anthropic’s $150 million and NVIDIA’s $1 billion for the Genesis Mission push in the same direction: putting these tools into researchers’ hands.
Sources
- Claude Dashboards and Claude Motion (Claude post)
- Dashboards and Motion launch thread (@claudeai)
- Claude Motion (@claudeai)
- Runway and Claude Motion (Runway)
- Luma and Claude Motion (Luma)
- Anthropic Cyber Mission (Anthropic)
- OSS Scanner (Anthropic’s Frontier Red Team)
- anthropics/oss-scanner repository
- Ultrafast for GPT-6.1 Sol (@OpenAIDevs)
- OpenAI API changelog
- OpenAI API pricing
- Ultrafast in Codex (documentation)
- Commitment to the Genesis Mission (Anthropic)
- Genesis Mission announcement (@AnthropicAI)
- NVIDIA commits 1 billion dollars to advance US science (NVIDIA)
- The missing map of the sky (Anthropic Science)
- Ultraviolet sky map (@AnthropicAI)
- 2026 usage policy update (Anthropic)
- Disrupting AI-enabled false front operations (OpenAI)
- Claude Haiku 5.5 in GitHub Copilot (GitHub Changelog)
- GitHub Copilot models and pricing
- Claude Haiku 5.5 in Devin (Devin blog)
- Claude Code 2.1.295
- Claude Code 2.1.294
- Codex CLI 0.162.0
- Antigravity changelog, CLI tab
- VS Code 1.141 release notes
- Delta release notes
- Delta post on moving away from Google Docs
- Gemini API release notes
- Nano Banana 2.1 model card
- Introducing Brand Kit (ElevenLabs)
- Brand Kit (@ElevenLabs)
- SemanTok (Stability AI)
- SemanTok paper on arXiv
- ElevenReader in Brazil (ElevenLabs)
- LightOnOCR-3 (Hugging Face blog)
- Falcon-ASR (Hugging Face blog)
- Carbon-A and the Carbon Annotation Database (Hugging Face blog)
- TermGrade (ai&)
- TermGrade collection on Hugging Face
- TRL v1.15.0
- llama.cpp in Windows ML (Microsoft Foundry on Windows)
- llama.cpp at the Windows event (@ggerganov)
- ML Drift (Google Developers Blog)
- google-ai-edge/ml-drift repository
- Together AI, IBM Cloud and NVIDIA (Together AI blog)
- B300 cluster on IBM Cloud (@togethercompute)
- ChatGPT release notes
- ChatGPT Go in Warp (@warpdotdev)
- Oracle (OpenAI case study)
- Pollo AI (OpenAI case study)
- Devin release notes for October 7
- Copilot CLI 1.0.94
- Gemini CLI v0.65.0-nightly.20261008
- Antigravity changelog, SDK tab
- Developer Knowledge API ecosystem (Google Developers Blog)
- AQuA, a quality agent for ADK agents (Google Developers Blog)
- Draft pull requests count toward pull request limits (GitHub Changelog)
- Triage role users or higher can now archive pull requests (GitHub Changelog)
- Screen readers can navigate timelines as lists (GitHub Changelog)
- Paper Reproductions with ML Intern (@abidlabs)
- Chat over SSH (@XciD_)
- huggingface_hub 2.2.0
- huggingface/swarmlab repository
- Darwin-27B-ZTC (Hugging Face blog)
- Superfluid (Hugging Face blog)
- funes 1.8.0 (Hugging Face blog)
- GLiNER and Indian languages (Hugging Face blog)
- Multilingual Terminal-Bench (Hugging Face blog)
- Communication times for GPU collective primitives (Hugging Face blog)
- facebookresearch/best-n-selection-video-reasoning repository
- Lessons from the KDD Cup 2026 (NVIDIA technical blog)
- Microduck autonomy (@matth_lapeyre)
- Albums on Suno (@suno)
- Ryan Serhant’s daily series (@HeyGen)
- Runway joins the Tech Coalition (Runway)
- SpaceXAI, founding patron (omarchy.org)
- AI guardrails (Perplexity)
- pplx-decider-v1.1-27b on OpenRouter
- Cohere’s new Ottawa office (Cohere)