ai-powered-markdown-translatorArticle translated from French to English using gpt-6.1-sol.
OpenAI will add an invisible watermark to ChatGPT and Codex text for its European Union users in the coming weeks, in response to the AI Act requirement to make generated text machine-identifiable, and is making this watermark available as an option to its API customers worldwide starting today. Among agents, memory is taking hold: Cognition gives Devin memory across sessions and publishes its format as an open standard, and Cohere launches North 2 with memory that retains context from one session to the next; GitHub, meanwhile, publishes ReviewBench, an open benchmark for AI code review. Among open models, Reflection AI introduces Beam, with 501 billion parameters, whose weights are promised later in October.
OpenAI’s invisible watermark: ChatGPT and Codex text marked in the European Union, a worldwide API option
October 5 — OpenAI publishes its response to European rules on text provenance. The AI Act requires generative AI providers to make the text they produce identifiable in a machine-readable way; OpenAI is responding in stages, warning from the outset that current text watermarking technologies have substantial limitations (significant limitations).
| Affected audience | What changes | Announced timeline |
|---|---|---|
| ChatGPT and Codex users in the European Union, all plans | Invisible watermark added to eligible text | In the coming weeks |
| API customers worldwide | Optional (opt-in) watermark on selected, unnamed models, disabled by default | Starting October 5 |
| Approved researchers and expert organizations | Access to the detector, on a case-by-case basis, through an application | Applications opened October 5 |
OpenAI is not making this a worldwide default setting, and is working with cloud partners to offer the same watermark on the OpenAI models they serve in the coming weeks. The technology, textGrain, adds an invisible statistical signal to the model’s word choices, without any visible mark or special character; according to OpenAI, it matches or exceeds the other approaches tested, including SynthID for text. A technical report has been published, and OpenAI plans to release the code. The detector, however, is not public at launch because of the risk of missed watermarks and false positives: access follows the European Commission’s Code of Practice (Code of Practice).
The published figures primarily show the limits of detection:
| Measurement conditions | Published detection rate |
|---|---|
| 200-token passages, psychology-type content, 1 % target false-positive rate | approximately 80 % |
| 400-token passages, psychology-type content, 1 % target false-positive rate | approximately 95 % |
| Mathematics-type content, same 1 % threshold | substantially lower (value shown only in a chart) |
| 400-token passages in English (ELI5 dataset), unedited | approximately 92 % |
| Same passages, 10 % of words replaced with synonyms | 66 % |
| Same passages, 25 % of words replaced | 17 % |
Watermarks have limits. They’re often undetectable, especially in short passages. Rewriting or translating text can completely remove the watermark. — @OpenAI on X
On quality, OpenAI measures no significant difference across eight GPT-6 Astra benchmarks without and with watermarking (94,44 % versus 93,94 % on GPQA Diamond, 53,90 % versus 56,06 % on Terminal-Bench 4.0). The post also clarifies what a watermark does not indicate: the amount of human work involved, ownership of or responsibility for the text, the user’s identity, or accuracy; and its absence does not prove that a human wrote the text. For images and audio, nothing changes: openai.com/verify and the Content Provenance API remain available to organizations, and SynthID has supplemented the C2PA metadata of generated images since May 19.
🔗 OpenAI’s post on text provenance in the EU
Coding agents: Devin remembers across sessions, Cursor, Qwen Code, Together Link and IBM Bob
October 5 — Cognition introduces two related features in Devin: Memory, which persists from one session to the next, and Dreaming, a daily “dream” that reorganizes it. Memory retains what the agent learns while working with each user: preferences, corrections, lessons learned from a project or workflow. These are short notes rather than session summaries, each linked to the session in which the lesson was learned, and this memory remains personal: it does not become an instruction shared across the entire organization.
The notes live in the Memory Drive, a persistent Git repository of Markdown files organized by repository, project or topic. A deliberately short MEMORY.md file brings together general preferences and an index of everything else: Devin receives it at the beginning of each session, then retrieves useful notes with the tools it uses to browse code, without loading the entire archive into its prompt. Each session works on its own Git copy: Devin merges updates from other sessions, a revision check rejects stale writes, and conflicts are reported instead of being silently overwritten.
Dreaming is a daily background session (at night, Cognition’s thread specifies): Devin rereads past conversations in light of its memory, merges overlapping notes, removes transient details, looks for lessons that had not been recorded and deletes memories that no session has used. Access is through devin.ai, under Customize → Memory; the post says nothing about Devin Desktop or Devin CLI and announces no pricing.
Cognition also publishes the format as an open standard, Agent Memory Repo, under the MIT license. The specification, open to contributions, describes each session’s loop (clone the memory, search, update without a human in the loop, push after each change) and the combination of multiple memories in a single session, each in its own repository with its permissions and history. The use cases mentioned include team memory and agent swarms (agent swarms):
In early experimentation we saw a swarm of Devins using the memory as a way to coordinate at large scale without us prompting them to (we promise they didn’t hack anyone) — @cognition on X
A “dream” that reorganizes an agent’s memory already existed at Anthropic (Claude Managed Agents, in May) and OpenAI (ChatGPT memory, in June); the change here is memory stored in files versioned with Git, published as a specification that other agents can adopt.
🔗 Memory and dreaming, Cognition’s post 🔗 Agent Memory Repo specification
Cursor: steer its SDK agents during execution
October 5 — Cursor expands its SDK, which is used to launch its agents from TypeScript or Python code. The run.steer() method inserts a message into an agent’s current turn; if a subagent is working at that moment, it moves into the background and continues. Background subagents also report back: their result returns to the parent agent as an additional turn in the same run, instead of being lost when the parent turn ends.
| SDK addition | Scope indicated by the changelog |
|---|---|
run.steer(), steering during execution | Local agents in TypeScript; on a cloud agent, the message is sent as a normal follow-up |
| Results returned by background subagents | Local agents, in TypeScript and Python |
| MCP annotations for custom tools (readOnlyHint, destructiveHint…) | TypeScript; hints only, which the SDK does not enforce |
| Replaceable system prompt, rules and skills still loaded | Local agents in TypeScript; access enabled account by account |
No pricing accompanies these changes.
🔗 Cursor’s thread on X 🔗 Cursor SDK changelog
Qwen Code v0.25.0: workspace agents, email channel and Mem0 memory
October 5 — Qwen releases v0.25.0 of Qwen Code, its command-line coding agent, the first stable release since v0.24.7 on September 29: 42 features, including 23 for the Managed Agent, 79 fixes and no known breaking changes. Three additions, all requiring explicit activation, stand out. Workspace agents now collaborate locally: the daemon starts and resumes their turns, and the Web Shell lets users manage agents and tasks and call on one with @agent from a conversation. These persistent agents now support the A2A 1.0 protocol through shares that expire and can be revoked. Finally, Mem0 memory connects directly to the CLI: simply specify an endpoint and a key, and Qwen Code registers the corresponding MCP server itself. An email channel also gives the agent its own mailbox, using IMAP and SMTP, and Code Mode executes in parallel the Bash calls grouped by the model. TypeScript SDK v0.1.18 and Desktop application v0.25.0 followed the same morning, without a tweet from Qwen.
🔗 Qwen Code v0.25.0 release notes
Together Link: Claude Code, Codex, OpenCode and Pi on open models
October 5 — Together AI launches Together Link in beta, running already installed coding agents on open models hosted by the platform. An installation command (macOS or Linux) connects Claude Code, Claude Desktop, Codex, OpenCode and Pi to its API using the account’s key; the documentation adds ChatGPT Desktop and warns that commands, routing and the model list may still change. In addition to Auto mode, four models are offered, all with 1M context tokens:
| Available model | Equivalent in Claude Code’s /model menu |
|---|---|
| Kimi K3 | Opus |
| GLM 5.3 | Fable |
| GLM 5.3 Flash | Sonnet |
| DeepSeek V4.1 Flash | Haiku |
Auto mode, selected by default, can assign difficult tasks to Opus 5.5 when the user provides an Anthropic key, but the post and documentation describe this routing differently: a choice made once per session to preserve the prompt cache according to the post, request-by-request selection restricted to Claude Code and Claude Desktop according to the documentation. Together AI claims a spending reduction of more than 50 %, without publishing a methodology, and displays each session’s cost alongside what it would have cost on Opus 5.5.
🔗 Together AI’s post 🔗 Together Link documentation
Self-hosted IBM Bob relies on NVIDIA’s Nemotron 3 Ultra
October 5 — IBM explains why the self-hosted (self-hosted) version of Bob, its agentic development assistant, relies on NVIDIA’s Nemotron 3 Ultra alongside Poolside Laguna S 2.1. This option, which became generally available the previous week, runs the assistant on the customer’s infrastructure, including disconnected (air-gapped) environments. The team evaluated models against four criteria (hardware footprint, coding accuracy, latency, openness of weights), testing them with Bob’s agent harness and using closed models from Anthropic, OpenAI and Google as references. Initially too verbose, Nemotron was tuned with NVIDIA: prompts, sampling parameters, reasoning settings and harness fixes. According to IBM’s internal tests, Nemotron 3 Ultra offers up to roughly 4 times lower latency on Blackwell GPUs than leading SaaS models accessed over the network, and strictly adheres to tool schemas; Laguna S 2.1 was selected for its performance-to-footprint ratio. The post publishes no accuracy scores.
🔗 IBM’s post on self-hosted Bob
ReviewBench: GitHub launches an open benchmark for AI code review
October 5 — GitHub launches ReviewBench in research preview (research preview), an open benchmark for AI code review agents. The dedicated site publishes the full dataset, leaderboard, methodology, judge prompt and configuration, and a self-service runner. The post is authored by Michelle Zhou and Alejandro Carderera de Diego, and its acknowledgments associate GitHub and Microsoft with the project.
The corpus brings together 219 pull requests from 187 public open source repositories, across 19 languages. Its distributions of languages and repository sizes mirror those of GitHub, established from 103,9 million pull requests, with a deliberate weighting toward medium and large changes. The ground truth (golden set) combines human reviewers, issues inferred from follow-up commits, deterministic analysis tools, and several leading LLMs; findings are validated by an LLM judge, Claude Sonnet 5 according to the post. Senior engineers who were not involved in building the dataset relabeled every finding: their verdict agrees with ReviewBench in 96,6 % of cases. “Grounded” recall (grounded), measured solely against the ground truth, serves as the primary comparison.
According to GitHub’s initial leaderboard, Copilot code review ranks first:
| Evaluated agent (configuration) | Grounded F1 | Grounded precision | Grounded recall | Run date |
|---|---|---|---|---|
| Copilot Code Review (Balanced) | 40,1 % | 87,8 % | 26,0 % | October 1, 2026 |
| Devin AI | 37,0 % | 84,0 % | 23,8 % | September 28, 2026 |
| Qodo | 35,1 % | 85,3 % | 22,1 % | September 28, 2026 |
| Codex (GPT-5.6 Sol · Ultra) | 32,3 % | 87,0 % | 19,9 % | July 19, 2026 |
| Cubic | 27,3 % | 85,5 % | 16,3 % | June 29, 2026 |
| Greptile | 27,2 % | 86,1 % | 16,2 % | June 16, 2026 |
| Cursor | 17,3 % | 87,7 % | 9,6 % | September 27, 2026 |
The site explains that these initial entries were produced by the ReviewBench team by running each vendor’s public product on the indicated date, and that the vendors neither ran nor verified them. Anyone can now submit their agent: the entrant provides a container image and model key, GitHub provides the judge, and testing on 25 pull requests precedes a full run in three passes; a score is published only after validation by a maintainer, and only if it improves the agent’s previous score or is its first entry.
GitHub also uses it to improve Copilot code review. In an A/B test, a multi-model review at the Lite level, combining several independent runs, increased recall by 13,6 % while reducing the cost per review by 8,0 %, in the direction predicted offline. The post contradicts itself on comment volume: +61 % according to the text, +25,0 % according to its chart.
🔗 GitHub’s post on ReviewBench 🔗 ReviewBench site and leaderboard
Cohere launches North 2 and forms a global alliance with PwC
October 5 — Cohere launches North 2, the new version of North, its enterprise AI agent platform, announced as “coming soon” on September 9 and then promised for October. The launch thread claims more than 15 new features (15+ new features), a count the post does not repeat. At the heart of the release is a new agent harness (agent harness), redesigned for multi-step automations and agents, on a platform that remains agnostic: Cohere models or the customer’s models.
| Functional area | New features cited by Cohere |
|---|---|
| Agents | Reusable agents and automations, created through prompts and shared across the organization; multi-agent orchestration; memory that retains context across sessions |
| Productivity | Skills (Skills), libraries (Libraries), and applications that produce presentations, dashboards, and documents; Automations, launched in late July, with ready-to-use templates and a visual editor |
| Connectors | Slack, SharePoint, OneDrive, Outlook, Exchange, Jira, Linear, Notion, and GitHub; financial data providers (PitchBook, FactSet, and others) only planned |
| Deployment and security | Self-hosted, VPC, hybrid, on-premises, or fully disconnected; SOC 2 Type 2, ISO 27001, and ISO 42001; autonomy policies with human approval of critical decisions |
| Governance | North Admin console, query and token usage tiers per user or group, alerts before limits are reached |
Cohere cites two customers, LG CNS in South Korea and Bell Cyber in Canada, and highlights improved throughput for its models on NVIDIA Blackwell and Hopper GPUs, without figures, with a quote from Kari Briski, head of generative AI at NVIDIA. No pricing or rollout schedule is provided: the post directs readers to book a demo with the sales teams.
🔗 North 2 launch post 🔗 North 2 launch thread on X
The global alliance with PwC, starting in Canada
October 5 — On the same day, PwC and Cohere announce a global alliance (global alliance) starting in Canada, in a PwC Canada press release datelined Toronto and shared in a short Cohere post. The division of roles is clear: Cohere provides North, its enterprise models, and its information retrieval tools; PwC brings its expertise in industry, regulation, risk management, and transformation, from selecting use cases to integrating AI with enterprise data and systems. The announced uses cover enterprise search, access to knowledge, in-depth research, decision support, and task automation, in private cloud, on-premises, or disconnected environments, primarily targeting regulated sectors. The press release quotes Domenic Marino and Annie Veillet of PwC Canada, and Aidan Gomez of Cohere; it provides no financial amount, customer, or timeline beyond Canada. Cohere had already formed an alliance with OpenText on September 16.
🔗 Cohere’s post on the alliance with PwC 🔗 PwC Canada press release
Open models: Beam, MetaEncoder-30B, and Gemma 4 in plant genomics
Three open models are in today’s news, at three different stages: one is promised, another has been uploaded without an announcement, and the last is being highlighted a month after its release.
Beam: Reflection AI’s open model with 501 billion parameters
October 5 — Reflection AI introduces Beam, its first open-weight model: a sparse MoE with 501 billion parameters, including 23 billion active parameters, designed for coding, reasoning, and agentic tasks, and trained end to end from scratch. Pretraining used 23 800 billion tokens; the reinforcement learning phase used 10 500 NVIDIA GB300 GPUs for four weeks and produced more than 100 million trajectories (rollouts).
According to Reflection AI’s tables:
| Evaluated benchmark | Beam | Nemotron 3 Ultra | GLM 5.3 | Kimi K3 | Qwen 3.8 Max | DeepSeek V4.1 Flash |
|---|---|---|---|---|---|---|
| SWE-bench Verified | 80,9 | 70,7 | — | — | — | — |
| SWE Bench Pro v1 | 65,5 | 46,4 | — | — | 67,7 | — |
| Terminal Bench v2.1 | 80,1 | 56,4 | 88,2 | 88,3 | 86,6 | 90,6 |
| DeepSWE v1.1 | 44,4 | — | 61,0 | 68,0 | 51,0 | 74,2 |
A dash indicates a score not reported in the post. Reflection AI claims reasoning scores comparable to those of GLM-5.2 with 3 to 4 times less inference compute, and acknowledges that Kimi K3 remains ahead in raw capability. Nothing is available to download yet: Beam is undergoing final evaluations and adversarial testing (red-teaming), early access requires registration, and the weights, technical report, model card, and developer tools are promised later in October.
🔗 Reflection AI’s post on Beam
MetaEncoder-30B, the decision encoder Meta uploads without an announcement
October 5 — Meta has uploaded MetaEncoder-30B to Hugging Face under its facebook organization, with no announcement found: the repository, created on September 30 and modified on October 5, had only two downloads at the time of our check. The model card presents a multimodal “System One” encoder, the format used by Jev’s decision models: it is given a task in natural language (question, instruction, state description, criteria) and a list of candidates, in text with supporting images and videos, and scores each candidate. Since the task and candidates are encoded separately, the comparison reduces to a dot product, and a nearest-neighbor index can cover millions of candidates without rerunning the model. MetaEncoder is a contrastive fine-tuning of Muse Glimmer 30B, the open-weight agentic model released by Meta in August, from which it inherits the Apache 2.0 license; downloading requires accepting terms, and vLLM serves it for text and images.
| Evaluation benchmark | Score according to the model card | Metric used |
|---|---|---|
| JEVBench (original / easy / hard) | 0,9444 / 1,0000 / 0,7387 | Accuracy |
| ImaJEV-Bench | 0,8958 | Accuracy |
| NanoBEIR | 0,6632 | NDCG_linear@10 |
| MMEB-V3 (image / video / visual documents) | 0,7897 / 0,6046 / 0,8138 | Hit@1 / Hit@1 / NDCG@5 |
🔗 MetaEncoder-30B model card on Hugging Face
Living Models combines Gemma 4 and BOTANIC-1 to decode plant DNA
October 5 — Google highlights the work of Living Models, a Paris-based AI biology lab, in an X Article by @GoogleAI and on its Gemmaverse showcase. Gemma 4 E4B orchestrates the analysis (terminal pipelines, data filtering, tool calls), running locally through Ollama on a single NVIDIA L4 GPU, while BOTANIC-1, a genomic foundation model pretrained on 320 plant species, evaluates the effects of mutations; proprietary genomes remain on-premises. The case study revisits a discovery made at INRAE by Adnane Boualem: changing a single letter in the melon’s CmEIN3 gene turns the flower from female to hermaphrodite. Among the 2 494 candidate point mutations, the validated mutation ranks first, with no ties, in less than four minutes according to the X Article.
| Tested configuration (20 runs) | Recall@1 |
|---|---|
| Without guidance | 0,00 |
| Expert guidance | 0,15 |
| With the BOTANIC-1 score | 0,90 |
The Botanic1 models (from 0,3 to 2 billion parameters) and the bioRxiv preprint date back to early September: the news is Google’s spotlight on them and the detailed results.
🔗 Google AI’s X Article 🔗 Google DeepMind’s Gemmaverse page
Advertising in ChatGPT: a visual format in testing and expanded measurement
October 5 — OpenAI is preparing a visual advertising format for ChatGPT: images will show the inspiration surrounding a product, its use, or the experience it makes possible. The first test will take place during image generation, with ads clearly labeled and separated from the image being created; it will begin later this month in the United States, with an initial group of advertisers. OpenAI reiterates that advertising does not influence ChatGPT’s answers, and that ChatGPT reaches 1,2 billion people each week according to the company.
Most of the announcements concern measurement: sending conversions through Hightouch, Tealium, and LiveRamp, ten named attribution partners (including AppsFlyer, Adjust, and Triple Whale), geographic incrementality experiments with Haus, Measured, and WorkMagic, a one-day view-through conversion window in reports, and a signal quality indicator (Event Quality Score). For brand safety, eligible advertisers can exclude phrases (Negative Phrases), and evaluation pilots are being prepared with DoubleVerify and Integral Ad Science. The pixel and Conversions API have been available since May 5.
| Result published by OpenAI | Published value | Measured by |
|---|---|---|
| WeightWatchers’ cost per acquisition compared with its paid search benchmark | −15,3 % | DV Rockerbox |
| Incremental purchases of the Dose brand made by new customers | 67 % | WorkMagic |
| Portland Leather visitors who came from ChatGPT Ads and were new to the brand | 93 % | Triple Whale |
🔗 OpenAI’s post on the new advertising format 🔗 Ads blog post on measurement
Perplexity’s October 5 changelog: browser extension for Computer, Wiley and Stripe Projects
October 5 — Perplexity publishes a product changelog with 18 sections, with no corresponding post on X at the time of our check. Eight revisit announcements already covered (Automations, Claude Opus 5.5, video generation, American Express Skills, Portable Computer on AMD, Fast Search, Agent API Profiles, Claude Fable 5.1), a ninth adds interactive 3D explanations to the October 1 charts, and nine items are new:
| New item | Platform or version | Audience |
|---|---|---|
| Extension for Chrome, Edge or Brave that lets Computer use the browser (Comet remains the default) | Perplexity for Mac 26.38.0 and later | Not specified |
| Computer tasks in separate tabs and windows | Perplexity for Mac 26.37.1 and later | Not specified |
| Effort mode (Light, Standard, High, Ultra) on mobile | iOS 26.38.0 and Android 2.100.0, and later | Pro, Max, Enterprise Pro, Enterprise Max |
| Wiley journals and books without a separate Wiley subscription | Perplexity and Computer | All plans, premium quotas vary |
| GPT-6.1 Sol, which also powers Computer’s Light effort level in place of GPT-6 Sol | Perplexity and Computer | Eligible paid subscribers |
| Guided conversation to launch a first Computer task | Web | New free accounts |
| Connecting an app from the input area (Connect an app) | Computer on the Web | Not specified |
| DocSend connector for investment data rooms (deal rooms) | Computer | Eligible DocSend account |
| Stripe Projects: Perplexity API key created from the Stripe CLI | Stripe CLI | Perplexity API developers |
With Stripe Projects, a command creates or links the project, generates the key and writes it to the .env file, a setup that Perplexity suggests handing over to Claude Code, Cursor or Codex.
🔗 Perplexity’s October 5 changelog
Generative creation: Suno Albums and HeyGen’s HyperFrames Studio
Suno adds albums
October 5 — Suno adds albums, previously absent from its platform. According to the accompanying announcement post, the feature turns users’ best tracks into complete projects (full-length projects), and a playlist (playlist) that served as an album can now be converted into an Album. Suno introduces the artists behind five initial albums: naenia (Oakwood///diskrot), Old Soul (Isaiah Wallace), phenomenal (KakerMix///diskrot), Lost In a Dream (Dream Relic) and VOID (ECHLO). The company specifies neither the eligible plans, the number of tracks per album, nor the exact launch date. In June, Suno was already asking its community about their listening experience across playlists, albums and radio stations.
🔗 Suno’s announcement on X 🔗 5 Albums You Can’t Miss, Suno blog
HyperFrames Studio, HeyGen’s video editor built for agents
October 5 — HyperFrames, HeyGen’s open source framework that turns HTML code into video, becomes a desktop application. HyperFrames Studio presents itself as a video editor built for agents from the ground up: users describe the video they want, the agent (Claude Code, Codex or a custom agent) delivers a first cut, then the human and agent work on the same timeline (timeline). Direction comes through gestures rather than prompts: circling an element in the image, pinning a comment to a word or making a cut yourself. The application exports in 4K and claims to offer hundreds of sources of inspiration and templates; it is free according to the HyperFrames account on X, and only a macOS version is available to download. It builds on July’s Studio Preview, which already allowed users to visually edit a video generated from code.
🔗 HyperFrames Studio announcement 🔗 HeyGen’s post
Briefs
- Warp Factories moves toward general availability — In an October 3 post about its fall retreat, Warp says Warp Factories, its agent infrastructure, is moving from early access to general availability, without a date or pricing; the team says using it reduced its cost per PR by 70% over the past month. 🔗 source
- Devin’s October 2 release notes — Microsoft Teams access settings are now enforced, one click selects all security findings of the same severity (including API v3), and suggested plugins show their skills, MCP, hooks and rules before installation; environment builds on large Perforce repositories now have 3 hours instead of failing after 10 minutes. 🔗 source
- TikTok Ads MCP in Replit — Replit adds TikTok Ads to its catalog of more than 450 integrations: once the account is connected, its Agent can create and manage TikTok ads, reach audiences and measure performance from the workspace, with no credits consumed for the connection. 🔗 source
- Cockle Finance on Replit — Replit pins a video about Cockle Finance, whose tool was built on Replit by Dan and his father Steve to track the influx of customers; according to Replit, Dan grew his business by 15% in three months, without specifying the metric. 🔗 source
- Copilot CLI 1.0.92 stable release — This version brings together the additions from prereleases 1.0.92-0 through -5; the only new addition is that, after signing in with Microsoft Entra, users choose which account to use,
/logoutcloses those OAuth sessions, and MCP servers protected by Entra refresh their credentials without intervention. 🔗 source - Codex CLI 0.160.1 — This patch for 0.160.0 contains just one backport: the Windows variables SYSTEMROOT, TEMP and TMP are preserved when launching remote stdio MCP servers whose remote environment is explicitly configured; it is the latest stable release, with no stable 0.161.0 released. 🔗 source
- Gemini CLI’s October 5 nightly — It contains no changes: built on the same commit as the October 3 nightly, it differs only in version numbers, while the stable (v0.62.0) and preview (v0.63.0-preview.0) channels remain unchanged. 🔗 source
- SpaceXAI TypeScript SDK 0.2.1 — Released on October 2, this version adds
toJson(schema), which validates structured output with a Standard Schema validator (Zod, Valibot or ArkType), and theretryBeforeOutputoption, which retries a streaming request (streaming) that failed before producing any output; a 429 without a Retry-After header now waits starting at one second, up to 30 seconds, instead of 250 milliseconds. 🔗 source - Cresta Conductor on the Claude Agent SDK — In Anthropic’s series about startups building with Claude, Cresta presents Conductor, a natural-language customer service agent builder built on the Claude Agent SDK; according to Cresta, it roughly halved initial deployment time in its first use cases, with no availability date or pricing provided. 🔗 source
- npm: expiration of trusted publishing configurations — Since October 2, a trusted publishing configuration (trusted publishing) that has not been used for any publication expires 48 hours after its creation, and npm also rejects tokens from GitHub Actions
issue_commentevents, as it does forpull_request_target. 🔗 source - npm: staged publishing creates packages — Since October 2,
npm stage publishcan create a package that does not yet exist, using a local session or a granular token, including one restricted to staging (stage-only); the first version awaits promotion by a maintainer before it can be installed. 🔗 source - Perplexity’s Agent API on Fast Search — The Agent API’s low, medium and high presets now use Fast Search for the
web_searchtool, dropping from $2.50 to $1 per 1,000 invocations and running about 800 ms faster; the xhigh preset retains standard search. 🔗 source - Perplexity’s RAG guide — Perplexity publishes a guide with nine examples of retrieval-augmented generation (retrieval-augmented generation, RAG) and seven architectures, citing itself for its web index and Instant Buy feature alongside third-party examples such as Zendesk or GitHub Copilot; no new product features. 🔗 source
- Cohere’s headcount — According to Cohere, the company has grown from around 400 employees at the start of 2026 to more than 800 today, a figure provided in a recruitment post. 🔗 source
- IsoVGRBench and Logbook at Meta — Without an announcement, facebookresearch publishes the code for IsoVGRBench, which evaluates video reward models on 16,000 pairs differing in only one aspect (when shown the same clip with its frames shuffled, these models respond almost at random), and for Logbook, which focuses on events in very long audio recordings. 🔗 source
- FiftyOne reads LeRobot v3 datasets — According to a community post by Harpreet Sahota, the open source toolkit FiftyOne now reads the LeRobot v3 format natively, without a converter or copying videos; the demonstration covers 497 episodes drawn from the 50 robot types in LeRobot’s community dataset. 🔗 source
- Ai2’s FetchMan-data — Published on October 1 without an announcement, this dataset contains 352,696 simulated episodes (52 million frames, 902 instructions) of object grasping by a Unitree G1 humanoid, generated in MolmoSpaces, in LeRobot v3 format and under the ODC-BY license. 🔗 source
- P(doom) on Hugging Face profiles — Hub users can display their P(doom) on their profile, their estimate of the probability that AI will cause an existential catastrophe; responses feed into a survey whose results will be published in two weeks, in aggregate and anonymized form only. 🔗 source
- hf-image extension — Hugging Face releases, without an announcement, a CLI extension that builds, pushes, pulls and runs container images on its cr.hf.co registry; uncompressed layers are deduplicated by Xet, so a 2 GB layer modified by 100 MB sends only about 100 MB according to the README. 🔗 source
- Three experiments by Milos Kotlar — In three community posts, Milos Kotlar makes a small transformer’s KV cache 2.71 times more compact with 97.85% agreement on the next token, and reduces the share of tokens with no expert in MoE routing from 13.84% to 8.09%, without any speed gain. 🔗 source
- Stephen Yu’s audit of an LLM judge — Stephen Yu audited, on 38,400 samples, the GLM-5.3 judge that scores factual fidelity in the GRPO reward for his 12B rewriting model: reliable at spotting that a fact has changed, noisy about its severity; counting faulty outputs reveals a 39% decrease that the initial count obscured. 🔗 source
- Sakana AI symposium in Tokyo — Sakana AI opens registration for a free symposium on October 26 at Hitotsubashi Hall in Tokyo, in English, with a talk by Jürgen Schmidhuber and an interview with David Ha, CEO of Sakana AI. 🔗 source
- NVIDIA Inception startups and breast cancer — NVIDIA presents four startups from its Inception program applying AI to the care pathway, including iSono Health and its FDA-authorized ATUSA 3D ultrasound scanner (about two minutes per breast compared with up to 45 minutes manually), Whiterabbit.ai, Ataraxis AI and SimBioSys; some of these technologies are not FDA-approved. 🔗 source
What this means
Text provenance is becoming a regulatory obligation. Until now, traceability for generated content mainly concerned images and audio, with watermarks and verifiable metadata; the AI Act extends the requirement to text, and OpenAI is responding in stages: watermarking applied by default to ChatGPT and Codex text in the European Union, an optional setting in the API, and a detector restricted to approved organizations. The published figures explain this caution: around 95% of 400-token passages are detected in psychology-related content, with a target false positive rate of 1%, but on 400-token passages in English, replacing 10% of words with synonyms reduces detection from around 92% to 66%, and rewriting or translation can erase the watermark. The watermark provides a clue to provenance, not proof, and its absence says nothing about the author.
Agent memory is taking hold and starting to become standardized. Devin keeps its lessons in a Git repository of Markdown files that Dreaming reorganizes each day, and Cognition publishes the format under the MIT license so other agents can adopt it; North 2 retains context between sessions; Qwen Code connects Mem0 directly to its CLI. The approaches differ, ranging from versioned files readable by humans to an external memory service and a feature built into a platform, but they address the same need: ensuring an agent does not start from scratch in every session. Cognition’s choice has a practical advantage: each note links back to the session in which it was learned, can be viewed in the interface and retains its Git history, making it possible to read and correct what the agent has retained.
Open models are entering coding tools through several routes. Together Link connects them to existing agents, including Claude Code and Codex, claiming to reduce spending by more than half; IBM runs the self-hosted version of Bob on Nemotron 3 Ultra for companies that keep development on their own infrastructure, including disconnected environments; Reflection AI presents Beam as an open model tailored for coding and agentic tasks, with 80.9 on SWE-bench Verified according to its own tables. The common argument has less to do with raw scores, where Reflection AI acknowledges that Kimi K3 remains ahead, than with cost, latency and control over data.
Many of today’s figures come from those publishing them. GitHub designed ReviewBench and produced the initial ranking in which Copilot code review comes out on top, without the other vendors running or verifying it; Nemotron 3 Ultra’s latency comes from IBM’s internal tests; Together Link’s savings and ChatGPT’s advertising results are those reported by Together AI and OpenAI. GitHub does, however, publish its dataset, judge and methodology, and opens submissions: each vendor can repeat the measurement with its own agent. That is what will establish whether this initial ranking holds up.
Sources
- Our approach to EU text provenance rules (OpenAI)
- OpenAI thread on watermarking limitations
- Memory and dreaming: how Devin learns from working with you (Cognition)
- Cognition thread on X
- Agent Memory Repo
- Cursor thread on the SDK
- Cursor SDK Changelog
- Qwen Code v0.25.0 on GitHub
- Together Link (Together AI)
- Together Link documentation
- Why IBM chose NVIDIA Nemotron for IBM Bob self-hosted
- ReviewBench: An open benchmark for AI code review (GitHub Blog)
- ReviewBench website
- North 2: Enterprise AI Without Compromises (Cohere)
- North 2 launch thread (Cohere on X)
- Cohere and PwC partnership (Cohere blog)
- PwC Canada press release
- Introducing Beam (Reflection AI)
- MetaEncoder-30B on Hugging Face
- Google AI X Article on Living Models
- Living Models on Gemmaverse (Google DeepMind)
- Building advertising for the way people use AI (OpenAI)
- More ways to measure ChatGPT Ads
- Perplexity Changelog for October 5
- Suno’s album announcement
- 5 Albums You Can’t Miss (Suno)
- HyperFrames Studio announcement
- HeyGen repost
- Build The Factory. Let It Run. Go Kayaking. (Warp)
- Devin release notes for October 2
- TikTok Ads MCP in Replit
- Cockle Finance on Replit
- Copilot CLI v1.0.92
- Codex CLI 0.160.1
- Gemini CLI Nightly for October 5
- SpaceXAI TypeScript SDK v0.2.1
- Cresta and the Claude Agent SDK (claude.com blog)
- Unvalidated npm trusted publishing configurations now expire
- npm staged publishing now supports creating new packages
- Perplexity API Changelog
- 9 RAG Examples (Perplexity blog)
- Cohere headcount
- IsoVGRBench repository
- FiftyOne and LeRobot v3 (Harpreet Sahota)
- FetchMan-data on Hugging Face
- P(doom) on the Hub (Hugging Face Changelog)
- hf-image repository
- KV cache under a KL budget (Milos Kotlar)
- LLM judge audit (Stephen Yu)
- Sakana AI symposium
- From Scan to Treatment Plan (NVIDIA)