Search

Aleph Alpha releases Kolibri under Apache 2.0, Meta expands its safety framework, OpenAI adds three sandbox sizes to its Agents API

ai-powered-markdown-translator

Article translated from French to English with gpt-6.1-sol.

View project on GitHub ↗

On Saturday, October 3, Aleph Alpha releases Kolibri, a 78.1-billion-parameter English-German open-weight model under the Apache 2.0 license. Cohere shares the release that same evening, while its combination with Aleph Alpha still awaits regulatory approval. The previous day, Meta expanded its safety framework to cover model training, and OpenAI’s Agents API added three sandbox sizes. Devin’s September 30 release notes make it a native Jira agent, Qwen-Image-2.1-Pro arrives via API at 0.04 dollars per image, and Apron’s developer explains how to predict an open model’s GPU memory requirements before renting a card.


Aleph Alpha releases Kolibri, a 78-billion-parameter English-German model under Apache 2.0

October 3 — On German Unity Day, Aleph Alpha releases Kolibri, an English-German open-weight language model. This Mixture-of-Experts model has 78.1 billion parameters, with 3.46 billion active per token, and its weights are available on Hugging Face under the Apache 2.0 license. Aleph Alpha targets sovereign workloads in regulated sectors: public administration, industry, and aerospace.

The architecture primarily targets serving costs: 6 active experts out of 384, and 40 of its 50 layers limited to a sliding window of 512 tokens. The model accepts up to 1 048 576 tokens, but it was trained up to 262 144, a length its model card recommends not exceeding. According to the post, a 123-billion-parameter version handled only 3 requests of 256k tokens on two H100s, compared with 18 for the selected size. Training used 768 B200 GPUs in Germany and Finland, processing nearly 24T tokens; German accounts for 21.3% of pretraining.

According to Aleph Alpha’s measurements (in-house harness, maximum reasoning effort), Kolibri outperforms Qwen3.5 35B-A3B and Nemotron 3 Super, which activates 12 billion parameters, in overall score, but the dense Qwen3.8 27B model remains ahead almost everywhere:

Benchmark (Aleph Alpha measurements)Kolibri 78B-A3,46BQwen3.5 35B-A3BNemotron 3 Super 120B-A12BMistral Small 4 119B-A6BQwen3.8 27B (dense)
Overall score (English)75,574,773,063,180,2
Overall score (German)70,869,867,961,479,9
GPQA Diamond (English)84,383,878,074,789,2
AIME 2026 (English)96,092,190,483,197,7
SWE-Bench Verified66,471,660,260,872,6

Aleph Alpha’s argument is therefore about the Pareto frontier between quality and serving costs, rather than first place. Trained to abstain, Kolibri also prefers to abstain or answer partially rather than get things wrong on 44% of the AA-Omniscience questions it fails to answer correctly. Developed in Germany with “no foreign control” (no foreign control), and designed with the AI Act and GDPR in mind, it can run on premises; a technical report of nearly 200 pages accompanies the release.

Cohere, whose definitive agreement to combine with Aleph Alpha remains subject to regulatory approval, shared the release that same evening, following congratulations from its CEO Aidan Gomez; Kolibri remains an Aleph Alpha model.

New Apache 2.0 model from @Aleph__Alpha. It’s really good. If you like that, you’re going to love what we build together. — @cohere on X

🔗 Aleph Alpha blog post · Weights on Hugging Face


Meta expands its safety framework to training and clarifies its rules for open weights

October 2 — Meta Superintelligence Labs updates its Meta Superintelligence Scaling Framework, which sets its safety requirements before training and then deployment. According to Meta, this version reflects commitments made this week at the White House, providing for internal controls verified by an independent team within the company and by an external auditor or evaluator.

Loss of control (loss of control) is now addressed during training and evaluation, because a model with strong cybersecurity capabilities can, according to Meta, exploit vulnerabilities in its environment. High-risk reinforcement learning training requires validated sandboxes, tamper-proof logs, including chain of thought, and automated monitoring capable of stopping the run.

For open weights, the framework accounts for possible modifications after release, such as removing refusals through fine-tuning. A board AI committee, announced for the coming months, will independently verify compliance with the framework. This framework is the former “Advanced AI Scaling Framework,” renamed with this version 2.1: the one under which Muse Spark was evaluated in April.

🔗 Developing Capable Models Responsibly (Meta)


OpenAI’s Agents API adds three sandbox sizes and environment reuse

October 2 — Steve (@stevendcoffey), who works on OpenAI’s API according to his X bio, lists the week’s updates to the Agents API, shared by @OpenAIDevs. Alongside three reminders from DevDay, four items are new. The first is confirmed by the documentation: the environment.container_size parameter, set when a session is created, selects the CPU and memory for the sandbox hosted by OpenAI.

Container sizeAllocated vCPUsAllocated memory
small11 Go
medium (default)24 Go
large416 Go

The other three appear only in the tweet: subagents configurable from the dashboard, reuse of an environment across multiple sessions, and improved reliability, with no measurement method published.

Overall better reliability: 99.97% turn reliability, fewer SSE disconnects, 20% faster tool calls 🪨 — @stevendcoffey on X

🔗 OpenAI documentation


Coding agents: Devin in Jira, Antigravity CLI 1.2.13 and 1.2.14

Devin becomes a native Jira agent and passes Devin Review findings to its sessions

September 30 — Devin’s September 30 release notes contain 25 sections. The main one makes Devin a native agent (native agent) in Jira: once an administrator installs the Devin for Jira app, assigning an issue to Devin or mentioning it launches a session, tracked in Jira’s agent panel. If Devin cannot start (missing permission, usage limit), Jira displays its explanation instead of a generic error.

On a pull request opened by a Devin session that is still active, Devin Review first passes its findings (findings) to that session: with automatic correction (Auto-fix), Devin fixes what it can, and only the remaining findings are published. Automations gain preliminary checks (preflight checks): a script runs after each trigger, before Devin, to validate, enrich, or skip the event. No pricing is mentioned.

🔗 Devin’s September 30 release notes

Antigravity CLI 1.2.13 and 1.2.14: retry delays, message queue, Remote Control without systemd

September 29 and 30 — Google’s Antigravity CLI changelog lists two versions. Version 1.2.13 follows the retry delay specified by the model’s API instead of waiting a fixed 5 seconds, and gives up immediately if that delay exceeds 30 seconds or if the limit reached is a daily or billing cap.

Version 1.2.14 (6 improvements, 3 fixes) adds the Queued Messages option in /config: by default (Queue), a message sent while the agent is working waits until the end of the turn; in Send Immediately mode, it interrupts the agent. On Linux machines without a systemd user service manager, such as most containers, Remote Control launches its daemon as a simple background process, which will restart neither after a crash nor at boot. The --json-schema option becomes strict: an invalid schema causes an error and exit code 1. The rollout is gradual, over a few days.

🔗 Antigravity changelog


Qwen-Image-2.1-Pro arrives via API on Model Studio and QwenCloud, at 0.04 dollars per image

October 2 — Alibaba adds Qwen-Image-2.1-Pro to the catalog of Model Studio, its API platform, for the International service scope, and QwenCloud gives it a listing marked “NEW,” with no tweet or blog post from Qwen. The model extends Qwen-Image-2.1, released with open weights on September 20: creation and editing in a single model, 7 billion parameters for visual generation, and native transparent images. The listing does not say whether the served version differs from the published weights.

Compared itemqwen-image-2.1-proqwen-image-2.0-pro
Price per output image0,04 dollar0,075 dollar
Free quota10 images100 images

The price drops by nearly half, but the free quota is ten times smaller. On QwenCloud, the model is capped at 20 requests per minute and 10 concurrent calls.

🔗 Model Studio model changelog · QwenCloud listing


Apron predicts GPU memory for 14 open models before rental, according to its author

October 3 — Vlad Ryzhkov, Apron’s developer, presents this open source tool on Hugging Face’s community blog. Before a GPU is rented, it predicts whether an open model will fit in memory and start under a specific version of vLLM, then verifies this on an actual card.

According to the author, predicted weight memory falls within 2% of the measurement for 11 of the 14 models started, on hardware ranging from an RTX 4090 to eight H200s. Four models failed before a fix, validated by a second startup, repaired them, and DeepSeek-V4.1-Flash uses 189 Gio of host memory by default, invisible to a check limited to the GPU.

The author cautions that there is only one measured startup per model and that none of this constitutes a ranking. The command-line tool is not ready for outside use, and its repository still describes itself as a specification without an implementation.

🔗 Vlad Ryzhkov’s blog post (Hugging Face)


Briefs

  • Copilot CLI 1.0.92-3 and Copilot SDK 1.0.17-preview.3 — The Copilot CLI preview adds a selector, opened with Ctrl+E before the conversation, to choose local or cloud execution; the SDK preview experimentally allows client tools to be replaced in an ongoing session. The latest stable version remains 1.0.91. 🔗 CLI · 🔗 SDK
  • Copilot code review via API — Copilot reviews can now be requested through the REST and GraphQL APIs, with an effort level set for each request, generally available for Pro, Pro+, Max, Business and Enterprise plans; the documentation estimates the cost of a review at $0.05 to $1 in IA credits for Lite and $0.25 to $5 for Balanced, the default effort level. 🔗 source
  • Gemini CLI nightly — The October 3 v0.64.0-nightly contains just one fix: Enter and Space reliably confirm selections in choice lists, including on terminals without the Kitty keyboard protocol, such as PyCharm’s terminal on Windows, where pressing Enter less than 30 ms after another key was interpreted as Shift+Enter. The stable and preview versions are unchanged. 🔗 source
  • DeepSeek Harness 0.2.1-alpha.1 — This 25th preview, still without a stable version, adds an experimental compatibility layer for Claude Code mods, intended, according to DeepSeek, to verify that their API is broadly a subset of what Harness plugins support, rather than to offer full compatibility. 🔗 source
  • GPT-6.1 Sol in Warp — On the evening of September 29, Warp made GPT-6.1 Sol available in its agentic terminal, usable with a Codex or ChatGPT subscription; the announcement provides neither pricing nor measurements specific to Warp. 🔗 source
  • End of grok-voice-transcribe-1.0 — On October 2, SpaceXAI retired the old version of its transcription model in the API: requests are redirected to grok-voice-transcribe-2.0, at the same price and with greater accuracy according to SpaceXAI. The deprecation had been announced on September 18 without a date. 🔗 source
  • September for OpenAI Developers — The @OpenAIDevs account recaps September’s developer announcements in an X Article, from DevDay on the 29th to GPT-6 Sol and Luna, without any new announcements; the only additional detail is that OpenAI’s Decisions API, in limited preview since September 29, is said to be coming to general availability soon, without a date. 🔗 source
  • Two ElevenLabs agent deployments — Banner Health is entrusting primary care appointment scheduling to ElevenAgents, with human handoff and HIPAA compliance; insurer Admiral describes putting its voice agents into production, with each change moving from 1% to 100% of traffic in a few hours instead of several weeks. Neither publishes any outcome figures. 🔗 Banner Health · 🔗 Admiral
  • Kling 4.0 in early access — The introduction to Kling 4.0, dated September 30, specifies that the full model is entering early access ahead of a broader rollout in October, while Kling 4.0 Flash has been available to annual Ultra subscribers since September 28; the 21:9 format is added, with no pricing, API or access criteria provided. 🔗 source
  • Runway Agent, Runway MCP and OpenAI’s dots — Runway’s changelog lists three additions without an announcement on X: Runway Agent uses Seedance 2.5’s Draft mode (480p drafts refined after generation), Ideogram 4.5 joins Runway MCP for paid plans, and Runway has been available in OpenAI’s dots since October 1. No credit costs are specified. 🔗 source
  • Stateless GitHub App tokens — Outside IA, GitHub is completing the rollout, begun on April 27, of the stateless ghs_APPID_JWT format for all new GitHub App installation tokens: around 520 characters instead of 40, with permissions and the one-hour expiration unchanged; the test header X-GitHub-Stateless-S2S-Token will be deprecated on November 30, 2026. 🔗 source
  • Live human feedback at Rapidata — Rapidata describes its Flows feature, which replaces Flow-GRPO’s reward model with real-time human pairwise comparisons, in a vendor post that publishes no measured training results. 🔗 source
  • Measuring dependency between agents — In an essay without experiments or measurements, Anouar Imel proposes evaluating multi-agent systems based on dependency between agents rather than their number, tracking accuracy, diversity of reasoning and coupling between agents at each turn. 🔗 source

What it means

Kolibri shows that an open model can hold its own without aiming for first place. Aleph Alpha itself publishes measurements in which the dense Qwen3.8 27B outperforms it almost everywhere, and highlights another criterion: serving many long requests on few GPUs, in both English and German, under a license that allows everything to be hosted on premises. For a public administration or industrial company bound by the AI Act and the RGPD, this tradeoff may matter more than a few benchmark points. Cohere, whose deal with Aleph Alpha is still awaiting regulatory approvals, is already promising more to come.

Serving costs are becoming a central issue for open models. Qwen-Image-2.1, released with open weights on September 20, returns through an API under the name Qwen-Image-2.1-Pro at $0.04 per image, nearly half the price of the previous generation, for those who prefer not to host it themselves. Apron’s post, meanwhile, highlights what self-hosting requires: a model can fail at startup because of a missing vLLM setting, or occupy 189 Gio of host memory that a check limited to the GPU will miss.

Meta is moving safety checks to before deployment. Risky reinforcement learning now requires validated sandboxes, tamper-proof logs and automatic shutdown, because a model with strong cybersecurity capabilities can exploit vulnerabilities in its own environment. The IA committee announced within the board of directors must also independently verify that these rules are being applied. As for the commitments made at the White House, which Meta says it reflects, the post conveys only their general intent: internal controls verified by an independent team within the company and by an external auditor or evaluator.

Finally, agents are becoming integrated into tools and operations. OpenAI lets users choose the sandbox size and reuse an environment across sessions, Devin integrates into Jira and addresses the findings from its own review before publishing them, and Copilot code review can be triggered through an API with an estimated cost for each effort level. Antigravity CLI respects the server’s retry delays and runs Remote Control in containers: operational settings more than spectacular features, but these are what matter when an agent runs continuously.


Sources