ai-powered-markdown-translatorArticle translated from fr to en with gpt-5.6-sol.
No new flagship model this Saturday, but tangible changes for those already using these tools. Microsoft is rolling out SpaceXAI’s Grok models in Copilot for Word, Excel, and PowerPoint the day after Grok Bot arrived in Teams, OpenAI is cutting the price of ChatGPT licenses for all US government agencies to zero while calling for federal legislation, and DeepSeek’s documentation has canceled the planned September 14 redirection of V4 Pro. On the developer tools side, Amp is becoming free for users who bring their own compute and models, while Gemini CLI and Vibe CLI are tightening their security controls.
Grok joins the Microsoft suite, from Copilot for Word, Excel, and PowerPoint to Teams
September 12 — Microsoft announces the rollout of SpaceXAI’s Grok models in Copilot for Word, Excel, and PowerPoint. The launch is targeted: it begins exclusively with customers in the Microsoft Frontier program, and Microsoft presents it as an expansion of the model choices available in Copilot. The @grok account shared the announcement less than an hour later.
Launching even more models in Copilot. Grok models from SpaceXAI are rolling out in Copilot in Word, Excel, and PowerPoint—starting with a focused release to customers in the Microsoft Frontier program. — @Microsoft365 on X
The distinction matters: SpaceXAI had already offered its own Grok add-ins for Word, Excel, PowerPoint, and Outlook since June and July, and Grok 4.6 has been available on Microsoft Foundry since late August. This time, Grok models are being integrated into Microsoft’s Copilot. The announcement specifies neither the Grok version involved, nor pricing, nor a timeline beyond the Frontier program.
🔗 Grok’s repost of the announcement on X
Grok Bot searches and acts in Microsoft Teams
September 11 — The day before Microsoft’s announcement, the @bot account said that Grok Bot, SpaceXAI’s persistent agent offering, can now search and act in Microsoft Teams on behalf of the user. A post published late in the afternoon on September 10, Pacific time, was already aimed at sales teams: Grok Bot connects to Salesforce, HubSpot, Gong, Clay, and Granola to track customer accounts, handle follow-ups, and conduct in-depth research. These connectors extend the features introduced in early September, from Grok Bot for Enterprise to form filling and message drafting. Neither post specifies a required plan, pricing, or usage limit.
🔗 Grok Bot’s Microsoft Teams announcement on X
OpenAI and the US government, between public procurement and a call for legislation
September 9 and 10 — Two OpenAI texts published in the middle of the week, which our previous editions did not cover, show the company engaging with US public authorities in two ways: it sells to the government, and it asks lawmakers to regulate it.
OneGov 2.0, a zero-dollar ChatGPT license for all government agencies
September 10 — OpenAI for Government signs a new multiyear agreement with the federal General Services Administration (GSA), continuing last year’s federal offering. The license, normally priced at 15 dollars per user per month, drops to zero with no minimum commitment, while usage is billed at half price. The offer now extends beyond the federal level: states, local governments, and tribal governments are eligible for the first time. More than one million public-sector employees already had access to ChatGPT through existing agreements; eligibility is expanding to approximately 23 million people, and GPT-6 Astra is among the tools offered.
| Agreement item | Announced terms |
|---|---|
| Per-user license | 0 dollars instead of 15 dollars per month, with no minimum commitment |
| Model usage | 50% discount |
| Eligible governments | Federal government, states, local governments, and tribal governments |
| Cyberdefense component | Daybreak Blue at 50% of the commercial rate, Daybreak Red on request at standard rate |
| Agreement term | 27 months, from October 1, 2026, to December 31, 2028 |
The cyberdefense component expands Daybreak for Frontline Defenders, a one-billion-dollar initiative announced the previous week. OpenAI reiterates that ChatGPT Enterprise does not use organizations’ data, whether inputs or outputs, to train its models. An introductory webinar is scheduled for September 14.
🔗 Expanding AI access and cyber defense for federal, state, local, and tribal governments
An op-ed calling for federal legislation and four California bills
September 9 — In an op-ed arguing that the political window is open, Chris Lehane, OpenAI’s Chief Global Affairs Officer, calls on Congress to adopt mandatory nationwide, capability-based regulation before the end of its session. These obligations should target the handful of labs developing the most capable systems, not startups or small developers, and safety policy must not become a policy against open-weight models. In the meantime, OpenAI is formally supporting four bills passed by the California legislature and sent to Governor Newsom. SB 1119 is among them: OpenAI had already announced its support for the bill in late August.
| California bill | Purpose of the bill |
|---|---|
| SB 813 | Framework for independent safety evaluations |
| AB 1405 | Standards applicable to AI auditors |
| SB 1119 | Protection of minors who use companion chatbots |
| AB 1864 | Safeguards against AI-enabled biological threats |
Some of these bills we did not endorse in the past, and are now supporting after reconsidering in light of the recent jump in capabilities we have seen. — Chris Lehane, Chief Global Affairs Officer, OpenAI
The op-ed also calls for industry standards to monitor misalignment: prompt written notification to affected organizations when a model, during development or evaluation, bypasses their security controls without authorization. For Astra, OpenAI says it monitors complete trajectories, including chains of thought, and requires an alignment evaluation before any broader internal deployment. According to the text, fully autonomous recursive self-improvement is not happening today and should not be pursued until it can be done safely.
🔗 The AI policy window is open. We need to act.
DeepSeek ultimately keeps its V4 Pro API after September 14
Observed on September 12 — DeepSeek’s API documentation now says the opposite of its September 10 announcement. That day, with the release of V4.1-Flash, DeepSeek scheduled the end of V4 Pro: starting September 14 at 04:00 UTC, all requests to deepseek-v4-pro were to be redirected to V4.1-Flash and billed at the Flash rate until the release of V4.1-Pro. A note on the Models & Pricing page now says that DeepSeek will continue providing the V4 Pro API after September 14 to meet user demand, with billing unchanged, and that the company will provide notice if anything changes again.
The same paragraph appears in the September 10 entry of the Change Log, in both English and Chinese. DeepSeek does not date this reversal and has not mentioned it on X, while the September 10 announcement post still displays the old schedule: two contradictory versions remain online.
| Billing item per million tokens | deepseek-flash, off-peak hours | deepseek-flash, peak hours | deepseek-v4-pro, off-peak hours | deepseek-v4-pro, peak hours |
|---|---|---|---|---|
| Cached input (cache hit) | 0.003 dollars | 0.006 dollars | 0.022 dollars | 0.044 dollars |
| Uncached input (cache miss) | 0.15 dollars | 0.30 dollars | 0.66 dollars | 1.32 dollars |
| Output tokens | 0.60 dollars | 1.20 dollars | 1.98 dollars | 3.96 dollars |
| Concurrent requests | 2,500 | 2,500 | 500 | 500 |
| Image understanding | yes | yes | no | no |
For API users, the automatic redirection will therefore not take place: the two models will coexist, each with a context window of one million tokens and a maximum output of 384K. Depending on the billing item, V4 Pro costs 3.3 to 7.3 times as much as Flash. Peak hours run from 01:00 to 04:00 and from 06:00 to 10:00 UTC, Monday through Friday, with all other times billed at half price.
🔗 Models & Pricing, DeepSeek API 🔗 DeepSeek API Change Log
Amp becomes free for users who bring their own compute and model subscriptions
September 12 — Amp has published a post titled Free Agent, dated September 13 on its website, that changes the billing model for its coding agent. Usage becomes free when users bring their own compute and their own model subscriptions or keys. Connecting a ChatGPT subscription no longer requires a monthly plan, and bringing personal API keys (Bring Your Own Key, BYOK) no longer incurs per-token fees or caps for any non-Enterprise customer. Amp now charges for orbs, its remote machines where agents work autonomously and in parallel, while runners installed on users’ own machines provide a free alternative. Inference remains available for purchase through Amp, with no added markup.
| Amp plan | Listed price | What the plan includes |
|---|---|---|
| Hobby (new) | Free | All features, usage-based paid orbs or free personal runners, BYOK with no fees or caps |
| Individual (Megawatt or Gigawatt) | 20 dollars per month for Megawatt | 45,000 minutes of orb time, unlimited public and private repositories |
| Teams (new) | No extra charge | Members on different plans, shared threads and portals, pooled credits, SAML/OIDC SSO |
| Enterprise | Custom pricing | Pooled credits with no paid seats, SCIM, spending caps, audit logs |
For teams, the entry cost disappears: anyone can be invited to a workspace for free, while heavy users can upgrade to a paid plan for usage discounts, averaging 60% for Megawatt and 65% for Gigawatt according to Amp. SAML/OIDC single sign-on becomes free for the entire workspace as soon as one member pays, and the zero- or minimal-data-retention policy, previously contractually guaranteed only to Enterprise customers, is extended to everyone. Finally, Megawatt and Gigawatt gain early access to eight additional providers, including OpenRouter, Amazon Bedrock, and Ollama Cloud.
ChatGPT desktop 26.908, Quick Chat from Pets and Appshots on Windows
September 11 — OpenAI releases version 26.908 of the ChatGPT desktop app, focused on Pets. These optional animated companions, which float above other windows on macOS and Windows, were previously used to track ongoing conversations. They now also serve as a starting point: from the controls beneath the pet, users can type a Quick Chat and send it with Enter without opening the main window. A global shortcut displays these controls—Option+Space on macOS and Windows+Alt+P on Windows—and Mini mode retains the same shortcuts and notifications, without an animal on screen.
The pet indicates four states: conversation in progress, awaiting a decision, completed, or blocked by an error, and prioritizes threads awaiting a response. Conversations launched from these controls are created outside any project, and custom pets remain stored on the device, without syncing to ChatGPT on the web.
Appshots, introduced on macOS in late August, are also coming to Windows: pressing both Alt keys simultaneously sends the foreground window to ChatGPT as a screenshot accompanied by any available text, into the conversation of the user’s choice. The shortcut can be customized, and on Windows the appshot opens in the main application.
🔗 ChatGPT announcement on X 🔗 ChatGPT app changelog
Gemini CLI, a nightly that isolates the sandbox and monitors build files
September 12 — After two nightlies with no code changes on September 10 and 11, Gemini CLI’s v0.61.0-nightly.20260912 brings two changes, both focused on security, continuing the nightlies of September 2, 4, and 5. The stable (v0.59.0) and preview (v0.60.0-preview.0) channels remain unchanged.
The first (PR #29250) targets indirect prompt injection through build files: malicious external content pushes the agent to modify a Makefile or a package.json, then run the build. The CLI now tracks build configuration files created or modified during the session, and any subsequent build or test command (npm run, make, cargo, blaze) requires explicit confirmation. The same rule applies when a command uses arguments taken from untrusted external content: Google Docs, Buganizer, pages retrieved through web fetch, or responses from MCP servers.
The second (PR #29214, 45 commits) separates the sandbox from the user’s configuration and credentials. Gemini CLI refuses to launch it from the home directory, the system root, or ~/.gemini, and refuses to mount these paths in a container, which now receives only a sanitized copy of the settings, without hooks or API keys. On macOS, Seatbelt profiles prohibit writing to ~/.gemini and sensitive files (oauth_creds.json, trustedFolders.json, .env), as well as reading credential stores and hook definitions.
🔗 v0.61.0-nightly.20260912 nightly notes
Vibe CLI 2.25.4 requires approval for risky shell syntax
September 12 — Mistral releases Vibe CLI 2.25.4, its fourth version in four days. This corrective release primarily addresses security: shell permission checks now require user approval for risky command syntax and options that could bypass workspace controls and the denylist. The release note links this hardening to four identifiers, from CVE-2026-87984 through CVE-2026-87987, and to residual variants of CVE-2026-87988. No security advisory appeared in the repository on September 12, and the severity of these vulnerabilities is not specified.
Smart approve, which version 2.25.1 already prevented from automatically approving git commands that might display secrets, changes behavior again. When a risky action is requested again, it no longer simply blocks it: it asks for confirmation and displays the reason. Calls it approves automatically now appear as a simple note rather than a warning, and the mode appears in the Vibe Desktop selector when enabled by its gradual rollout.
The remaining changes improve everyday reliability: an invalid connector no longer prevents the others from loading, file mentions using @ work across all roots added through --add-dir, and ACP clients such as Zed and JetBrains always display the active mode.
🔗 2.25.4 release notes on GitHub
Wringer, a 4-billion-parameter reasoning model at 2.655 bits per weight
September 12 — weiciao wu publishes Wringer on the Hugging Face community blog, a quantization method tested on InternScience/Agents-A1-4B, a 4-billion-parameter reasoning model with a hybrid Qwen3.5 architecture. It is an individual contribution, not a laboratory publication. The model body is reduced to 2.655 bits per weight and retains an average of 93.7% of its original scores on IFEval, HumanEval, and GSM8K, while the curve for GGUF quantizations of the same model falls, by interpolation, to 52.7% at the same bitrate.
The recipe has three stages. Closed-form quantization, using GPTQ on the full covariance matrix, already provides 81.4% retention. Fill then freezes the codes and adds a rank-128 adapter to each linear layer, trained through self-distillation on about 100 million tokens, taking 7 hours and 30 minutes on a single GPU. Wring finally removes the adapter by recalculating the codes to absorb its contribution: the same bit budget, with no adapter at inference time.
| Tested configuration | Body bits per weight | IFEval score | HumanEval score | GSM8K score | Average retention |
|---|---|---|---|---|---|
| bf16, reference | 16 | 92,79 | 94,51 | 95,53 | 100 % |
GGUF, IQ2_M class | 2,763 | 73,38 | 63,41 | 76,65 | 75,47 % |
| Wringer without training | 2,655 | 75,97 | 70,73 | 83,70 | 81,44 % |
| Wringer, two cycles (published model) | 2,655 | 88,17 | 85,37 | 93,48 | 94,65 % |
GGUF, IQ2_XXS class | 2,558 | 34,38 | 20,12 | 36,54 | 32,16 % |
The 93.7% figure corresponds to the initial evaluation of a single cycle; the published model on the Hub, with two cycles, reaches 94.65%. The limitations are clearly stated: the 2.655 bits cover only the body’s linear weights, and the entire model weighs 4.69 bits per weight with the embedding retained in bf16. Inference reverts to bf16 in vLLM due to the lack of native low-bit kernels, so no memory or speed gain has yet been demonstrated at runtime, and the evaluation is limited to one model, three benchmarks, and a comparison with GGUF. The project took about five weeks on a single GPU, with an agent running the experiment chains while a person selected the directions to pursue.
🔗 Wringer on the Hugging Face blog
In brief
- Cognition and Perplexity entrust end-to-end testing to GPT-6 Astra — two API case studies. Devin tests the iPhone game Otter Run and returns a recording from a simulator, along with a report distinguishing successful checks from untested areas; Perplexity has Astra build a small program that mimics responses from other services to test an application end to end. 🔗 source
- OpenAI Developers brings together thirteen projects built with GPT-6 Astra — among them, 2,234 anatomical parts to explore in 3D, Manhattan recreated in Unreal Engine, an old train drawing turned into a Blender model containing 3,295 editable objects, and a ray tracer written in C++ and ported to Swift on the iPhone GPU. 🔗 source
- autofinetune entrusts dozens of Gemma fine-tuning experiments to an agent — on the Google Developers Blog, Antigravity CLI and Gemini 3.7 Flash modify the training script, run Tunix on Cloud TPU, and retain only the winning commits. Results: 20 experiments in a few hours on FunctionGemma 270M, followed by 40 GRPO experiments over two to three days on Gemma 3 1B, for about 10% more reward. 🔗 source
- ToolGrad generates tool-use data starting from the answer — a catch-up from September 10: Google Research first constructs a valid chain of API calls, then writes the request, with a success rate close to 100% across ToolBench’s more than 16,000 APIs. A Gemma 3 12B fine-tuned on this data scores 83.1 on BFCL, compared with 83.2 for gemini-2.5-pro and 74.4 for GPT-5, the leading models at the time of the article. 🔗 source
- Gemini API documentation arrives in Google AI Studio — a catch-up from September 10: an initial preview, rebuilt from scratch and available at ai.studio/docs, which Logan Kilpatrick describes as designed for humans and agents alike. The messages do not explain how agents use it. 🔗 source
- Three Google-backed Android XR projects at the Venice Film Festival — the 100 ZEROS team supported Andy Serkis’s NEVATARS and Galápagos: The Last Eden, narrated by Margot Robbie, where Gemini powers conversations with the viewer, as well as the trailer for Sedona, a psychological thriller from Asylm Studios, automatically converted from 2D to 3D. The post provides neither a specific model nor an availability date. 🔗 source
- Claude Code 2.1.270 fixes a regression from 2.1.269 — read-only git commands run by the Bash tool no longer ask for authorization again after a certain amount of session time. This is the version’s only change. 🔗 source
- Smart reports in beta on Claude Enterprise — a catch-up from September 10: Claude analyzes a sample of a team’s transcripts covering up to 28 days and writes a report on the work completed, its cost, friction points, and repeated tasks that could be turned into shared skills. Ten free reports per month during the beta, with the feature disabled by default. 🔗 source
- Copilot metrics measure the VS Code Agents window — now generally available, new fields count daily active users, sessions, and messages in this window over 1- and 28-day periods, by enterprise, organization, or user. Access is restricted to owners, billing managers, and roles with the View Copilot Metrics permission. 🔗 source
- GitHub relaunches ticket sales for Universe 2026 — a short video highlights the programmable badge, presented as a mini-computer and reserved for in-person attendees: an in-person pass costs $1,399 for October 28 and 29 in San Francisco, while virtual attendance is free. A promotion with no product announcement, following similar pushes on August 22 and September 7. 🔗 source
- Fugu Ultra v2 arrives on OpenRouter — Sakana AI’s high-end orchestrator costs $5 per million input tokens and $30 per million output tokens there, with a context window of one million tokens. Prompts longer than 272,000 tokens are billed at a higher rate, and tokens consumed by the orchestration itself are billed like ordinary tokens. 🔗 source
- ChordStream, one 64-bit token per chord for symbolic music models — specification published on the Hugging Face community blog: each token carries the chord’s key, degree, bass, and position, extracted from MIDI files or REMI streams, alongside the captioned-ChordStream dataset. The extractor remains heuristic, and no trained model is presented. 🔗 source
- Qwen3.8-27B runs on Cerebras at about 1,850 tokens per second — a 64k-token context with free access and 128k with paid access, priced at $0.99 per million input tokens and $1.49 per million output tokens. Cerebras gives it a score of 34 on the Artificial Analysis Intelligence Index. 🔗 source
What it means
The day’s first thread is distribution. Grok models are entering Copilot for Word, Excel, and PowerPoint through Microsoft, while SpaceXAI already offered its own add-ins for these applications, and Grok Bot arrived in Teams the day before, after connecting to business tools such as Salesforce and HubSpot. OpenAI is targeting the same territory with ChatGPT for desktop, where a conversation starts from a floating companion and Windows can now send a window to ChatGPT with two keystrokes. Microsoft, moreover, presents Grok’s arrival as an expansion of model choice in Copilot: the model becomes an option within the tool that is already open, rather than merely a separate application users must seek out.
The second thread is political. In two days, OpenAI published an agreement that reduces the ChatGPT license cost to zero for about 23 million U.S. public-sector employees, from the federal government to tribal governments, and an op-ed calling on Congress to pass a mandatory law based on model capabilities. The two texts reinforce each other: the company becomes a nationwide government supplier and proposes placing obligations on the handful of laboratories developing the most capable systems, rather than on startups or open weights. Its support for four California bills, including some it had not previously backed, fits within what the op-ed calls reverse federalism: state laws converging while awaiting action at the federal level.
The third thread concerns the price of usage. Amp stops charging for tokens users bring themselves, resells inference without a markup, and shifts its billing to the remote machines where agents run. DeepSeek, which intended to withdraw V4 Pro in favor of a cheaper V4.1-Flash, is keeping it at the same price in response to demand, even though it costs 3.3 to 7.3 times more: some users value a specific model more than the lowest price. Sakana, for its part, charges for the tokens used by its own orchestration, while Cerebras sells Qwen3.8-27B on its speed. Finally, Wringer is a reminder that compression to 2.655 bits per weight does not yet reduce the bill as long as inference reverts to bf16.
The final thread concerns the security of command-line agents, and it revolves around the same mechanism: approval requests. Gemini CLI requires one when a build file has been modified during the session or when a command uses arguments from a web page or an MCP server, and it isolates the sandbox from the user’s credentials. Vibe CLI 2.25.4 makes approval mandatory for shell syntax that bypassed its controls, citing four CVEs and variants of a fifth, and changes smart approve from blocking to asking for confirmation. Claude Code 2.1.270 illustrates the downside: a regression caused it to request authorization again for simple read-only git commands. Between the vulnerability that lets something through and the prompt that appears too often, tuning these requests is becoming a visible part of vendors’ work.
Sources
- Microsoft 365 on X, Grok models in Copilot
- Grok on X, sharing the Copilot announcement
- Grok Bot on X, Microsoft Teams
- OpenAI, OneGov 2.0 for U.S. government agencies
- OpenAI, Chris Lehane’s op-ed on regulation
- DeepSeek, Models & Pricing page
- DeepSeek, API Change Log
- Amp, Free Agent post
- ChatGPT on X, Pets and Appshots
- ChatGPT, application changelog
- Gemini CLI, September 12 nightly
- Mistral, Vibe CLI 2.25.4 release notes
- Hugging Face, Wringer
- OpenAI, Cognition case study
- OpenAI Developers on X, projects built with GPT-6 Astra
- Google Developers Blog, autofinetune
- Google Research, ToolGrad
- Google AI Studio on X, Gemini API documentation
- Google, three Android XR projects at the Venice Film Festival
- Claude Code, 2.1.270 release notes
- Claude, application release notes
- GitHub Changelog, VS Code Agents window
- GitHub on X, Universe 2026 ticket sales
- Sakana AI on X, Fugu Ultra v2 on OpenRouter
- Hugging Face, ChordStream
- Qwen on X, Qwen3.8-27B on Cerebras