ai-powered-markdown-translatorTranslated article from fr to en with gpt-5.4-mini.
A big day for agentic infrastructure building blocks: Anthropic is releasing computer use, the browser tool, the Skills API and the Files API from beta on Claude Platform, while GitHub Copilot lands in Slack and Microsoft Teams. On the model side, DeepSeek is publishing its first multimodal model and Pika Labs is launching a voice synthesizer claiming unprecedented value for money. In between, there is a massive egocentric dataset, a study on benchmark bias in speech recognition, and around ten notable updates from Google DeepMind, GitHub, Z.ai, Qwen, Harvey, and video/audio generation tools.
Claude Platform: computer use, browser tool, Skills API and Files API generally available
August 20 — Anthropic announces the move to general availability of four building blocks for creating agents on Claude Platform: computer use, the browser tool, the Skills API and the Files API. The stated goal is to reduce the number of back-and-forth steps needed to automate a task in an application that does not have an API, and to enable the creation of Claude Managed Agents on versioned skills and reusable files.
The new computer_toolset_20260801 lets Claude chain multiple actions per turn (click, type, keyboard press, screenshot capture) instead of one round trip per isolated action. Anthropic says early access customers saw 20 to 40% fewer round trips per task. The browser tool is also evolving with browser_toolset_20260801: instead of targeting pixels — fragile at the slightest layout change — Claude now receives the page structure and can act on element references.
On the API side, the Skills API lets you upload a team’s procedure once, version it, and pin it to a specific version_id or to latest. The Files API gets a expires_in_seconds parameter, a 5x higher rate limit (500 requests per minute), and a quota of 1 TB per organization. Anthropic has also released an AG-UI adapter for Managed Agents, which maps each conversation thread to a managed session and streams text, tool calls, and reasoning to the interface.
| Building block | Today’s update |
|---|---|
| Computer use | computer_toolset_20260801, up to -40% fewer round trips per task |
| Browser tool | browser_toolset_20260801, element-reference targeting |
| Skills API | Versioning, pinning to version_id or latest |
| Files API | 500 requests/minute (x5), 1 TB quota per organization |
🔗 Claude Platform announcement
GitHub Copilot arrives in Slack and Microsoft Teams
August 21 — GitHub Copilot arrives in public preview on Slack and Microsoft Teams. The principle is the same in both cases: mentioning @GitHub in a direct message, channel, or thread opens a cloud agent session capable of answering questions about code and GitHub activity, sorting bug reports, investigating failures, implementing changes in a secure cloud sandbox, and then opening pull requests with a link back to the original conversation.
On Slack, the integration relies on Slack Code, Slack’s new agentic offering launched the same weekend: GitHub Copilot can open dedicated code channels where the team reviews diffs and inspects render previews without polluting the original conversation. On Teams, a discussion becomes a collective agent session that everyone can follow and guide; the work continues asynchronously in the cloud sandbox while the team does something else.
Availability is limited to organizations on the Copilot Business or Enterprise plan. The administrator enables the Copilot cloud agent policy and installs the app for Slack or Teams; repository administrators can require approval before merging pull requests generated by the agent. Usage consumes AI credits, and the cloud sandbox is billed separately with configurable budgets.
🔗 GitHub Copilot in Slack · 🔗 GitHub Copilot in Teams
DeepSeek-V4-Flash-Vision-Exp: DeepSeek’s first multimodal model on the API
August 21 — DeepSeek puts DeepSeek-V4-Flash-Vision-Exp online, an experimental multimodal model available under the identifier deepseek-v4-flash-vision-exp. On purely text capabilities — agents, reasoning, world knowledge — the model matches DeepSeek-V4-Flash. The novelty lies elsewhere: on multimodal agent benchmarks, DeepSeek reports a sharp jump over V4-Flash, with performance approaching Opus-4.8.
Multimodal support on the API side is complete: mixed text + image input (base64, external URL, or via a new free Files API launched the same day), compatible with the three existing formats (Chat Completions, Messages, Responses). Pricing follows the V4-Flash grid, with image tokenization capped at 384 tokens per image. The Files API lets you upload an image once and reference it by file_id, saving bandwidth for repetitive agentic workflows. DeepSeek Harness 0.1.1, the open-source agent harness launched on August 13, was updated the same day to support the new model natively.
| Item | Detail |
|---|---|
| Model | deepseek-v4-flash-vision-exp |
| Text capabilities | Parity with DeepSeek-V4-Flash |
| Multimodal capabilities | Close to Opus-4.8 on multimodal agent benchmarks |
| Image pricing | Up to 384 tokens/image, V4-Flash pricing |
| Supported APIs | Chat Completions, Messages, Responses |
| Harness | DeepSeek Harness 0.1.1 (native support) |
DeepSeek-V4-Flash-Vision-Exp is now live on the DeepSeek API Platform! This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge. On multimodal agent benchmarks, V4-Flash-Vision-Exp makes a major leap over V4-Flash, bringing multimodal agent performance close to Opus-4.8. — @deepseek_ai on X
Pika Speech: 3B TTS at 1.2 seconds per minute, up to 9x cheaper than ElevenLabs v3
August 21 — Pika Labs expands its Pika Audio lineup with Pika Speech, a 3-billion-parameter text-to-speech model. The main selling point is speed: a real-time factor (RTF) of 0.02, meaning one minute of studio-quality 48 kHz speech is generated in about 1.2 seconds in long-form local tests, with requests possible up to five minutes of continuous speech.
On rhythm control, Pika Speech breaks new ground: instead of stretching or compressing audio afterward — the standard method in most TTS systems — the model controls rhythm and duration directly thanks to an “end-of-speech latent” (EOS latent), an anchor placed at the target final frame. Moving this anchor earlier compresses the delivery; moving it later lets the text breathe. On the architecture side, the team combines flow matching and distribution-matching distillation, FlashAttention-3, and “token packing” that makes five sentence segments cost about twice the price of one instead of five times.
On cost, Pika claims an advantage of up to 9x versus ElevenLabs v3, 4.5x versus Cartesia and ElevenLabs Turbo, and 2x versus Fish Audio — without detailing the exact methodology beyond those ratios. The model is accessible via the Pika API Club.
Let’s talk about Pika Speech, a 3B text-to-speech model at RTF 0.02. […] Pika Speech generates one minute of studio-quality 48 kHz speech in about 1.2 seconds—and supports requests up to five minutes. That efficiency makes it up to 9× more cost-efficient than ElevenLabs v3, 4.5× than Cartesia and ElevenLabs Turbo, and 2× than Fish Audio. — @pika_labs on X
EgoSuite-Open100K: the largest open human egocentric dataset
August 21 — Lightwheel AI, in partnership with Hugging Face, open-sources EgoSuite-Open100K, presented as the largest fully annotated human egocentric dataset ever released openly. The stated goal is 100,000 hours of first-person human data; the first 10,000 hours are already available, with the rest to be published in stages.
The dataset covers more than 15,000 tasks spread across more than 15,000 real-world scenes — kitchens, bedrooms, warehouses, assembly lines — grouped into 7 environment categories and 128 scene types. Each sequence is annotated with hand pose, body pose, and subtask-level semantics, and part of the corpus also includes wrist-mounted camera coverage. A notable licensing point: unlike many research datasets, EgoSuite-Open100K is licensed for commercial training, not just academic use.
The dataset is presented by its creators as the first public building block of a broader data infrastructure for “Physical AI” (robotics and embodied AI) — the idea that scaling physical models now depends on access to massive, shared human data.
🔗 Lightwheel AI announcement · 🔗 Hugging Face relay
Claude Security moves to Claude Mythos 5
August 21 — Anthropic is running Claude Security, its vulnerability scanning tool, on Claude Mythos 5, in public beta for all Claude Enterprise customers. The pitch: benefit from Anthropic’s most capable security model on its codebase without separate model access — scans are billed as standard token usage on the existing plan.
In practice, the tool points at a GitHub repository and Mythos 5 analyzes the code by tracing data flow across files and reasoning about interactions between components. Each result comes back with a CWE category, a confidence level, a severity assessment, and a suggested fix, which opens directly in Claude Code on the web. Anthropic is also launching a Defender Advantage Fund with $35 million in credits for open-source security, and plans to expand its Cyber Verification Program in the coming weeks.
🔗 Claude Security announcement
GPT-5.6 Sol pricing: OpenAI lowers API prices, GitHub Copilot follows
August 21 — OpenAI is cutting GPT-5.6 Sol API and credit prices by more than 20% over the next three months. The change is already active on the API side and is rolling out in Codex and ChatGPT Work during the day. For Codex users on token-based plans, purchased credits now go further; included usage in Pro, Plus, and Business subscriptions remains unchanged — so the reduction mainly benefits API developers and token-billed users.
GPT-5.6 Sol at -50% in GitHub Copilot
In parallel, GitHub is relaying a separate promotion: -50% on GPT-5.6 Sol in GitHub Copilot until September 3, usable in the app, CLI, and IDE. This is a one-off Copilot discount, not to be confused with OpenAI’s permanent price cut described above — two different companies, two distinct mechanisms on the same model.
🔗 OpenAI price cut · 🔗 GitHub Copilot promotion
GLM-5.3 (Max) climbs on Code Arena: WebDev
August 20 — LMArena says that Z.ai’s GLM-5.3 (Max) moved the Pareto frontier (quality/cost ratio) in the Code Arena: WebDev ranking, dedicated to evaluating models on web development tasks. With 1597 points, the model would rank 2nd among open-weight models (once released) and 8th overall.
| Model | Price ($/M tokens) | Code Arena: WebDev points |
|---|---|---|
| GLM-5.3 (Max) | $3.65 | 1597 |
| Qwen3.8 (Max) | $5.00 | — |
| Gemini-3.7-flash-high | $2.86 | lower score |
| DeepSeek-v4-flash-high | $1.10 | lower score |
At $3.65 per million tokens, GLM-5.3 (Max) remains in a price range comparable to Qwen3.8 (Max) while outperforming Gemini-3.7-flash-high and DeepSeek-v4-flash-high on this ranking — another signal of Z.ai’s competitiveness in agentic web development, alongside already known results on Terminal-Bench and the general Agent Arena.
Harvey launches Tenet, a legal model built on Kimi K3
August 20 — Harvey presents Tenet, its first post-trained model for law, built on a Kimi K3 base (Moonshot AI) post-trained with Fireworks AI on a corpus of public legal data, synthetic data, and human expert data simulating long-duration legal work.
Harvey says this post-training increases Tenet’s full success rate by 82% on the LAB benchmark and by 22% on LAB Contracts, compared with the base Kimi K3 — with state-of-the-art results on LAB Contracts and 2nd place on LAB overall. Tenet’s operating cost would be less than a quarter of the most performant foundation models on the market. In addition, Harvey trained three specialist sub-agents: M&A diligence, table review, and firm knowledge.
Georgia Tech traces the origins of Olmo 3’s social reasoning
August 21 — A Georgia Tech team leveraged the fact that Ai2’s Olmo 3 stack is fully open — weights, Dolma 3 pretraining corpus (1.26 billion documents), infrastructure — to statistically link model capabilities to the texts that produced them, an experiment impossible on a closed model.
The method relies on influence functions, applied to about 5.68 million documents sampled from Dolma 3. Main result: texts rich in dialogue and interpersonal writing weigh more heavily on social reasoning performance than on factual knowledge — and when the most influential literary documents were removed, performance on the SocialIQA benchmark dropped significantly, confirming causality.
”Benchmark optimization” bias detected in open source ASR models
August 21 — A Hume AI team publishes on the Hugging Face blog a study on a methodological bias in speech recognition: some models would seem optimized to reproduce public benchmark reference transcripts rather than faithfully transcribe the audio.
Across 11 open source ASR models tested, about 40% of the VoxPopuli excerpts analyzed would contain errors in their references, and 6 out of 11 models reproduced those errors instead of correctly transcribing the audio. On LibriSpeech, masked-number recovery rates reached 30 to 40%, a sign that some models “fill in” the text rather than hearing it. The conclusion recommends using secret evaluation sets and a stricter separation between training and test data.
SenseNova-U1.5-8B-MoT: open image model without VAE or DiT
August 21 — SenseNova is releasing an open-weights image generation model whose reported performance would be comparable to Nano Banana 2. The technical interest lies in the architecture: no VAE, no separate text encoder, no DiT — the model is based on a “mixture of transformers,” where text and image tokens use distinct weights and attend to each other, with denoising happening directly in pixel space rather than in a compressed latent space.
This approach differs sharply from the standard stack of text-to-image diffusion models, and its open weights will allow the community to independently verify the reported performance comparisons.
Diffusers 0.40.0: Modular Diffusers leaves experimental status
August 21 — Hugging Face releases Diffusers 0.40.0. The most structural change is the move out of experimental status for Modular Diffusers, the library’s modular pipeline system. The release adds integration for several recent open models (LTX2.5, MiniMax H3, MiniMax Music 3, Wan Animate 2), support for tensor parallelism for a selection of models, and two new quantization backends (SDNQ, Nunchaku-Lite).
🔗 Diffusers 0.40.0 announcement
Google DeepMind explores AI in games with Fenris Creations
August 21 — Google DeepMind announces a research partnership with the studio Fenris Creations, building on 15 years of work using video games as an experimentation ground for AI, from mastering Atari to Grandmaster level on StarCraft II. After SIMA, which taught agents to understand 3D worlds, DeepMind wants to explore navigation in real human dynamics within a persistent universe.
The partnership targets four open challenges: continual learning (acquiring new skills without forgetting previous ones), deep memory systems going far beyond current context windows, long-horizon planning (weeks, months, years), and multi-agent dynamics (cooperation, negotiation, economy, emergent behaviors). The long-term goal is twofold: create new game experiences with developers, and reuse these advances for real-world problems and scientific discovery.
🔗 Google DeepMind announcement
Runway Ruby: SDR to 16-bit HDR video conversion
August 21 — Runway launches Ruby, a model that converts SDR (standard dynamic range) video into 16-bit HDR, outputting EXR sequences or 10-12 bit ProRes/HEVC, in the BT.2020 color space with PQ or HLG support — the standards used in professional post-production and HDR broadcast.
Ruby works equally well on video uploaded by the user and on output generated by Runway, with a 30-second limit. The feature is available now for Max and Enterprise plans.
NVIDIA AVO claims 100% on the ARC-AGI-3 benchmark
August 21 — NVIDIA presents AVO, a general-purpose code agent, with a score of 100% on ARC-AGI-3, an interactive reasoning benchmark. According to NVIDIA, AVO completed the 183 levels across the benchmark’s 25 public environments, without explicit instructions, stated rules, or preformulated objectives.
AVO operates through a continuous loop of inspection, planning, implementation, and evaluation, relying on memory, tools, and execution feedback to build on what it learns over time — an approach meant to preserve progress on long tasks rather than starting from zero with each new context window.
HeyGen Look Packs: change outfit and background while keeping the face
August 21 — HeyGen launches Look Packs, a feature that changes the outfit and background of an avatar video while faithfully preserving the person’s face — the argument being that most generative models also alter facial features when changing appearance.
The intended use is personal brand consistency across different professional contexts: real estate, legal, fitness, education, wellness. A B2B content creator can thus produce videos in outfits and settings suited to each sector while remaining recognizable from one video to the next. This announcement fits into HeyGen’s brisk pace over the past few days (Retouch on August 20, Brand Kits and 4K upscaling on August 18-19).
Amp launches Explain Usage to explain token consumption
August 21 — Amp adds Explain Usage, a feature that lets you directly ask Puck, the built-in assistant, to analyze its own consumption of tokens, credits, and orb size. The user asks a natural-language question — “Which threads used the most tokens today?” — and Puck reads the personal usage data and per-thread metrics to respond.
There are two entry points: ask Puck directly, or click the dedicated button on the usage pages. For users who prefer raw data, Amp also exposes this information via CLI (amp usage --details, amp threads usage <thread-id> --details).
v0 (Vercel) adds Vercel Connect to connect to 100+ services
August 21 — Apps and agents built in v0 can now connect directly to Slack, GitHub, Salesforce, and more than 100 other third-party services via Vercel Connect, without the developer having to manage the authentication flow themselves.
The mechanism relies on reusable team connectors associated with short-lived tokens: once a connector is configured at the team level, each new app or generated agent can reuse it instead of re-requesting full authentication every time. This positions v0 as a platform capable of producing agents integrated with a company’s software stack, not just isolated interfaces.
Shared threads in Codex and ChatGPT Work
August 20 — Codex and ChatGPT Work introduce “shared threads”: a read-only link that exposes the full reasoning of a session — process, decisions, context — rather than just a final result. The goal is to avoid reconstructing pull request context from scattered screenshots during a code review, deep technical exploration, or project handoff between teammates.
This is a feature that fits into the continuity of recent announcements around ChatGPT Sites (multi-user editing, URL customization) and strengthens Codex/ChatGPT Work’s positioning as a team tool.
Transparent backgrounds in preview for GPT-Image-2
August 20 — The GPT-Image-2 API now allows, in preview, the generation of images with transparent backgrounds. The practical benefit is to produce reusable assets — product visuals, graphic design elements, website mockups, marketing campaigns — that can then be placed on any background without manual cutout cleanup.
Briefs
- Zed — Antigravity installs in one click from Zed’s ACP Registry. 🔗 Tweet
- trackio 0.36 — Hugging Face releases an update adding support for a new loggable data type, with no further details. 🔗 Tweet
- “Full-stack” AI according to Google DeepMind — Paige Bailey breaks down Google’s end-to-end approach into five layers: infrastructure, security, research, models/tools, products. 🔗 Google article
- GitHub — better handling of blocked users: search by name, sorting, and pagination for long lists. 🔗 Changelog
- GitHub — pinning views, projects, and milestones in the repository sidebar reaches general availability, with several related improvements (reaction avatars, dashboard density). 🔗 Changelog
- Manus extends free access to Manus 1.6 Lite and Manus 1.6 until August 28 (SGT), subject to the daily quota. 🔗 Tweet
- Manus reminds notified users to save their data before August 23 at 7:59 a.m. (SGT), before its transition into an independent company. 🔗 Tweet
- Genspark launches a promotion until September 30: Design + GenTeam at -90%, GenMail free, Slides Ultra Mode at -50%. 🔗 Tweet
- NVIDIA details the split between an agent’s harness (what it tries to do) and the infrastructure (what it can actually do) in a deep dive on agentic stack security. 🔗 Tweet
- NVIDIA publishes the recap of the Seattle Spark Hack Series on DGX Spark: winning projects in disaster response, health, and offline survival with Nemotron and NemoClaw. 🔗 Tweet
- NVIDIA Cosmos 3 — a Cosmos Labs video shows post-training of Cosmos 3 with partners Aigen and Linker Vision. 🔗 Tweet
- Luma Labs will speak on creative agents at Summer Signal, organized with GMI Cloud on September 14. 🔗 Tweet
- Synthesia — its CEO is quoted in a Forbes Tech article on measuring AI adoption by outcomes rather than by usage. 🔗 Tweet
- Qwen3.8-27B continues its community momentum: new NVFP4 + DFlash2 inference recipes in the SGLang cookbook, lighter versions released by Unsloth, 1st place on the Cline code agent leaderboard in four days, and an NVIDIA DGX Station reporting more than 2713 tokens/second peak aggregate in BF16 service. 🔗 SGLang · 🔗 Unsloth · 🔗 Cline · 🔗 DGX Station
- Qwen3.8-Max is now usable in the Kilo Code coding tool for interface generation. 🔗 Tweet
- Kimi Work publishes a second tutorial for financial analysts: investor dashboard, financial model updates, batch report generation. 🔗 Tweet
- GLM-5.3 continues its progress: score of 69 on the official DeepSWE leaderboard, availability on the WorkBuddy platform, and an extension of Build Week by Z.ai until August 23 at 6:00 p.m. (PT) with 100 million free tokens for the first 50,000 new ZCode sign-ups. 🔗 DeepSWE · 🔗 WorkBuddy · 🔗 Build Week
- GPT-5.6 Sol — three third-party platforms (Router, Cloudflare AI Gateway, OpenCode Zen) are each offering -50% until mid-September, in addition to the discounts already covered at Devin, OpenRouter, and v0. 🔗 Cloudflare AI Gateway
What this means
Today’s Anthropic-GitHub announcement points to a broader trend: agentic infrastructure is leaving the experimental stage. The GA of computer use and the browser tool on the Claude Platform, combined with the arrival of Copilot in Slack and Teams, signals that major vendors are no longer just selling models — they are selling agents capable of acting inside the tools teams already use every day, without needing a dedicated API integration each time. This is as much a change in distribution as in capability: the agent now goes where the work happens, rather than requiring the work to come to it.
On the model side, DeepSeek-V4-Flash-Vision-Exp confirms that multimodality is becoming an expected standard rather than a differentiator — even players positioned around cost efficiency (V4-Flash) are now integrating it natively, with performance approaching the best closed models. Competition on code agent benchmarks (GLM-5.3 on Code Arena: WebDev, Qwen3.8-27B on Cline) illustrates the same dynamic on the development tools side: the price gaps between open and closed models are narrowing, which is pushing OpenAI to cut its API prices by more than 20% on GPT-5.6 Sol in the process.
The move toward full openness — not just weights, but data and training infrastructure — continues to bear scientific fruit: the Georgia Tech study on Olmo 3 simply would not have been possible on a closed model, just as the Hume AI study on ASR benchmark bias is a reminder that a leaderboard score does not guarantee real transcription quality. These are two useful reminders at a time when the industry is multiplying record-score announcements.
Finally, the multiplication of massive commercially licensed datasets — EgoSuite-Open100K today, after several similar announcements in recent weeks — confirms that access to real human data, not just synthetic data, is becoming a strategic issue for robotics and embodied AI, just as compute has been for large language models.
Sources
- Claude Platform — GA computer use, browser tool, Skills API, Files API
- GitHub Copilot in Slack
- GitHub Copilot in Teams
- DeepSeek-V4-Flash-Vision-Exp
- Pika Speech
- Pika Speech architecture details
- EgoSuite-Open100K
- EgoSuite Hugging Face relay
- Claude Security on Mythos 5
- GPT-5.6 Sol price cut (API)
- GPT-5.6 Sol -50% in GitHub Copilot
- GLM-5.3 (Max) on Code Arena: WebDev
- Harvey Tenet
- Georgia Tech traces the origins of Olmo 3’s social reasoning
- Benchmark optimization bias in ASR
- SenseNova-U1.5-8B-MoT
- Diffusers 0.40.0
- Google DeepMind × Fenris Creations
- Runway Ruby
- NVIDIA AVO on ARC-AGI-3
- HeyGen Look Packs
- Amp Explain Usage
- v0 Vercel Connect
- Shared Codex / ChatGPT Work threads
- GPT-Image-2 transparent backgrounds
- Zed × Antigravity
- trackio 0.36
- Full-stack AI according to Google DeepMind
- GitHub — blocked user management
- GitHub — sidebar pinning
- Manus — extended free access
- Manus — backup reminder
- Genspark — promotion
- NVIDIA — agentic stack security
- NVIDIA — DGX Spark Seattle hackathon
- NVIDIA Cosmos 3 — post-training
- DGX Station — Qwen3.8-27B
- Luma Labs — Summer Signal
- Synthesia in Forbes
- Qwen3.8-27B — SGLang cookbook
- Qwen3.8-27B — Unsloth
- Qwen3.8-27B — 1st place Cline
- Qwen3.8-Max — Kilo Code
- Kimi Work — financial analyst tutorial
- GLM-5.3 — DeepSWE score
- GLM-5.3 — WorkBuddy
- GLM-5.3 — Z.ai Build Week
- GPT-5.6 Sol — third-party discounts