Search

GPT-5.6 Sol unifies ChatGPT, Kimi K3 reaches general availability in Copilot, WeatherNext predicts cyclones five days ahead

Article generated by artificial intelligence
GPT-5.6 Sol unifies ChatGPT, Kimi K3 reaches general availability in Copilot, WeatherNext predicts cyclones five days ahead

ai-powered-markdown-translator

Article translated from fr to en with gpt-5.4-mini.

View project on GitHub ↗

On August 6, 2026, OpenAI unifies the ChatGPT experience around a shared GPT-5.6 Sol for both the fast and reasoning modes for paid subscribers, with 68% fewer factual errors, while GPT-5.6 Luna becomes unlimited and free for chats. On the same day, Kimi K3 arrives in general availability in GitHub Copilot, Google DeepMind publishes WeatherNext in Nature, its open model that predicts cyclones five days ahead, and Alibaba opens the public beta of Wan3.0 for 30-second video generation. Meta rounds out the week with Artificial Analysis benchmarks placing Muse Spark 1.2 near the frontier and a tally of five gold-level results at international science Olympiads, while around fifteen announcements cover agentic code tools (Amp, Devin, Replit, Warp, Cursor), the open-weight ecosystem (Hugging Face, Sakana, Together AI), and media generation (ElevenLabs, Suno).


ChatGPT switches to unified GPT-5.6 Sol, GPT-5.6 Luna becomes free unlimited

August 6 — OpenAI unifies the chat experience in ChatGPT around an updated version of GPT-5.6 Sol for paid subscribers. Until now, Plus and Pro users manually switched between a fast model (Instant) and a deep reasoning model. From now on, GPT-5.6 Sol handles both modes in a single experience, with a reasoning-effort slider adjustable on the fly — a direct simplification of the interface, in response to user feedback about confusion between the two former modes.

On reliability, OpenAI provides a concrete figure: on its high-stakes factuality evaluation (finance, medicine, law), the new version of GPT-5.6 Sol produces 68% fewer responses containing factual errors compared with GPT-5.5 Instant. This is the core argument of the update, presented as a gain in reliability rather than a jump in raw capabilities.

On the free offering side, Free and Go users also gain expanded access to GPT-5.6 Luna: unlimited text conversations starting August 7, and the appearance of a “Think” button to trigger more reasoning on difficult questions — a feature previously reserved for paid plans.

Tracked itemNumeric detail
Factual errors in finance/medicine/law-68% vs GPT-5.5 Instant
GPT-5.6 Sol + slider (Plus/Pro)available since August 6
Unlimited GPT-5.6 Luna chats (Free/Go)available since August 7

One scope note: this update concerns only the Chat experience in ChatGPT — the GPT-5.6 Sol version powering ChatGPT Work and Codex is unchanged by this announcement.

🔗 Official OpenAI announcement


Kimi K3 available in general release in GitHub Copilot

August 6 — GitHub announces the general availability of Kimi K3, Moonshot AI’s open-weight model, in Copilot. Rollout is gradual across the Pro, Pro+, Max, Business, and Enterprise plans, and the model is accessible from all Copilot surfaces: Visual Studio Code, Visual Studio, Copilot CLI, the cloud agent, the GitHub Copilot app, github.com, GitHub Mobile, as well as the JetBrains, XCode, and Eclipse plugins.

📣 @Kimi_Moonshot’s Kimi K3, an open-weight model, is now generally available and rolling out in GitHub Copilot. The model shows frontier-level abilities on agentic coding with highly cost-effective pricing. It is hosted by @FireworksAI_HQ. — @github on X

The model is hosted via Fireworks AI, with usage-based billing aligned to the provider’s pricing — the changelog does not specify exact prices or a context window. For Business and Enterprise organizations, Kimi K3 is disabled by default: administrators must explicitly enable the model policy before teams can access it. This addition is part of Copilot’s multi-model strategy, which had already integrated Kimi K2.7 Code earlier in August.

🔗 GitHub changelog


WeatherNext: Google DeepMind predicts cyclones five days ahead, open weights

August 6 — Google DeepMind publishes in Nature the results of WeatherNext, its AI model dedicated to forecasting tropical cyclones. The model achieves state-of-the-art accuracy for predicting a storm’s track and intensity, offering on average 24 hours of extra lead time compared with previous methods to prepare.

Predicting cyclones accurately can help save lives - and every hour of lead time counts. Published in @Nature, our AI model WeatherNext achieves state-of-the-art accuracy in forecasting a storm’s track and intensity, giving us a critical extra 24 hours to prepare on average. — @GoogleDeepMind on X

During Hurricane Melissa, WeatherNext provided early predictions of its Category 5 landfall, with five days’ lead time and 80% confidence. The model, trained on several years of global atmospheric data and nearly 5,000 historical cyclones, generates each 15-day probabilistic forecast scenario in under a minute on a TPU chip. Google DeepMind now provides 1,000 probabilistic forecasts per storm, accessible via the WeatherLab tool.

Measured indicatorMeasured value
Extra preparation lead time+24h on average
Hurricane Melissa alert (Category 5)5 days ahead, 80% confidence
Historical cyclones in training~5,000

A notable point for the ecosystem: the model’s code and weights are open on GitHub, freely available for academic research as well as operational forecasting.


Wan3.0 (Alibaba) enters public beta with native 30-second videos

August 6 — Alibaba launches the public beta of Wan3.0, the new generation of its video generation model, with three highlighted axes: native video generation of up to 30 seconds in a single pass, a visual rendering described as “reality-grade” (more expressive characters, better consistency of references from one scene to another), and an Omni-Reference function that extends the model’s understanding beyond text, image, audio, and video: Wan3.0 can now read structured documents (doc, xls, ppt, pdf, etc.) and generate a video directly from their content.

The beta is available now on Alibaba Cloud Model Studio and Qwen Cloud, with a coming arrival on the members-only Wan app.

Video resolutionAPI price (per generated second)
480p$0.05
720p$0.10
1080p$0.20

Full API access is expected to open gradually over the next few weeks. This launch positions Wan3.0 against recent competitor announcements in long-form video generation with fine multimodal control — FLUX 3 Video from Black Forest Labs, Seedance 2.5 from ByteDance, MiniMax H3 — a segment where the ability to generate 20- to 30-second clips with rich reference control becomes a central differentiator.

🔗 @Alibaba_Wan announcement


Meta’s Muse Spark 1.2 approaches the frontier, according to Artificial Analysis

August 5-6 — Independent analysis firm Artificial Analysis publishes its full evaluation of Muse Spark 1.2, Meta’s model launched alongside the Muse Code agent. On the AA Intelligence Index, it scores 54 (xhigh variant), up 3 points from Muse Spark 1.1 (51) and 11 points from Muse Spark 1.0 (43, released in April) — its third release in four months. This score places it effectively tied with GPT-5.5 (xhigh, 55) and Grok 4.5 (high, 54), just behind the current leading models: Claude Opus 5 (max, 61), Claude Fable 5 (max, 60), GPT-5.6 Sol (max, 59), and Kimi K3 (max, 57).

The main improvement is in agentic knowledge work, the weak point identified at the launch of Muse Spark 1.1: the GDPval-AA v2 Elo score rises from 1371 to 1631, placing the model 5th among all evaluated models, ahead of Claude Opus 4.8 (1588).

Measured indicatorMuse Spark 1.1Muse Spark 1.2
AA Intelligence Index (xhigh)5154
GDPval-AA v2 Elo13711631
Terminal-Bench 2.178%80%
Cost per task (Intelligence Index)$0.29$0.40

On cost, Muse Spark 1.2 remains among the most efficient models at its intelligence level, at 0.40pertaskforunchangedAPIpricing(0.40 per task for unchanged API pricing (1.25 / 4.25permillioninput/outputtokens).OnlyGrok4.5andGPT5.6Sol(medium)arecheaperinitsgroup.OntheValsIndex,MuseSpark1.2alsoentersthetop5atonly4.25 per million input/output tokens). Only Grok 4.5 and GPT-5.6 Sol (medium) are cheaper in its group. On the Vals Index, Muse Spark 1.2 also enters the top 5 at only 0.69 per test — three times cheaper than Kimi K3, ten times cheaper than Claude Fable 5, Opus 5, and GPT-5.6 Sol.

🔗 Artificial Analysis thread


Meta claims gold-level results at five international science Olympiads

August 6 — Meta Superintelligence Labs (MSL) announces that it entered internal versions of the Muse Spark family into five high-level international science competitions in 2026, to measure real progress in reasoning rather than on synthetic benchmarks — a cumulative tally after a first physics medal already reported in July.

To understand whether we’re making genuine progress on reasoning, we entered our AI models in five international STEM Olympiad competitions this year. The results: Asian Physics Olympiad (APhO): Perfect score on the theory exam — gold medal. International Physics Olympiad (IPhO): Perfect score — gold medal. International Mathematical Olympiad (IMO): Gold medal, top 4% of human participants. International Chemistry Olympiad (IChO): Gold-medal level. Romanian Masters of Mathematics (RMM): Gold-medal level. — @shuchaobi on X

Science competitionResult achieved
APhO (physics, Asia)Perfect theoretical score — gold
IPhO (physics, international)Perfect theoretical score — gold
IMO (mathematics)Gold, top 4% of human participants
IChO (chemistry)Gold-medal level
RMM (mathematics, Romania)Gold-medal level

Three of these competitions (APhO, IPhO, IMO) were live contests, graded by official judges according to the same scoring rules as for human candidates. Meta specifies that the models used had no access to external tools — no web search, no code interpreter, no calculator — with multi-agent orchestration in parallel reasoning.

🔗 @AIatMeta relay


AI development tools: Portals, remote sandboxes, model routing

Five announcements narrow the gap between local development and remote or multi-model environments this week.

Amp launches Portals — live app preview in Orbs

August 6 — Amp adds Portals, available now in all orbs (Amp’s remote development environment), a feature that exposes HTTP services running inside an orb directly in the browser. Until now, live testing the changes made by the agent required setting up a VPN or juggling ports. Technically, the agent creates or detects a .amp/services.yaml file and runs amp orb services ensure to launch the application, which is then exposed through a secure HTTPS connection. Portals are accessible to anyone with access to the thread, allowing multiple team members to see the same changes live, even when the orb wakes up after being put to sleep.

🔗 Amp — Portals

Devin Outposts now runs on Vercel Sandbox

August 5 — Cognition expands Devin Outposts (launched at the end of July to run Devin natively on any computer) with an integration to Vercel Sandbox, available now. The agent can now build and test an application in an isolated microVM on Vercel’s infrastructure, with access to Docker, connection to private networks via VPN, and resumption from a filesystem snapshot that preserves the repository, dependencies, and build state intact. This integration brings Devin Outposts closer to the kind of ephemeral cloud-environment offerings seen from Amp (Orbs) or Warp, with the advantage of relying on Vercel’s infrastructure already used by many front-end teams.

🔗 @cognition on X

Replit runs its security scans during the build, with Semgrep

August 5 — Replit announces, in partnership with Semgrep, that its security scans now run during application build rather than only at the end — an evolution that directly responds to a user request. These scans are enabled by default and are integrated into the full Replit Auto-Protect feature set, which automatically secures applications generated by the agent. Replit says its Agent already runs more than 100,000 scans per day, a figure that illustrates the scale at which the tool is used for production vibe coding, with vulnerability detection and fixes before deployment.

🔗 @Replit on X

Warp introduces custom model routers

August 5 — Just over a week after the launch of the Warp Agent CLI, Warp adds custom model routers, available now to all users. It is now possible to define your own routing rules in natural language (“send simple tasks to Minimax”, “send bug fixes to Kimi K3”), and each request is then automatically directed to the model deemed most appropriate. These routers add to the automatic routing already built into the CLI and work with open-weight models, in BYOK (bring your own key), or via a SuperGrok subscription — a direct response to the trend seen in several competing tools to orchestrate multiple models depending on the task rather than betting on a single frontier model.

🔗 @warpdotdev on X

Cursor details Cursor Router — up to 68% savings with no loss in quality

August 6 — Cursor publishes a technical post detailing how Cursor Router works, its automatic model selection system launched on July 22 — so not a launch, but an in-depth article with new data. The router works in two stages: “Compass” first estimates task complexity by predicting user satisfaction probability, then a taxonomy-based classifier (domain, task type, modifiers such as visual changes) selects the most performant frontier model for that specific category. The router distributes requests across Grok (routine tasks), Sol/GPT-5.6 (planning), Opus (execution), and Fable (debugging) — Opus 5 was added after the initial launch. Cursor provides concrete numbers: “Auto Intelligence” mode would deliver higher user satisfaction than Fable alone at a cost 68% lower, and “Auto Balance” mode would outperform Opus 4.8 for 41% less cost.

🔗 Cursor — how Cursor Router works


Open-weight ecosystem: inference, storage, and autonomous research

Three announcements strengthen the infrastructure around open-weight models, on both the hosting and research sides.

Baseten becomes official inference provider on Hugging Face

August 6 — Baseten, an AI infrastructure platform specializing in serverless inference, joins Hugging Face’s Inference Providers. The integration makes it possible to launch Baseten-hosted inference directly from model pages on the Hub, via the web interface or the Python and JavaScript SDKs. Initial support covers three major open-weight models: DeepSeek V4 Flash, Kimi K3, and GLM-5.2, with other task types planned later. PRO subscribers receive $2 in monthly inference credits usable across all providers, while free accounts have more limited quotas. The integration exposes an OpenAI-compatible API format via Hugging Face’s router, simplifying provider switching without rewriting calling code.

🔗 Hugging Face — Baseten

Ai2 expands its partnership with Hugging Face — storage increased to ~2 petabytes

August 5 — Ai2 (Allen Institute for AI) expands its partnership with Hugging Face to accelerate the dissemination of its open science work. The storage allocated on the Hub is nearly tripled, reaching about 2 petabytes, and traffic is no longer subject to standard rate limits — a significant gain for downloading the largest datasets and multi-checkpoint models. Ai2 says its assets have been downloaded more than 50 million times since spring 2024, and that artifacts related to its MolmoAct 2 model surpassed 400,000 downloads in three weeks. The partnership covers the entire Ai2 portfolio: the Olmo and Molmo models, the OlmoEarth Earth observation model, as well as several evaluation datasets. According to Ai2, the organization publishes more new artifacts on the Hub each year than any other structure tracked by Hugging Face, with more than 900 models and 1,200 datasets to its name.

🔗 Ai2 — Hugging Face partnership

Sakana Marlin — autonomous research agent and the launch of Marlin Insights

August 6 — Sakana AI highlights Sakana Marlin, described as a “Virtual CSO” (virtual chief science officer) capable of reasoning and conducting research autonomously for several hours. The company says this capability is built on two research components published by its teams: AB-MCTS, accepted as a spotlight at NeurIPS 2025, and AI Scientist, published in Nature. Sakana AI has also launched Marlin Insights, a series of research projects produced by Marlin; the first issue analyzes the outcome of the Strait of Hormuz crisis and why the United States, despite its military superiority, struggles to find an exit path. This launch illustrates Sakana AI’s strategy of turning its fundamental research into autonomous analysis products aimed at enterprise, continuing its recent commercial deployments (Daiwa Securities, Namazu).

🔗 @SakanaAILabs on X


GitHub Copilot and the ecosystem: open data, reusable design, slash commands

Q1 2026 update to the GitHub Innovation Graph

August 5 — GitHub publishes the quarterly update (Q1 2026) of its Innovation Graph, an open data project that tracks global development activity. The tweet thread highlights an acceleration of Git push activity in several countries outside traditional markets, as well as a 16% quarter-over-quarter increase in cross-border open source collaboration in Q1 2026, with notable gains for the European Union, India, and Singapore since 2020.

Country trackedGit pushes (Q1 2020 → Q1 2026)
India4.4M → 45M
Vietnam630K → 5.1M
Pakistan240K → 3.5M

All datasets remain open and downloadable, making it possible to explore these dynamics in greater detail by country and by period.

🔗 @github thread — Innovation Graph

Genspark Design — reusable design system across projects

August 6 — Genspark launches Genspark Design, a feature available now that automatically turns an already built project (colors, typography, components) into a reusable design system for subsequent projects. The goal is to avoid starting the visual identity from scratch for every new project: a landing page generated afterward inherits the same style as an existing app, without manual rework. This announcement adds to Genspark’s recent efforts around interface generation and office suites (GenOffice), and targets users who chain together multiple AI-generated projects while seeking visual consistency without repeated configuration.

🔗 @genspark_ai on X

Slash command guide in the GitHub Copilot app

August 6 — GitHub publishes a guide detailing the six slash commands available in the GitHub Copilot desktop app, accessible by typing “/” in the chat area. They cover the whole workflow: planning, stress-testing an approach, automated implementation, cross-review by another model, creating interactive interfaces, and multi-repo orchestration — standing apart from the CLI slash commands, which target the terminal rather than the desktop app.

Slash commandFunction provided
/planBreaks down the task before coding
/sparDevil’s advocate to challenge an approach
/autopilotTurns a plan into code
/rubber-duckIndependent review by another model
/orchestrateCoordinates work across multiple repositories

🔗 GitHub — slash command guide


Media generation: dubbing and audio transparency

ElevenLabs launches Dubbing v2 on ElevenAPI

August 6 — ElevenLabs makes version 2 of its dubbing engine (Dubbing v2), which already existed in the product, generally available on ElevenAPI, allowing developers to integrate AI dubbing directly into their own applications. Technically, Dubbing v2 conditions directly on the original audio performance rather than on a flat transcript, preserving the tone, emotion, and phrasing of the source voice. The model adapts phrasing for natural delivery in more than 90 languages, with “sync-aware” translation that automatically aligns the beginning and end of each line without manual adjustment. For enterprise use cases, a granular editing mode lets users import their own transcript and regenerate only the modified segments rather than the entire dub — with the whole pipeline (translation, voice cloning, dubbing, synchronization) exposed through a single API.

🔗 @ElevenLabs on X

Suno rolls out transparency tools and updates its community rules

August 6 — Suno publishes a post signed by its CEO Mikey Shulman, laying out the principles the company says it wants to follow to “responsibly” build the future of AI-generated music. Concretely, the announcement introduces a rollout of watermarking and audio fingerprinting technologies to identify tracks generated on the platform, presented as resistant to tampering without degrading the listening experience, as well as upcoming download restrictions aimed at limiting mass distribution to streaming platforms while preserving professional and creative uses. The announcement is accompanied by an update to the Community Guidelines, which now explicitly forbids recreating existing songs, using a person’s voice or likeness without consent, and spam or automated moderation evasion practices, with sanctions ranging from warnings to permanent bans.

🔗 @suno on X


Agents and plugins: open standards and subagent models

Agent Plugins — open standard for agent plugins

August 6 — OpenAI announces Agent Plugins, a new open standard developed with AWS, Cursor, GitHub, VS Code, and Vercel to package Agent Skills and MCP server configurations into a common format. The goal is to let developers build a plugin once and reuse it across multiple compatible agent clients, rather than duplicating the integration for each tool. At launch, the listed compatible clients are Codex, ChatGPT, Cursor, GitHub Copilot, Kiro, and VS Code — a broad scope covering both the OpenAI ecosystem and several direct competitors in the code-agent space. The tweets do not detail the exact technical mechanism (manifest format, publishing process) or a general availability date beyond the announced launch.

🔗 @OpenAIDevs on X

Perplexity Computer rolls out GPT-5.6 Terra and Luna as subagent models

August 6 — Perplexity rolls out two new GPT-5.6 models within Computer, its subagent and automation platform: Terra and Luna. Terra becomes the default model for all Computer subagents and also serves as the orchestrator model; Luna is positioned as the main model for scheduled automations. According to Perplexity, Terra is designed for complex goal-oriented work, a profile suited to the subagent role that must break down a task and execute it autonomously. On the WANDR benchmark, Terra reportedly scores 11 points higher than Claude Sonnet while delivering a tenfold cost reduction. This introduction of role-specialized models illustrates an increasingly differentiated agent architecture at Perplexity.

🔗 @perplexity_ai on X


Briefs

  • Together AI claims Kimi K3 is nearly twice as good as Claude Fable 5 on the Harvey LAB-AA legal benchmark — With no methodological details published, this should be taken as an inference provider marketing claim. 🔗 source
  • Hugging Face highlights Cadena, a 3D mesh decompiler into editable CAD — Spaces demo that converts a frozen 3D file into an editable CAD program step by step, for example to widen a hole without rebuilding the model. 🔗 source
  • Hugging Face CEO defends layer-based AI regulation — Clément Delangue supports the new U.S. regulatory framework distinguishing model weights, APIs, and applications, which he sees as protective of open source. 🔗 source
  • Together AI details its multi-datacenter architecture with 99.9% uptime — Traffic is split live across two sites simultaneously, with the ability to absorb the complete loss of a datacenter without service interruption. 🔗 source
  • Gemini Omni: free generation of 10 videos extended until August 11 — The offer, which was supposed to end on August 4, remains available until August 11, 2026, 11:59 p.m. Pacific Time. 🔗 source
  • GenOffice available on Linux — Genspark’s open source AI office suite, already on macOS and Windows, is now compilable and available on Linux. 🔗 source
  • Flux 3 available in Genspark’s AI video agent — Integration of Flux 3 (Black Forest Labs) with native audio and multilingual dialogue, for clips up to 20 seconds in 1080p. 🔗 source
  • Qwen3.8-Max climbs the public rankings — Alibaba claims 5th place in the Artificial Analysis Intelligence Index and 1st place in the Agentic Index, after a 2nd-place finish in the Image-to-WebDev Arena the day before. 🔗 source
  • Qwen-Image-3.0-Pro deployed on Qwen Cloud and fal.ai — Alibaba’s image generation model, already ranked 5th worldwide, becomes directly accessible on Qwen Cloud and on the third-party platform fal.ai. 🔗 source
  • OpenAI and the American Psychological Association partner on AI and youth mental health — Partnership to develop evidence-based recommendations and safeguards around the well-being of young users. 🔗 source
  • OpenAI publishes Signals data on global ChatGPT usage — New country-level indicators on adoption and usage trends, with no detailed figures available at this date. 🔗 source
  • Cohere partners with the University of Waterloo for an AI transformation certificate — The new “AI Transformation & Change Management” certificate is intended to prepare students for responsible AI adoption in business. 🔗 source

What it means

The week illustrates a clear shift: frontier intelligence is becoming the default option rather than an extra paid choice. OpenAI is unifying ChatGPT around GPT-5.6 Sol, which handles both speed and deep reasoning, while making Luna unlimited and free for Free and Go accounts. GitHub is following a similar logic on the developer side by rolling out Kimi K3, an open-weight model, in Copilot across all of its paid offerings. In both cases, the novelty is not a leap in raw capabilities but a change in distribution: models that were previously reserved for premium tiers or custom integrations are becoming the standard configuration, shifting competition toward perceived reliability and access cost rather than raw benchmark scores.

Measuring AI capability is itself becoming a structured competitive arena. Meta’s Muse Spark 1.2 is nearing the frontier according to Artificial Analysis, at a per-task cost that remains among the lowest in its category, while Meta Superintelligence Labs claims gold-level results across five international science Olympiads without tool use — a public demonstration designed to persuade beyond the usual leaderboards. At the same time, the infrastructure supporting these open models is becoming more professionalized: Baseten becomes an official inference provider on Hugging Face for DeepSeek V4 Flash, Kimi K3, and GLM-5.2, Ai2 is nearly tripling its storage on the Hub, and Sakana AI is turning its fundamental research into an autonomous research product with Marlin. The common thread: value is moving from the isolated model to the ecosystem that distributes, hosts, and verifies it.

Agentic coding tools are converging on the same idea: no single model dominates every task anymore, so routing has to be smart. Cursor explains how its router combines complexity estimation and taxonomy-based classification to cut costs by up to 68% without losing perceived satisfaction; Warp introduces natural-language routing rules; and Amp and Devin extend their remote environments (Portals, Vercel Sandbox integration) to bring cloud development closer to the feel of local development. Replit, for its part, is moving its Semgrep security scans earlier in the build rather than later, at more than 100,000 scans per day at this scale. Together, these announcements sketch a generation of tools where intelligence is no longer the only lever: orchestration, default security, and infrastructure integration matter just as much.

Beyond language models, the week shows AI taking hold both in directly impactful uses and in large-scale creative production. Google DeepMind’s WeatherNext, published in Nature and available as open weights, provides an average of 24 hours of additional lead time for forecasting a cyclone — an application where the public interest clearly takes precedence over commercial competition. By contrast, Alibaba’s Wan3.0, ElevenLabs, and Suno illustrate the rapid maturation of media generation: native 30-second video with multimodal control, voice-performance-conditioned dubbing now generally available via API, and watermarking tools deployed in response to concerns about the authenticity of generated content. These two directions, open science and commercial creation, are moving at the same pace but with very different responsibility stakes.


Sources