ai-powered-markdown-translatorTranslated article from fr to en with gpt-5.4-mini.
MiniMax launches H3, a general-purpose multimodal model with native stereo sound and announced open weights “in the coming days.” Together AI claims spectacular growth in its volume of served tokens, while GitHub tightens the security of npm tokens after a wave of account takeover attacks. Around these three topics, the day is especially dense on the image and video generation front (Runway, Wan Video, Luma, Black Forest Labs, Ideogram), code-agent tooling (Devin, Replit, GitHub Actions), and model transitions at OpenAI and Google.
MiniMax H3: a general-purpose multimodal model, open weights announced
July 31 — MiniMax launches H3, described as a general-purpose multimodal generation model. Unlike previous generations (Hailuo 01 and 02), which separated tasks by specialized model, H3 unifies these capabilities in a single model able to understand a context blending text, image, video, and audio. Its most striking feature is native stereo sound generated jointly with video, on sequences of up to 15 seconds at 2K resolution.
MiniMax emphasizes an aggressive price-performance positioning: at 2K, the price per second would be less than one third that of “mainstream” models, and at 768p, less than half the 720p price of those same models. H3 relies on several proprietary building blocks (Contextual Omni Representation, H3-VAE, H3-Omni Transformer, In-Context Regeneration) and handles editing in a generalized way: image-to-image, image-to-video, audio-to-audio, video-to-video motion transfer, with precise rendering of text and brand elements — a point that is often weak in video generation models.
MiniMax also announces the upcoming release of the model weights, presented as a way to accelerate hardware compatibility and support the open source community in a sector historically dominated by closed models. The launch comes with day-0 deployment at around fifteen partners, including Runway, Pika, Leonardo.Ai, OpenRouter, and Krea.
| Characteristic | Detail |
|---|---|
| Resolution | 2K by default |
| Maximum length | 15 seconds |
| Audio | Native stereo, jointly generated |
| Price at 2K | Less than one third of mainstream models’ price/sec |
| Price at 768p | Less than half the price of mainstream 720p |
| Open weights | Announced “in the coming days” |
MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities — Today, we’re launching MiniMax H3, a general-purpose multimodal generation model. H3 understands unified context across text, images, video, and audio, generating video with native stereo sound, up to 15 seconds at 2K resolution. — @MiniMax_AI on X
🔗 Related thread — @MiniMax_AI on X
Grok Imagine Video 1.5 available on Runway
August 1 — Runway continues expanding its catalog of third-party models by integrating Grok Imagine Video 1.5, the video version of xAI’s generation model, directly accessible from its platform. The integration fits into Runway’s strategy of aggregating frontier models from multiple vendors — it had already welcomed MiniMax H3 the day before — within a single creative production platform, rather than relying solely on its own Gen models.
🔗 Announcement — @runwayml on X
Wan Video (Alibaba) launches a Realtime mode for sketch and text
July 31 — Alibaba has gone live, in the wan video Labs section, with a Realtime mode that turns a sketch into a photorealistic scene as the drawing progresses — stick figure, rough rectangle, a few strokes — without an explicit generation step or a “generate” button. The same day, Alibaba introduced a second Realtime mode, this time for text: the image is edited live as the user types a description, sentence by sentence, without a loading screen. These two features bring the image-generation experience closer to a continuous interaction, like drawing or word processing, rather than a classic prompt/generation cycle.
🔗 Announcement — @Alibaba_Wan on X
Luma launches Layers, layer-based image editing powered by Uni-1
July 30 — Luma rolled out Layers within Luma Agents, powered by its Uni-1 model. The idea: ask the agent to extract the layers of an existing image (generated by Luma or imported by the user), then modify a specific element through simple conversation while preserving the rest of the image intact — whether it is the featured product, the background, or text. Luma illustrates the feature with two examples: a full redesign of a rebranding poster from a simple flat JPG, and the addition of an isolated object in a scene without touching the lighting or set. The goal is to enable targeted, iterative edits rather than regenerating an entire image to change a single detail.
🔗 Announcement — @LumaLabsAI on X
Black Forest Labs reveals a preview of FLUX 3
July 30 — Black Forest Labs (BFL), the publisher of the FLUX models, gave a first public preview of its next model, FLUX 3, via a temporary 48-hour access window on Nous Research’s Hermes agent. This agent chains together several generated plans to produce a full short film from a single end-to-end prompt. The announcement is accompanied by a short-film contest organized by Nous Research throughout the week, and an in-person presentation of FLUX 3 at a BFL event in San Francisco the following Friday. At this stage, BFL has shared neither a general availability date nor technical details on the model’s capabilities.
Ideogram launches P-Image-Ideogram, a Pareto-optimal model family with Pruna AI
July 30 — Ideogram launched, in collaboration with Pruna AI, a new image model family called P-Image-Ideogram, optimized simultaneously for quality, speed, and cost. The model offers four quality modes (Very low, Low, Medium, High), native generation in 1K and 2K, and a starting price of $0.003 per image, with latency announced between 3 and 8 seconds depending on the mode. On the Design Arena ranking (blind human votes), Ideogram claims a new Pareto frontier for preference/cost and preference/speed: at 1K, the most economical mode costs 0.3¢/image for a generation of about 3 seconds, or roughly 8.6x faster than the average of the evaluated models. Immediate rollout on around fifteen partner platforms, including ComfyUI, Runware, Replicate, Leonardo.Ai, and Together AI.
| Quality mode | Resolution | Price at 1K | Approximate latency |
|---|---|---|---|
| Very low | Native 1K/2K | 0.3¢/image | ~3 s |
| Low | Native 1K/2K | 0.75¢/image | ~3 s |
| Medium | Native 1K/2K | not specified | up to 8 s |
| High | Native 1K/2K | not specified | up to 8 s |
🔗 Announcement — @ideogram_ai on X
Devin launches Outposts, native app execution on any computer
July 31 — Cognition, the publisher of Devin, announced a new feature called Outposts. It allows the Devin agent to run and test applications natively directly on any computer, rather than relying solely on isolated cloud environments. This capability continues Cognition’s recent announcements around native execution — native iOS support via macOS cloud agents, announced the day before — and aims to broaden the range of tasks Devin can handle end to end, including on hardware or software configurations specific to the user’s workstation.
🔗 Announcement — @cognition on X
Replit rolls out Model Selector, Follow-up Tasks, and a new usage page
July 31 — In its weekly “This Week in Replit” recap, Replit details several new features. Model Selector is now available for Core and Pro plans, with a broader selection including open-weight models in addition to proprietary models — the user keeps control, or lets Replit recommend automatically. Follow-up Tasks becomes available to all users this week: after finishing a task, the Agent suggests next steps that can be activated with one click and run in the background. Two adjustments round out the week: a rebuilt usage page offering near-real-time tracking by date, resource, project, workspace, group, and user, and an extended user session (default logout increased from 7 to 14 days).
| Feature | Status |
|---|---|
| Model Selector | Live, Core and Pro plans, open models included |
| Follow-up Tasks | Rolled out to all users |
| Usage page | Rebuilt, near-real-time tracking |
| Session length | Extended from 7 to 14 days |
Together AI: scaling open inference
Together AI is multiplying announcements around dedicated inference this week, on two distinct fronts: a striking growth figure on usage, and a detailed technical article on autoscaling the underlying infrastructure.
Token volume served multiplied by more than 10,000
August 1 — Together AI is showing a striking growth figure: its volume of served tokens would have gone from 30 billion per month to 400,000 billion (400 trillion) per month, or more than 10,000x. The company explains this progress by the migration of large-scale workloads to open-weight models rather than closed proprietary APIs. The figure is presented without a precise reference period, which limits its comparative value.
Josh had to stop @vipulved and ask him to repeat the number: @wolfejosh Together AI went from serving 30B tokens a month to 400T. Over 10,000x growth as AI-native companies and enterprises move scaled workloads to open model. — @togethercompute on X
Eight metrics to drive dedicated inference autoscaling
July 31 — Together AI publishes a technical article on autoscaling its Dedicated Model Inference. The starting point: classic metrics inherited from the web world (CPU/GPU utilization) do not reflect the real pressure on an LLM inference engine — a GPU can show 60% utilization while its queue is already saturated, because utilization measures compute intensity, not waiting load. The article details eight inference-native metrics, grouped into three families (concurrency, SLO, efficiency), and recommends inflight_requests as the default metric. It also quantifies the real cost of a cold start on an H100 GPU: about 86 seconds for a base model, 145 seconds for an 18 GB fine-tune, and about 2.5 minutes to add a replica.
| Measured step | Duration on 1×H100 |
|---|---|
| Base model deployment | 86 s |
| Fine-tune deployment (18 GB) | 145 s |
| Scaling from 1 to 2 replicas | ~2.5 min |
Ai2 — an infinity-gram-based study traces borrowed language in AI-written books
July 31 — Ai2 (Allen Institute for AI) highlights a study built on its infinity-gram engine, a tool that indexes massive text corpora and counts the exact frequency of any phrase, making it possible to trace back to the likely sources of a formulation. Tuhin Chakrabarty’s team (Stony Brook University) isolated “rare expressions” — phrases appearing in at most five Google Books volumes and absent from the indexed web — and compared their rate of presence across self-published books with high detected AI content, self-published books without detected AI, and award-winning literature. Result: books with high AI content show a significantly higher proportion of rare expressions, suggesting a stronger dependence on patterns from training data.
| Evaluated group | Rare-expression coverage |
|---|---|
| Top 200 self-published, AI strongly detected | 41.6% |
| Top 200 self-published, AI not detected | 37.2% |
| Awarded or nominated literature (reference) | 19.1% |
Gemini 3.5 Flash Cyber, a restricted-access model dedicated to cybersecurity
July 31 — Google quietly added a new member to its Flash model family: Gemini 3.5 Flash Cyber. Unlike the consumer versions announced this month (3.6 Flash, 3.5 Flash-Lite), this model targets a very specific use case — detecting and fixing software vulnerabilities at scale — and remains an economical model optimized for that single use case. Access is deliberately restricted to governments and trusted partners, making it a defensive cybersecurity tool rather than a commercial product open to everyone. No dedicated product page, benchmark, or architecture details accompany the announcement: the only available source is a biweekly recap thread from @GoogleAI, which limits how much can be said about it for now.
The same recap thread also announces Collections in Gemini Notebook (formerly NotebookLM), a new way to group notebooks. Again, just one sentence, with no associated screenshot or documentation.
🔗 Recap thread — @GoogleAI on X
Copilot deprecates Gemini 2.5 Pro and Gemini 3 Flash
31 July — GitHub is removing Gemini 2.5 Pro and Gemini 3 Flash from across its Copilot experiences (Chat, inline edits, ask and agent modes, code completions), with no action required from users — the models will disappear automatically from the selectors. Users are encouraged to switch to Gemini 3.1 Pro (Preview) and Gemini 3.6 Flash. On the administration side, Enterprise organizations must verify that their model policies allow these replacements before their teams encounter them again in the Copilot Chat selector on VS Code and github.com. This deprecation illustrates the rapid pace of model rotation in Copilot, where previous-generation Gemini versions give way to Google’s newest iterations.
npm restricts GAT tokens from bypassing 2FA
31 July — GitHub is limiting what npm granular access tokens (GAT) configured to bypass two-factor authentication (2FA) can do. Until now, a leaked token of this kind could be used to take full control of an account: create other tokens, add a maintainer, change a package’s access rights. From now on, these sensitive operations require an interactive 2FA challenge, even for a token normally exempt from 2FA. The measure takes effect immediately; a second step is planned for January 2027, when these tokens will also lose the ability to publish packages directly, pushing automation teams toward trusted publishing (OIDC-based). Only npm tokens are affected — classic GitHub tokens (PAT, GitHub App, Actions) continue to work unchanged. This is a direct response to the wave of npm package takeover attacks seen in recent months across the open source ecosystem.
Self-repository syntax for referencing actions in GitHub Actions
30 July — GitHub Actions is introducing a new syntax, $/, for referencing an action or reusable workflow located in the same repository as the calling workflow. Until now, this kind of internal reference either required using ./ paired with an explicit checkout step, or hard-coding a version. With $/, the reference automatically points to the exact commit being run, with no extra checkout step, and works everywhere ./ was used: workflow steps, composite actions, nested composition, and reusable workflow calls. This new feature requires runner version 2.336.0 or later. Its main benefit: it lets organizations enforce policies requiring actions to be pinned to full commit SHAs while keeping internal references automatically synchronized with the ref being executed.
OpenAI removes GPT-5.4 and GPT-5.4 mini from ChatGPT as of August 31
31 July — OpenAI announced that GPT-5.4 and GPT-5.4 mini will be removed from the ChatGPT interface for users signed in with an account, effective August 31. Both models will remain accessible via the OpenAI API and in authenticated Codex sessions using an API key — a distinction that affects ChatGPT app/web users, but not developers using the API or Codex with an API key. This announcement continues the GPT-5.6 (Luna/Terra) pricing update at the end of July, which speeds up the shift of the ChatGPT lineup toward the newest models.
🔗 Announcement — @OpenAIDevs on X
ImageGen in Codex CLI: new lightbox and canvas
30 July — OpenAI Devs has updated the ImageGen feature built into Codex CLI with a new lightbox-style interface and a dedicated canvas. This change aims to simplify exploring and refining images generated directly within the Codex workflow, without changing context. It is an ergonomics improvement rather than a new model, but it directly affects the CLI tool that OpenAI is tracking as a priority.
🔗 Announcement — @OpenAIDevs on X
Briefs
- Amp — orbs start 20% faster, Puck assistant switched to GPT-5.6 Sol — Amp announces that orbs (remote machines) start about 20% faster and switches its Puck assistant to GPT-5.6 Sol following very critical user feedback. 🔗 source
- Hugging Face deploys a free public endpoint for DeepSeek-V4-Flash-0731 — A Hugging Face employee puts up, on the day of release, a free tokenless, OpenAI-compatible endpoint to test the model. 🔗 source
- Sakana AI — 5th place in the OSINT Diver 2026 CTF — A Sakana AI defense/intelligence team, using an OSINT agent based on Fugu-ultra 1.1, finishes 5th out of more than 850 participating teams. 🔗 source
- Kimi K3 on Together AI — lowest price and best cache hit rate according to OpenRouter — Data from the OpenRouter dashboard show that Together AI offers one of the lowest prices for Kimi K3, with one of the best prompt cache success rates. 🔗 source
- An open video model (Marlin-2B) makes 370 hours of public archives searchable for about 10 dollars — A community member indexes 1,864 films from the Prelinger Archives with Marlin-2B on Hugging Face Jobs, generating 23,148 searchable moments. 🔗 source
- GitHub surveys developers on their use of accessibility tools — Survey open until August 31, 2026, among developers with disabilities about their use of accessibility technologies. 🔗 source
- Seedance 2.5 also announced on Luma — Luma teases the upcoming arrival of Seedance 2.5 on its platform, with no date specified. 🔗 source
- Ideogram Object Remover available on Runware — Ideogram’s object removal tool claims the best FID score on the OmniEraser benchmark, with no prompt required. 🔗 source
- SGLang adds support for Inkling-Small on NVIDIA DGX Spark cluster — Measured throughput of 24 tokens/second at concurrency 1, on 2 DGX Spark systems connected in ConnectX-7. 🔗 source
- ChatGPT desktop — Activity view and Voice shortcut via the animal icon — New Activity view bringing together pending conversations and project updates, plus a shortcut to ChatGPT Voice. 🔗 source
- Ten advances in mathematics and theoretical computer science — OpenAI shares ten results on open problems in geometry, cryptography, and complexity; detailed content unavailable, source blocked to robots. 🔗 source
- Advancing responsible AI in Europe — OpenAI presents its safety, transparency, and provenance practices in support of AI governance in Europe in the face of the AI Act; detailed content unavailable, source blocked. 🔗 source
- Univé builds an AI-ready workforce with ChatGPT Enterprise — Case study of Dutch insurer Univé on the rollout of ChatGPT Enterprise; detailed content unavailable, source blocked. 🔗 source
- OpenAI says it dismantled a Cambodia-based scam network — Use of ChatGPT for investment, romance, gambling, and identity theft scams; detailed content unavailable, source blocked. 🔗 source
- avatarin builds a 24/7 retail agent with GPT-Realtime — Multilingual customer support for Yamada Denki: 30,000 users in two weeks, 92% positive feedback according to the official summary; detailed content unavailable, source blocked. 🔗 source
What this means
Media generation is accelerating its push toward broad distribution and conversational editing. MiniMax H3 leads the way with a unified text/image/video/audio model and open weights promised “in the next few days,” shipped day-0 with around fifteen partners. The same movement can be seen at Ideogram (P-Image-Ideogram, four quality modes and pricing starting at $0.003 per image), Wan Video (Realtime mode without a “generate” button), and Luma (Layers, editing one layer at a time through conversation): the added value is no longer based only on the raw quality of a model, but on iteration speed and multi-platform integration from launch. Black Forest Labs, with a 48-hour FLUX 3 preview reserved for a third-party agent, takes the opposite approach — controlled teaser rather than immediate general availability.
Open model inference is changing scale, and usage is being documented as much as the technology. Together AI claims more than 10,000x growth in its token-serving volume in one month, driven by enterprise migration toward open models, and alongside that publishes a detailed technical article on the limits of classic autoscaling metrics for LLM inference. In addition, the Ai2 study built on infini-gram sheds unexpected light on this same wave of open models: based on self-published books, it documents how close AI-generated texts still are to the formulations encountered during training — a reminder that inference scalability does not erase the deeper questions about the originality of generated content.
Tooling for code agents and supply-chain security are advancing in parallel, not always in the same direction. Devin expands its execution scope to any computer with Outposts, Replit multiplies autonomy features (Follow-up Tasks, Model Selector), and GitHub Actions simplifies the referencing of internal actions with the $/ syntax. At the same time, GitHub is tightening npm security by imposing an interactive 2FA challenge even on tokens configured to bypass it, in direct response to the wave of account takeovers seen in recent months — a sign that the growing automation of agents is accompanied by a parallel hardening of access controls.
Two players are tightening access to some of their models, for different reasons. OpenAI is removing GPT-5.4 and GPT-5.4 mini from ChatGPT for signed-in users starting August 31, while keeping them available via API — a lineup clarification rather than a pure removal. Google, for its part, is reserving Gemini 3.5 Flash Cyber, an economical model dedicated to detecting and fixing software vulnerabilities, for governments and trusted partners only: a sign that some labs are now explicitly distinguishing between restricted defensive-use models and public-facing models.
Sources
- MiniMax — H3 announcement on X
- MiniMax — Follow-up thread
- Runway — Grok Imagine Video 1.5
- Alibaba Wan — Realtime mode
- Luma — Layers
- Black Forest Labs — FLUX 3 preview
- Ideogram — P-Image-Ideogram
- Cognition — Devin Outposts
- Replit — This Week in Replit
- Together AI — Growth in tokens served
- Together AI — Dedicated inference autoscaling
- Ai2 — infini-gram study on books
- GitHub — Deprecation of Gemini 2.5 Pro / Gemini 3 Flash
- GitHub — Restriction on npm 2FA tokens
- GitHub — Self-repository syntax
- OpenAI — GPT-5.4 removal
- OpenAI — ImageGen in Codex CLI
- Amp — Orbs and Puck
- Hugging Face — DeepSeek-V4-Flash-0731 endpoint
- Sakana AI — OSINT Diver 2026 CTF
- Kimi K3 on Together AI
- Marlin-2B — Prelinger Archives
- GitHub — Accessibility survey
- Seedance 2.5 on Luma
- Ideogram Object Remover on Runware
- SGLang — Inkling-Small on DGX Spark
- ChatGPT desktop — Activity view
- OpenAI — Ten advances in mathematics
- OpenAI — Responsible AI in Europe
- OpenAI — Univé
- OpenAI — Disrupting a scam network
- OpenAI — avatarin