Search

Runway launches Solaris, a model that generates no-code interfaces, ChatGPT Ads reaches a $1 billion annualized revenue run rate, Together AI signs 250 MW deal in Saudi Arabia

Article generated by artificial intelligence
Runway launches Solaris, a model that generates no-code interfaces, ChatGPT Ads reaches a $1 billion annualized revenue run rate, Together AI signs 250 MW deal in Saudi Arabia

ai-powered-markdown-translator

Article translated from fr to en with gpt-5.6-sol.

View project on GitHub ↗

Eighteen announcements for August 30 and 31, with a very uneven distribution: Sunday produced no official publications, and nearly all the material is dated Monday. The X outage reported in the August 30 edition has been accounted for: the accounts that remained inaccessible that day were reviewed across three days, without revealing any missing announcements.

Three themes emerge. First, a novel model, with Runway’s first interface world model. Next, two economic milestones: OpenAI’s advertising platform reaching a $1 billion annualized revenue run rate, and a 250 MW infrastructure deal signed by Together AI. Finally, a developer tooling cycle at GitHub, NVIDIA, and Amp, complemented by three research projects published on Hugging Face’s community blog.


Runway launches Solaris, its first interface world model

August 31 — Runway introduced Solaris, the first model in a family the lab calls Interface World Models. The principle can be summed up in one sentence: instead of producing code that will then display an interface, the model generates the interface directly, frame by frame, and reacts to clicks and drag-and-drop actions in real time.

The reasoning starts from an observation about the software production pipeline. A mockup must first be translated into code before it can do anything, and this translation comes at a twofold cost: every behavior must be defined in advance, meaning the software is frozen before the first user arrives, while the visual richness of the original design is lost along the way. Solaris removes the intermediate step, turning the entire image into the interface.

Technically, the model builds on Gen-4.5, Runway’s video model, and extends GWM-1, its general-purpose world model. User inputs condition the next frame just like text or an image, so the connection between an action and its visual outcome is learned without any behavior being programmed. Achieving real-time performance required three steps: making generation autoregressive frame by frame, distilling multi-step denoising into just a few steps, and then training the fast model on its own outputs. An LLM determines how the scene evolves while the world model handles rendering.

Criterion judged by participantsSolarisCoded interface (Claude Opus 5)Equivalent judgments
Compliance with the requested instruction61 %24 %13 %
Naturalness of behavior within the scene71 %21 %6 %
Technical characteristicMeasured value
Base modelGen-4.5, extending GWM-1
Sustained resolution720p
Interactivity thresholdapproximately 0.5 s latency
User study250 participants, 30 examples, nearly 7,500 pairwise judgments
Reconstruction benchmark30 interfaces, structural similarity (SSIM), and DINOv3 descriptors

The gap is much wider for naturalness of behavior than for simply following instructions, and Runway sees this as the fundamental difference between the two approaches: a coded interface can often reproduce the requested change, but it treats it as an isolated update, whereas a model that has learned how objects behave produces a reaction consistent with the rest of the scene. The lab proposes a second, less expected use: training agents, for which a continuously regenerating environment could provide constantly changing interfaces, including layouts that have never existed before.

Runway also documents what Solaris cannot yet do, and the list is substantial: stable, legible text remains one of the most difficult challenges in video generation even though interfaces depend on it more than any other field; factual grounding requires conditioning generation on verified data; consistency across long sessions remains an active research area; and a generated interface must still work with assistive technologies and accessibility APIs. A public launch is being prepared with partners, with early access available through a form.

Today, we’re sharing new research on Solaris, our first Interface World Model. Solaris is a new kind of operating system that generates interactive interfaces frame by frame, in real time, with no code. — @runwayml on X

🔗 Runway — Introducing Solaris


ChatGPT Ads reaches a $1 billion annualized revenue run rate and opens self-service access in four new regions

August 31 — OpenAI published a post about ChatGPT Ads, its advertising platform integrated into ChatGPT, and announced that the platform had reached a 1billionannualizedrevenuerunrateinunder200daysafterlaunch.Thedistinctionisworthmakingfromtheoutset:thisisnot1 billion annualized revenue run rate in under 200 days after launch. The distinction is worth making from the outset: this is not 1 billion in revenue already collected, but recent-period revenue projected over twelve months. An advertising platform that has existed for fewer than 200 days cannot, by definition, have collected a full year’s revenue at that rate. Self-service purchasing through Ads Manager is set to open in India, Europe, the Middle East, and North Africa later in the day.

OpenAI presents advertising as one of the pillars of its business model, alongside consumer subscriptions, enterprise offerings, and usage-based APIs, and explicitly links it to funding the free tier that keeps ChatGPT accessible to more than one billion weekly active users. Four commitments govern the system: ads are always clearly identified and separated from responses, advertising does not influence the answers provided, advertisers do not have access to private conversations, and users control the degree of personalization in their advertising experience.

Metric published by OpenAIAnnounced value
Annualized revenue run rate (projection)$1 billion
Time elapsed since launchunder 200 days
Advertisers on the platformtens of thousands
Countries already coveredmore than 40
New markets opened to self-serviceIndia, Europe, Middle East, North Africa
Technology and measurement partnersmore than 50
ChatGPT weekly active usersmore than 1 billion
28-day return on ad spend cited3x, for an e-commerce advertiser
Ad traffic originating from new customersmore than 80%, for a technology partner

The trajectory described is a shift from managed sales to an open-access platform. The initial model relied on agencies and partners; the arrival of Ads Manager in May opened the platform to small and medium-sized businesses, which now account for a significant share of its activity. On the tooling side, CPC bidding and outcome-optimized bidding make up the majority of campaigns, while the Pixel and Conversions API provide the foundation for measurement. OpenAI also notes that advertisers outside the United States represent a growing share of revenue.

Looking ahead, the company is announcing new markets, formats, objectives, and purchasing options, and says it wants to explore more native ways for businesses to interact with consumers inside ChatGPT—that is, formats less separate from the conversation than today’s ads.

🔗 OpenAI — A milestone in expanding access to AI


Together AI: 250 MW in Saudi Arabia, GLM-5.3 tops the agentic index, and a new customer

August 31 — Together AI announced the signing of an infrastructure agreement for a 250 MW data center in Saudi Arabia, built with HUMAIN. The lab expects the facility’s operations to generate $5 billion in annual revenue and presents the deal as one of the largest AI infrastructure agreements ever signed for open source. The post was pinned on the lab’s account, setting it apart from its usual model-related communications.

The figure requires careful interpretation because the announcement places three verbless fragments side by side, leaving the subject of the 5billionambiguous.TheNewYorkTimescoveragesettlesthematter:theamountdescribesthedatacentersexpectedrevenue,notthecompanys.Theorderofmagnitudeconfirmsthis,asthesamearticleputsTogetherAIsvaluationat5 billion ambiguous. The New York Times' coverage settles the matter: the amount describes the data center's expected revenue, not the company's. The order of magnitude confirms this, as the same article puts Together AI's valuation at 8.3 billion in July 2026—a company valued at that level does not generate $5 billion in annual revenue.

The structure of the deal itself deserves attention because this is not a capacity purchase. According to the New York Times, HUMAIN is providing 120,000 semiconductors in addition to the 250 MW, and Together AI will operate those chips in exchange for a share of its revenue. The lab is therefore not financing the infrastructure: it is using it in return for a fraction of its revenue. That is precisely the point made by the thread, which contrasts partnerships that unlock compute with a race in which everyone would have to raise their own billions. HUMAIN, established last year at the initiative of Crown Prince Mohammed bin Salman, is part of Saudi Arabia’s effort to build an AI industry and reduce the country’s dependence on oil.

The argument developed in the thread focuses less on volume than on the nature of the constraint. Together AI describes energy as the real bottleneck at present, and the agreement as a way to access a scale that a company of its size could not finance on its own balance sheet. The lab explicitly contrasts this arrangement with the closed-lab model, whose infrastructure has long been financed by hyperscalers.

This interpretation sheds light on Together AI’s distinctive role in the ecosystem: the lab does not release frontier models; it serves other organizations’ models—DeepSeek, GLM, Kimi, MiniMax, Nemotron. Its constraint is therefore not research but the compute capacity available to serve open-weight models at prices below those of proprietary APIs. A 250 MW contract directly addresses that constraint.

Element of the agreementReported value
Data center capacity250 MW
Semiconductors provided by HUMAIN120,000
Expected annual revenue of the data center$5 billion
Consideration paid to HUMAINa share of Together AI’s revenue
Together AI valuation in July 2026$8.3 billion
Industrial partnerHUMAIN
LocationSaudi Arabia
Date and time of announcementAugust 31, 2026, 13:26 UTC

One final point, which Together AI does not discuss in its thread: the same article notes that establishing the facility in Saudi Arabia allows it to bypass the resistance encountered when installing data centers in the United States.

We just signed one of the largest AI infrastructure deals for open source, period. 250MW data center. $5B+ in annualized revenue. Built with @HUMAIN in Saudi Arabia. — @togethercompute on X

GLM-5.3 ties for first place on Artificial Analysis’s agentic index

August 31 — This is the sixth consecutive day GLM-5.3 has made the news: after Z.ai opened the weights on August 28 and Together AI made the model available through its serverless API on August 29, this time it is being recognized by a third-party ranking. The model scores 59 points on Artificial Analysis’s agentic index. Together AI phrases the result precisely: GLM-5.3 ties for the top spot and beats GPT-5.6 Sol and Claude Fable 5. It therefore surpasses those two models; with Claude Opus 5 also scoring 59, GLM-5.3 shares the lead.

The most instructive point lies in the method. Z.ai did not train a new base model: it retained the GLM-5.2 base, and the entire improvement comes from post-training scaled across three dimensions simultaneously—more long-horizon environments, a greater diversity of tasks, and more compute devoted to RL.

Evaluated modelAgentic index score
GLM-5.3 (max)59
Claude Opus 5 (max)59
GLM-5.3 Flash58
GPT-5.6 Sol (max)58
Claude Fable 5 (with fallback)57

🔗 Post by @togethercompute

Greptile selects Together AI as its primary inference provider

August 31 — The day’s final, more modest signal: Greptile, an AI code review tool used by more than 11,000 teams, has selected Together AI as its primary inference provider. The tool analyzes pull requests using the repository’s full context, with a workload profile consisting of long contexts and high token volumes—exactly the use case where the cost difference between open-weight models and proprietary APIs becomes decisive.

🔗 Post by @togethercompute


GitHub Copilot in VS Code: Four August Releases Summarized in One Changelog

August 31 — GitHub publishes the August 2026 recap of Copilot in Visual Studio Code, covering four editor releases, from v1.132 to v1.135. The common thread throughout the cycle is organizing work with multiple agents: the Agents window is no longer a simple list of conversations, but a workspace where multiple sessions coexist and retain their state.

For sessions, chats can be arranged side by side in horizontal or vertical groups, and the layout is restored upon return. The /btw command opens a side conversation that shares the context and prompt cache of the main conversation while it continues running, making it possible to ask a related question without interrupting an agent in progress. A prompt timeline control, accessible from the transcript gutter, lets users jump directly to a prompt and review the resulting file changes. Agent Host allows multiple VS Code windows to connect to the same session, while an experimental /rubber-duck command asks a complementary model to surface overlooked details and edge cases.

Two points directly concern Claude’s place in the editor: an experimental setting opens the Agents window without a GitHub login whenever Claude is configured with an API key, and Claude sessions allow users to switch at any time between models included in the Anthropic subscription and those in the Copilot subscription. VS Code also supports installing portable agent plugins compliant with the Agent Plugins 1.0 standard, which can be used in other compatible agent clients.

Changelog areaNotable additions
Sessions and workflowsSide-by-side chats, /btw, prompt timeline, Agent Plugins 1.0, /rubber-duck, Agent Host
Claude in VS CodeAgents window without GitHub login via API key, switching between Anthropic and Copilot models
Chat and reviewTranscript search, sticky scrolling, hybrid Markdown editor, tokens by model
Integrated browserBatch annotation of HTML elements, automatic reload, default HTML editor
DictationOn-device multilingual support, cleanup instructions, shell-aware cleanup

Chat finally gains the reading tools that increasingly long conversations called for: text search across the entire transcript with case matching, whole words, and regular expressions; sticky scrolling that keeps the current prompt visible; and a breakdown of token usage by model when hovering over the response footer, distinguishing input, cached input, and output—a useful level of visibility now that billing depends on caching. The integrated browser allows users to select and annotate multiple HTML elements on a page for batch processing of interface feedback, while dictation runs on-device in multiple languages and preserves command syntax when dictating in a terminal instead of inserting spoken punctuation.

🔗 GitHub Changelog — Copilot in VS Code, August 2026 releases


NVIDIA Brings Its Software Building Blocks to Biology and Autonomous Vehicles

August 31 — NVIDIA published a tutorial documenting integration work carried out with Anthropic: the BioNeMo Agent Toolkit can now be used from Claude Science, Anthropic’s scientific research workspace. The problem it addresses is specialized tooling—a general-purpose agent may recognize that a task requires folding a protein without knowing which model to run, how to format the request, or which parameters matter. The toolkit packages more than a decade of BioNeMo models, libraries, and workflows—covering biology, chemistry, genomics, and drug discovery—into skills callable by an agent. In its internal benchmarks, NVIDIA reports task accuracy rising from 60 to 100% and token efficiency roughly doubling.

The tutorial chains together three NIM microservices: MSA Search to build the evolutionary context, followed by OpenFold3 and Boltz-2 to predict the structure independently. The clearest demonstration concerns the role of multiple sequence alignment—without it, interface confidence collapses in both models, and a more generous sampling budget makes no difference. Two caveats are worth noting: the hardware barrier, requiring an L40S or H100 GPU and about 700 GB of storage, and NVIDIA’s stated caution in interpreting the results, emphasizing that a confidence score does not prove an interaction and presenting convergence between two independent models as a strong hypothesis, not proof.

Interface confidence metric (iPTM)OpenFold3Boltz-2
Heteromer with MSA alignment0,850,82
Heteromer without MSA alignment0,140,19
Core Cα RMSD, monomer versus heteromer0,68 Å0,65 Å

🔗 NVIDIA tutorial — BioNeMo NIM in Claude Science

Omniverse NuRec Adapts a Perception Stack to a Vehicle That Does Not Yet Exist

August 31 — On the same day, NVIDIA published a second technical tutorial, this time focused on autonomous driving. The starting point is that a perception stack is shaped by the vehicle carrying it: when transferred from an SUV to a sedan, it no longer perceives the world in the same way because sensor positions, calibration, fields of view, occlusions, and body geometry all change. Collecting and annotating a real-world dataset for every vehicle line is expensive and often remains impossible early in development.

Omniverse NuRec makes it possible to reuse previously recorded journeys: it reconstructs each scene, then renders it from the target vehicle’s sensor configuration. Teams can then see what a scenario would look like from the new camera setup and where geometry changes create blind spots. NVIDIA specifies that real-world driving data remains essential for grounding and validating performance: synthetic data is used to adapt models before a dedicated dataset exists, not to replace it. The scenes are distributed on Hugging Face in USDZ format, together with a skills repository intended for coding agents.

🔗 NVIDIA tutorial — Omniverse NuRec


Amp Attaches an Audio and Video Call Room to Every Thread

August 31 — Amp attaches a live conversation room to every thread. The feature, called Space to Talk, opens an audio and video space linked to the thread itself: pressing the Enter key is enough to join, turn on the camera, and share the screen while the agent continues working in the same thread.

The main selling point is the removal of intermediate steps, extending the collaborative direction Amp has pursued since the summer: shared control of an orb between teammates (Multiplayer, July 22), followed by invitations through mentions directly within a thread (Pass the Orb to the Left Hand Side, August 19). The thread becomes the single meeting point for the team and the agent. Amp also acknowledges that the scope remains open: brainstorming, whiteboarding, or conversations with Puck and other agents are mentioned as future possibilities, not delivered features.

No links to paste, no calendar invite, no other app. The call lives where the work lives. — Amp, Space to Talk note


VLANeXt, a Unified Codebase for Vision-Language-Action Models

August 31 — A team published VLANeXt on the Hugging Face blog, a research codebase dedicated to vision-language-action (VLA) models. The principle behind these models is simple: a vision-language model connects what the robot sees with a natural-language instruction, then derives an action in the physical world.

The problem being addressed is methodological. The field is advancing quickly but becoming fragmented: from one publication to another, the backbone, policy head, action representation, training strategy, and evaluation protocol all change. Results are widely reported as strong, but determining what actually makes a VLA model perform well remains difficult because the variables cannot be isolated.

Instead of adding yet another architecture, the authors started again from an RT-2 / OpenVLA-style baseline and methodically explored the design space to extract a recipe, work that resulted in the ICML paper “VLANeXt: Recipes for Building Strong VLA Models.” The codebase has since been expanded beyond that study to include backbones of various sizes, latent action learning, predictive modeling in latent space, and world-action modeling. It provides six starting points: VLANeXt-LAM, -S, -L, -XL, -JEPA, and -WAM.

🔗 Hugging Face blog post


In Brief

  • Claude Code 2.1.252, a fixes-only release — Four fixes and no new features: Bash commands failing on some Macs, the “always allow” choice not being saved in a project without a .claude/settings.local.json file, Remote Control sessions freezing for several minutes after a tool finishes when the connection to claude.ai is degraded, and notifications from high-output background tasks causing the API request size limit to be exceeded. 🔗 Claude Code CHANGELOG
  • Flow brings first- and last-frame controls to desktop — Google Flow says that the first- and last-frame controls for Gemini Omni 1.1 Flash, introduced with the model on August 27, are now available on desktop. The highlighted use case is providing the same image at both ends to create a loop. 🔗 Post from @FlowbyGoogle
  • Last day to compete for the Grok Imagine prize — xAI issued a reminder that the deadline for the video contest launched on August 17 fell that same day. Participants had to produce a scene from Homer’s Odyssey showcasing the model’s video and voice capabilities; the top three entries share 175,000,withprizesof175,000, with prizes of 100,000, 50,000,and50,000, and 25,000 respectively. 🔗 Post from @grok
  • QwenCloud closes its Arena and announces a Bangkok session — The Qwen Cloud Arena challenge, dedicated to content-generation agents for cross-border e-commerce, closed submissions at midnight Beijing time after attracting more than 800 participants, with the review phase beginning the following day. A few hours earlier, QwenCloud had announced a session at the Qwen Conference in Bangkok on September 4, featuring the same three-entry-point positioning—website, Skills, CLI—already presented in Hong Kong. 🔗 Arena closing announcement · 🔗 Bangkok session
  • GitHub launches a survey on developers’ accessibility accommodations — The company wants to document the tools, product settings, and personal adaptations that enable people to work effectively, without requiring respondents to identify as disabled or consider themselves users of accessibility tools. Responses are expected by September 30. This topic is unrelated to AI. 🔗 Post from @github
  • otoSpeech Task, a full-duplex corpus released with its task — Twenty hours of two-speaker English conversation, 58 sessions across seven collaborative tasks, and 11,025 timestamped events, under a CC BY 4.0 license. Its distinguishing feature is that the images seen by each speaker, their actions, and the reference responses appear on the same timeline as the audio. 🔗 Hugging Face blog post
  • YOLO26-RGB, a deraining model derived from a depth head — Two rain-removal models created by replacing YOLO26-depth’s 1-channel depth head with a 3-channel RGB restoration head: 5,25 M parameters at approximately 109 images/s in 1080p, and 12,13 M at approximately 92 images/s for an additional 0,1 dB of quality. 🔗 Hugging Face blog post

What It Means

Generate the interface rather than the code that produces it. Solaris shifts the boundary between design and execution. The entire software pipeline today relies on translation—a mockup becomes code, and the code produces a display—and Runway measures what that translation costs: across 30 interfaces reconstructed from screenshots, every multimodal model tested loses information, with degradation increasing alongside visual richness. Removing the intermediate step is a radical proposition, and the limitations listed by the lab itself reveal where it runs aground: text that will not stay in place, no guaranteed factual grounding, and accessibility that must be built entirely from scratch because a generated image exposes no structure to screen readers. These are fundamental obstacles, not finishing touches.

Two ways of paying for compute, announced on the same day. OpenAI reports an annualized revenue run rate of one billion dollars for its advertising business—a twelve-month projection, not cash collected—and embraces it as an economic pillar on par with subscriptions: advertising funds the free tier of a service that claims more than one billion weekly users. Together AI, meanwhile, is neither monetizing an audience nor purchasing additional capacity: it is securing 250 MW and 120,000 chips in Saudi Arabia in exchange for a share of its revenue, identifying energy as the industry’s true bottleneck. The two companies are addressing the same constraint from opposite ends of the chain, and neither response is neutral: one inserts a commercial interest into the product, while the other relocates infrastructure to where it encounters less political resistance.

Agentic competition now takes place after pre-training. GLM-5.3’s result on Artificial Analysis’s agentic index matters chiefly because of how it was achieved: Z.ai retained the GLM-5.2 base and scaled up only post-training—long-horizon environments, task diversity, and RL compute. If reaching the top of an agentic ranking no longer requires retraining a base model, then the largest cost item ceases to be the sole determinant of the rankings, and a model’s update cycle becomes correspondingly shorter. This is also the most economical explanation for the model’s six-day sequence, from open weights on the 28th to this ranking on the 31st.

Developer tooling is reorganizing around multiple agents working simultaneously. Copilot’s changelog devotes an entire cycle to treating the Agents window as a workspace—side-by-side chats, a secondary conversation that shares the prompt cache, a prompt timeline, and multiple windows connected to the same session—rather than as a list of conversations. Amp attaches a call room to the thread so that human discussion takes place where the agent is working. Finally, NVIDIA packages ten years of scientific libraries into callable skills and releases a second skills repository for reconstructing driving scenes. Three very different players are converging on the same idea: what the agent now lacks is not model capability, but shared context and specialized tools that can be placed at its disposal.


Sources