Search

Microsoft's ThinkingBox benchmark arrives on OpenEnv with 20 runs per task, the Exgentic dataset gathers 10,000 traces in OpenTelemetry format

Article generated by artificial intelligence
Microsoft's ThinkingBox benchmark arrives on OpenEnv with 20 runs per task, the Exgentic dataset gathers 10,000 traces in OpenTelemetry format

ai-powered-markdown-translator

Article translated from fr to en with gemini-3.8-flash-high.

View project on GitHub ↗

Overnight from Saturday to Sunday, Microsoft and Hugging Face connected ThinkingBox to OpenEnv: this benchmark replays 507 business tasks 20 times each and evaluates agents on the final state of the database, where repetition alters the rankings. On Sunday, October 4, three authors described Exgentic Agent LLM Traces, 10,000 agent runs in OpenTelemetry format published under the Hub’s Exgentic organization. In brief, Claude Code 2.1.289 and Copilot CLI 1.0.92-4 strengthen their guardrails, Feyn introduces MultiMatte, a matting model built on Meta’s SAM 3, and ChatGPT Enterprise manages its plugins in the Admin console as of October 1.


Microsoft and Hugging Face connect ThinkingBox to OpenEnv: 20 runs per task, and the rankings shift

October 3 — In a blog post presented as joint between Microsoft and Hugging Face, Tuhin Kundu (Microsoft) describes ThinkingBox, a benchmark that judges AI agents on what they leave behind in the database, not on their responses. Built by the Microsoft Copilot Studio team with Toloka, it is not new (research paper from August 20, dataset published in late August): what is new is its arrival behind OpenEnv, the open agent environment interface hosted by Hugging Face, and the publication of results for 18 models.

Its 507 synthetic business workflows are replayed 20 times each from a clean state, and deterministic judges compare the final state of the database to the expected one. Results according to the authors; their costs, estimated from OpenRouter list prices (Anthropic’s for Claude Opus 5.5), are a comparative indicator, not an invoice:

Evaluated modelOverall pass@1Tasks passed 20 out of 20 times (out of 507)Cost per reliable task
Claude Opus 5.567.16%2417.80 dollars
Claude Opus 566.50%24113.30 dollars
GPT-5.465.36%1286.80 dollars
GPT-5.6 Sol61.91%829.76 dollars
GPT-6 Astra58.31%2317.45 dollars
Kimi-K3 (open weights)57.37%6820.68 dollars

Repetition changes the rankings: Kimi-K3, the best open-weights model, solves 476 of the 507 tasks at least once, but only 68 on every one of the 20 runs, compared to 241 for both Claude Opus 5 and Claude Opus 5.5. According to the team, about four out of five failures are due to tool use rather than reasoning. The code is under the MIT license, and the data under CDLA-Permissive-2.0.

A trajectory is a claim. Database state is the evidence. Repetition is the trust test. — Microsoft and Hugging Face post

🔗 ThinkingBox-Bench dataset · ThinkingBox in OpenEnv


Exgentic Agent LLM Traces: 10,000 agent runs preserved in OpenTelemetry format

October 4 — On the Hugging Face community blog, Lena Dankin, Elron Bandel, and Michal Shmueli-Scheuer present Exgentic Agent LLM Traces, a dataset published under the Hub’s Exgentic organization: 10,000 agent runs and 242,000 model calls, each preserved according to OpenTelemetry GenAI conventions, along with the score and cost of each run.

These runs cover six benchmarks (SWE-bench, BrowseComp Plus, AppWorld, and three variants of τ²-bench) and up to five agent designs. On average, Claude Opus 4.5 consumes 2.4 million tokens per run, compared to 552,000 for GPT-5.2.

Two caveats: the five models are from earlier generations (DeepSeek-V3.2, Kimi-K2.5, GPT-5.2, Claude Opus 4.5, and Gemini 3 Pro), and the post presents as a release of the day a repository created on June 10, with no declared license in its card.

🔗 Blog post · Dataset on Hugging Face


In Brief

  • Claude Code 2.1.289 — This 27-entry release fixes four cases where deny and ask permission rules were not applied; a user-installed plugin can no longer overwrite the connection tool descriptions of a managed MCP server, and 18 other entries cover mods and plugins. 🔗 source
  • Copilot CLI 1.0.92-4 and Copilot SDK 1.0.17-preview.4 — The CLI, in preview (latest stable: 1.0.91), adds the copilot config subcommands to list, read, set, and delete a setting, and no longer passes the ambient GITHUB_TOKEN to sandbox shells unless explicitly configured; the SDK receives the OnSubagentStart and OnSubagentStop hooks to enrich a sub-agent’s initial prompt, then inspect or replace its response. 🔗 CLI · 🔗 SDK
  • Feyn’s MultiMatte — This matting model is text-guided: name the object to keep, and the model removes the rest, with an alpha matte that better renders hair and blurred areas. Built on Meta’s SAM 3 and fine-tuned with a LoRA adapter, it achieves an S-measure of 0.901 on DIS-VD according to the authors, compared to 0.667 for SAM 3; its weights, under Apache 2.0 according to their card, have been online since August 20. 🔗 source
  • ChatGPT Enterprise plugins in the Admin Console — On October 1, the Enterprise and Edu release notes announced that owners and admins of ChatGPT Enterprise workspaces now manage plugins and their marketplaces in the Admin Console, on the Plugins and Marketplaces pages; apps remain in Workspace settings > Apps, with existing permissions. 🔗 source

What This Means

ThinkingBox changes the question asked of an agent: no longer whether it can accomplish a task, but whether it accomplishes it every single time. On a single trial, Kimi-K3 is within one point of GPT-6 Astra; over twenty, it succeeds every time on only 68 tasks, compared to 231. For anyone entrusting an agent with operations that modify a database, this second figure is what counts, and the benchmark measures it on the final state of the database rather than what the agent claims to have done.

Cost follows the same logic. When evaluated only against tasks passed on every trial, GPT-5.4 emerges as the cheapest (6.80 dollars per reliable task), and not GPT-5.6 Sol, which is nonetheless the cheapest per successful trial (0.127 dollars); Kimi-K3, meanwhile, costs 20.68 dollars per reliable task, three times more than GPT-5.4. As the authors point out, these amounts remain a comparative indicator rather than an invoice.

The authors of Exgentic Agent LLM Traces start from a related limitation: a benchmark reduces a run to a score, whereas a trace shows the chosen tools, context growth, and error recovery. With each call preserved in a standard format, they see a way to debug an agent against a baseline, replay a step with another model, or test inference infrastructure under a real agent workload; one team is already doing so with inference-perf, the Kubernetes SIG benchmark tool, and its results are scheduled for later release.


Sources