Search

ESPnet triples its open speech corpus with YODAS v3, Mistral serves GLM 5.3 through its API, JEV-27B and OpenDecider claim parity with Jev

Article generated by artificial intelligence
ESPnet triples its open speech corpus with YODAS v3, Mistral serves GLM 5.3 through its API, JEV-27B and OpenDecider claim parity with Jev

ai-powered-markdown-translator

Article translated from fr to en with gpt-6-sol.

View project on GitHub ↗

ESPnet releases YODAS v3, an open speech dataset containing 1.1 million hours in more than 100 languages. It is three times larger than the first version, and all audio is distributed at 48 kHz. Mistral now serves Z.ai’s open-weight model GLM 5.3 through its API, with 1 million tokens of context and a 99.5% SLA. Meanwhile, JEV-27B and OpenDecider, two open decision models, claim parity with Jev on benchmarks run by their authors. MiniMax launches M3.1-Flash-Preview without benchmarks or API pricing. The briefs also cover a Copilot CLI preview that reads Claude Code rules and Space Bunny Alpha, a free stealth model on OpenRouter.


ESPnet releases YODAS v3, 1.1 million hours of open speech at 48 kHz

September 27 — The ESPnet team (William Chen and Shinji Watanabe) releases YODAS v3, a multilingual speech dataset containing 1.1 million hours, three times as much as the first version. All audio is distributed at 48 kHz, with metadata recording the effective sample rate, since much web audio is upsampled.

YODAS v3 featurePublished value
Languages coveredMore than 100, including 34 with over 1,000 hours
Multichannel audioMore than 70% with at least two distinct channels
AnnotationsWord-level timestamped transcripts; English translations for more than half of the multilingual data
Size and license60 TB in OPUS format, CC BY 3.0

According to the authors, this volume is enough to train a Whisper-level model from scratch. The first version, downloaded more than 2.5 million times, was used to train IBM Granite, NVIDIA Canary and Parakeet, and Ai2’s OLMoASR. A related paper appears in the Interspeech 2026 proceedings.

🔗 YODAS v3 on the Hugging Face blog


Mistral makes Z.ai’s GLM 5.3 available through its API, with 1M tokens of context and a 99.5% SLA

September 26 — Z.ai launched GLM 5.3 on August 14 and released its weights on August 28. The news is its arrival on the Mistral API, announced by @MistralDevs. Available as a public preview under the identifier zai-glm-5-3, the model is served unchanged on Mistral’s infrastructure, with 1 million tokens of context, 128k output tokens, more than 100 tokens per second and a 99.5% SLA on the priority tier.

Price per million tokensMistral APITogether AI (August 29)
Input$1.40$1.40
Cached input$0.14$0.26
Output$4.40$4.40

Mistral announced on August 11 that it was opening its platform to third-party models, starting with GLM 5.2. GLM 5.3 arrived in Vibe Code on September 21 as the web app’s default model.

🔗 Announcement by @MistralDevs · Model page


Open decision models: AutoTrust AI’s JEV-27B and OpenDecider take on Jev

Two community posts from September 27 release open decision models under the Apache-2.0 license. These models answer the small, structured questions agents ask—true or false, a choice among several options, or a rating—quickly and with a probability for each answer. They arrive four days after Together AI’s Tev1-4B-experimental. Both are compared with Jev 1.13, TypeSafe AI’s closed, hosted model; the scores below come from the authors’ own measurements.

AutoTrust AI’s JEV-27B: parity with Jev 1.13 on average, according to its authors

September 27 — AutoTrust AI, which says it has no connection to TypeSafe AI, releases JEV-27B, a student of Jev 1.13 distilled on an open corpus. Its training recipe keeps Qwen3.8-27B frozen and adds a detachable block of 108.9 million parameters, trained in about 9.2 hours on a B200. The base model remains intact, scoring 78.0% on HumanEval. Scores measured by the authors on September 26:

Decision benchmarkJEV-27BJev 1.13 (API)
JevBench (public dataset)88.7087.18
Kev83.7585.52
OpenJev text73.8972.96
Nimble92.9191.84
VitaminC77.4678.46
MASSIVE-en87.7187.14
Average across six84.0783.85

Its median latency is 137 ms per decision on a B200. The authors acknowledge that parity holds only on average and that the student inherits Jev’s blind spots; they plan to submit it to independent leaderboards.

🔗 AutoTrust AI’s post

OpenDecider: a 400-million-parameter nano model ahead of Laya and Jev on typed-decisions

September 27 — Manjunath Janardhan releases OpenDecider, a family of decision models: a nano model of about 400 million parameters built on the Ettin encoder, a 4-billion-parameter small model built on Qwen3-4B, and MLX versions for Mac. They are distilled from two open teachers, Qwen3-235B and DeepSeek V4.1 Flash, for less than $30 in GPU costs.

Model evaluatedtyped-decisions (2,000)Unseen general decisions (200)Calibration error
OpenDecider-nano0.7960.6800.092
OpenDecider-small0.6720.7350.087
Laya (specialized checkpoint)0.7660.5700.162
Jev 1.13 (API)0.7540.7300.164

According to the author’s measurements, the small model matches Jev on unseen decisions, with nearly half the calibration error. Jev remains ahead on Laya’s application suite (0.774), particularly on phishing (0.90 versus 0.63). The predictions and a reproducible benchmark are published; the results come from a single seed.

🔗 OpenDecider post


MiniMax launches M3.1-Flash-Preview in MiniMax Code and the Token Plan, without benchmarks or API pricing

September 27 — MiniMax introduces M3.1-Flash-Preview as its latest text model, in preview. It debuts in the morning in MiniMax Code, the company’s coding agent, then arrives that evening in the Token Plan, MiniMax’s platform subscription:

MiniMax-M3.1 Flash Preview is now live on the Token Plan! Faster, lighter, and built for teams running high-volume, latency-sensitive workloads, now available under your existing Token Plan subscription, no extra setup required. — @MiniMax_AI on X

MiniMax has published no benchmarks, API pricing, model size, context window or weights: its model release notes still end on July 31.

🔗 Debut in MiniMax Code, on X


Briefs

  • Copilot CLI reads Claude Code rules — In preview version 1.0.89-5, released on September 27 (no stable 1.0.89 has been released yet), Copilot CLI supports Claude Code rule files stored in .claude/rules as custom instructions. The release notes specify neither the recognized format nor the order of precedence. 🔗 source
  • mistral CLI 0.7.0 — Released without an announcement on September 26, this version of Mistral’s command-line tool adds mistral apps deploy, which deploys an application to the platform through production; mistral apps publish, which publishes it to the Vibe catalog; and mistral evals for offline evaluations. 🔗 source
  • Perplexity’s Agent API moves to GPT-6 — Catching up on September 25: the low and medium presets move from openai/gpt-5.6-luna to openai/gpt-6-luna, and the high preset from openai/gpt-5.6-sol to openai/gpt-6-sol. Prompts, reasoning effort, tools, token budgets and step limits remain unchanged. 🔗 source
  • Space Bunny Alpha on OpenRouter — Catching up on September 23: OpenRouter lists a free stealth model whose developer is unnamed, with 1 million tokens of context and text, image and video inputs. The third-party provider operating it may retain prompts and responses but does not use them for training. No official benchmark is available. 🔗 source
  • A 4B judge for financial figures — Tony Esposito releases FinCalc-NLI, a synthetic dataset with labels calculated by code across ten banking metrics, and openjev-fincalc-4b, a judge fine-tuned in 40 minutes on an A100 to verify numerical claims. It scores 94.4% across 4,000 cases, versus 63.2% for the best open judge tested. 🔗 source
  • Fine-tuning a model on a 16 GB Mac — Across ten LoRA fine-tuning runs, Harshit Kumar Gupta measures PyTorch MPS at 2.2 to 5.7 times the speed of Apple MLX, but limited to about 1.5 billion parameters, while 4-bit MLX fits a 3-billion-parameter model. The LoRA rank and optimizer differed between the two stacks. 🔗 source

What it means

Open data remains the quiet infrastructure behind open models. According to its authors, the first version of YODAS helped train models from IBM, NVIDIA and Ai2. Version 3 triples the volume, moves to 48 kHz and multichannel audio, and retains the CC BY 3.0 license—enough, they say, to train a Whisper-level model from scratch. Today’s decision models follow the same pattern: JEV-27B learns from an open corpus, OpenDecider distills two open teachers for less than $30, and both are released under Apache-2.0.

An open-weight model can be sold as a catalog entry, with hosting providers competing on service. GLM 5.3 has the same price outside the cache at Mistral and Together AI: $1.40 for input and $4.40 for output per million tokens. Mistral highlights cached input at $0.14 versus $0.26 at Together AI, throughput above 100 tokens per second and a 99.5% SLA. Having hosted third-party models since GLM 5.2, Mistral is positioning itself as a hosting provider as well as a model developer. The model also changes behind the scenes at Perplexity: requests through the Agent API presets now run on GPT-6 without the developer selecting that model.

The question is who does the measuring. JEV-27B and OpenDecider claim parity with Jev on benchmarks their authors ran themselves. They do at least publish limitations, predictions or a reproducible benchmark, and their own tables show Jev ahead on several tests, from VitaminC to phishing. At the other end, MiniMax launches a model without any verifiable figures: no benchmark, API pricing or context window. Independent measurement remains to be done in both cases; AutoTrust AI says it plans to submit to independent leaderboards.

At both Mistral and MiniMax, a model first reaches the company’s coding agent: GLM 5.3 became the default model for Vibe Code on the web on September 21 before arriving on the Mistral API, and M3.1-Flash-Preview debuts in MiniMax Code before the Token Plan. Instructions, meanwhile, are beginning to travel between agents: Copilot CLI preview 1.0.89-5 reads .claude/rules files written for Claude Code, so the same repository can guide both agents with the same rules, at least in this preview.


Sources