Skip to content

MIT-Licensed Ornith-1.5 models achieve self-improvement breakthrough, rival Claude Opus

MIT-Licensed Ornith-1.5 models achieve self-improvement breakthrough, rival Anthropic, OpenAI
Share this article

Ornith AI announced the release of Ornith-1.5, a family of open-source language models that represents a fundamental shift in how AI models are trained. For instance, rather than relying on a fixed set of human-curated tasks and manually designed harnesses, Ornith-1.5 continuously generates new training tasks, discovers effective strategies for solving them, and improves the policy through reinforcement learning.

The self-improvement loop that changes everything

The training cycle proceeds in three stages, and it goes like this:

  • Given an environment or codebase, high-level instructions, and access to the model’s previous task-solving history, the system proposes progressively harder tasks that expose capability gaps. 
  • The model then generates or refines a task-specific scaffold: the instructions, tools, decomposition strategy, and orchestration used to approach the problem. 
  • Finally, the policy produces a solution rollout, and reward is propagated across all three stages.

Ornith AI scores generated tasks on validity, difficulty, and novelty. A proposed task must be executable and verifiable, sit near the model’s current capability frontier, and differ enough from previous work to really teach it something new, adding useful training signal.

The target solution success rate is set at 0.2, favoring tricky tasks the model usually fails while preserving enough successful rollouts to make the reinforcement learning actually stick.

Ornith AI has released Ornith-1.5, a family of MIT-licensed models spanning 9B to 397B parameters that learn by generating their own training tasks, achieving performance on par with Claude Opus 4.8 on key coding and reasoning benchmarks.
Source: Ornith AI

Three models, three deployment strategies

Ornith-1.5 spans three model scales: 397B Mixture-of-Experts, 35B MoE – activating 3B parameters per token – and 9B dense. All released under the Massachusetts Institute of Technology License for unrestricted commercial and research use. 

The flagship 397B model achieves 86.1 on Terminal-Bench 2.1 and 56.0 on DeepSWE, matching Claude Opus 4.8 at 85.0 and 59.0, respectively, while outperforming GLM-5.2 and DeepSeek-V4-Flash-0731 across the board.

Ornith AI has released Ornith-1.5, a family of MIT-licensed models spanning 9B to 397B parameters that learn by generating their own training tasks, achieving performance on par with Claude Opus 4.8 on key coding and reasoning benchmarks.
Source: Ornith AI, 397B model

Now, the 35B MoE model significantly outperforms similarly sized peers like Qwen 3.6-35B, and despite activating only 3 billion parameters per token, it successfully beats dense models like Meta’s Muse Glimmer-30B and Gemma 4-31B by wide margins on agentic coding benchmarks: 68.5 versus 51.7 and 42.1 on Terminal-Bench 2.1.

Ornith AI has released Ornith-1.5, a family of MIT-licensed models spanning 9B to 397B parameters that learn by generating their own training tasks, achieving performance on par with Claude Opus 4.8 on key coding and reasoning benchmarks.
Source: Ornith AI, 35B MoE model

Remarkably, the edge-deployable 9B model delivers strong results, achieving 47.0 on Terminal-Bench 2.1 and 70.6 on SWE-Bench Verified, matching or exceeding much larger models.

Ornith AI has released Ornith-1.5, a family of MIT-licensed models spanning 9B to 397B parameters that learn by generating their own training tasks, achieving performance on par with Claude Opus 4.8 on key coding and reasoning benchmarks.
Source: Ornith AI, 9B model

The mobile AI revolution: 9B model runs on your phone

Quantized versions such as GGUF, MLX, FP8, and NVFP4 are available from day one, enabling deployment on mobile devices.

Ornith-1.5-9B-Mobile marks a giant step toward bringing frontier-level AI onto devices we use every day. The quantized 9B model can run on iPhones and Android devices through Ollama, AtomicChat, and LM Studio, delivering 70.6 percent on SWE-Bench Verified from a smartphone.

This compression without catastrophic degradation means developers can now build genuinely capable AI assistants that operate entirely offline, keeping things private while matching or beating those giant, cloud-based models. It’s honestly a game-changer for what our devices can do on their own.

About The Coin Headlines

The Coin Headlines strives to bring trust into crypto media. At a time when every soundbite and headline can move the markets from red to green and vice-versa, The Coin Headlines promises to bring verified, credible and timely news and analysis from the world of crypto, blockchain, Web3, tech and markets. Founded in 2026, The Coin Headlines is based in the UAE with a team of experienced journalists and editors covering breaking news and updates from around the world.

From covering the biggest events to interviewing some of the most popular KOLs in the industry, The Coin Headlines keeps you informed of the latest trends and insights.

At The Coin Headlines our focus is clear: Real-time news updates, market movements, whale transfers, macroeconomic trends, tech and AI and geopolitical breaking news. The news we report goes through a strict editorial audit before its published to ensure the readers only get verified and credible information. We realize the world of crypto is dynamic, volatile, and many times, confusing. At The Coin Headlines we break down these complex issues into simple articles which cater to not just the experienced trader but also the student and first-time investor who wants to understand the space before committing to it.