Thinking Machines Ships Inkling, 975B Open-Weights MoE

Thinking Machines Lab logoThinking Machines LabImportantJuly 16, 2026Models
What happened
Thinking Machines Lab released Inkling, an open-weights 975B MoE model with Apache 2.0 license, native multimodality, and controllable thinking.
Why it matters
Inkling is a customization platform masquerading as a model launch — fine-tuning-ready, broad rather than SOTA, with a self-improvement demo that closed the loop in 27 minutes.
What to do
If you need a customizable base model for specialized workflows, evaluate Inkling on your own data via Tinker. If you need frontier performance out of the box, stick with GPT-5.6 or Fable 5.

Thinking Machines Lab released Inkling on July 15, 2026 — and it's not trying to beat GPT-5.6 (Thinking Machines Lab, 2026(opens in new tab)). Mira Murati's Thinking Machines Lab shipped a 975B-parameter MoE model (41B active) under Apache 2.0 with full weights on Hugging Face (Hugging Face, 2026(opens in new tab)). The real product isn't the model itself — it's the controllable thinking dial and the Tinker fine-tuning platform that ships alongside it.

What happened — Inkling ships

Inkling is a broad generalist Mixture-of-Experts model with native multimodality: text, images, and audio in, text out. It's competitive across agentic coding, reasoning, and multimodal tasks but explicitly not SOTA on any one benchmark.

The headline innovation is a controllable thinking effort dial. Developers can trade performance for token efficiency on the same model — at maximum thinking effort, Inkling matches Nemotron 3 Ultra on Terminal Bench 2.1 at roughly one-third the tokens.

Weights are on Hugging Face in both original and NVFP4 formats for NVIDIA Blackwell systems. Inference endpoints are live through TogetherAI, Fireworks, Modal, Databricks, and Baseten. Thinking Machines also previewed Inkling-Small — 276B total, 12B active — with similar reasoning and agentic performance at lower cost. Full weights drop after testing completes.

On safety, Inkling's FORTRESS adversarial score of 78.0% is the strongest among open-weights models tested. On Design Arena's Agentic Web Dev leaderboard, it scored 1,257 — tied with Claude Opus 4.6 and ahead of Gemini 3.5 Flash (1,254) (Thinking Machines Lab, 2026(opens in new tab)).

Why Inkling matters

Inkling isn't a frontier challenger. It scored 29.7% on HLE (vs 53.3% for Fable 5) and 77.6% on SWE-Bench Verified (vs 95.0% for Fable 5) (Thinking Machines Lab, 2026(opens in new tab)). Those gaps are real and deliberate — Thinking Machines is betting that customization, not raw performance, is the wedge.

The launch demo makes the thesis concrete. Inkling wrote its own fine-tuning job, ran it on Tinker, and converted itself into a lipogram model that never uses the letter "e" — completing the self-improvement loop in roughly 27 minutes (Thinking Machines Lab, 2026(opens in new tab)). That's the pitch: a base model that can specialize itself into your workflow without an external training pipeline.

"We want to make customization accessible for more use cases," the company wrote in its launch post. "Picking the right base model to fine-tune is a qualitative judgment that combines measurable benchmarks with the unique feel of a model."

The Tinker platform is the distribution channel. Inkling is available there with 50% introductory pricing (Thinking Machines Lab, 2026(opens in new tab)), and the self-improvement demo shows the platform can handle real fine-tuning workloads in under half an hour.

What changes for you

If you need a customizable base model for specialized agentic or reasoning workflows, evaluate Inkling on your own data. The Tinker platform is designed to make that evaluation fast — upload your data, run a fine-tuning job, and test the result.

If you need frontier performance out of the box, Inkling is not your model. Stick with GPT-5.6 Sol or Fable 5 for production workloads where every benchmark point matters.

Inkling carries our Conditional verdict: a compelling base for organizations that need to own and fine-tune a multimodal model, not for teams wanting peak off-the-shelf performance.

BenchmarkInklingFable 5
HLE29.7%53.3%
SWE-Bench Verified77.6%95.0%
FORTRESS (safety)78.0%
Design Arena (Agentic Web)1,257

FAQ

Is Inkling free to use?

Yes — full weights are on Hugging Face under Apache 2.0. Hosted inference is available through Tinker (50% introductory pricing), TogetherAI, Fireworks, and others.

How does Inkling compare to GPT-5.6 or Fable 5?

It doesn't — and that's the point. Inkling trails closed frontier models by 20+ points on HLE and SWE-Bench, but its open-weight license and Tinker fine-tuning platform make it a customization base, not a turnkey SOTA model.

Affected tools & models

Never need to catch up again

The weekly delta — only verdict changes and act-now items. No digest filler.

By subscribing you agree to our Privacy Policy. Unsubscribe anytime.