I

Inkling

Thinking Machines Lab · Released Jul 2026

Conditional

Inkling is the strongest reason yet to take open-weight multimodal AI seriously. It delivers competitive reasoning and coding with best-in-class open-weight safety and controllable thinking effort, all under Apache 2.0. But Thinking Machines is honest that it's not the strongest overall — factuality lags (SimpleQA 43.9%) and self-hosting demands serious hardware (2TB VRAM). Conditional: a compelling base for organizations that need to own and fine-tune a multimodal model, not for teams wanting peak off-the-shelf performance.

Is it right for you?

Good for

  • Open-weight (Apache 2.0) multimodal base — full weights downloadable, modifiable, and deployable on your own infrastructure
  • Safety-conscious deployments — best open-weight FORTRESS scores (78.0% adversarial, 95.9% benign refusal)
  • Token-efficient reasoning via controllable thinking effort — uses 1/3 the tokens of Nemotron 3 Ultra for equal Terminal Bench performance
  • Domain-specific fine-tuning on Tinker — proven uplift in financial reasoning with Bridgewater Associates
  • Native multimodal input (text, image, audio) — rare among open-weight models at this capability level

Not good for

  • Pure factuality and knowledge recall tasks — SimpleQA Verified only 43.9%, well behind closed models at 70%+
  • Budget-constrained self-hosting — BF16 checkpoint requires ~2TB aggregate VRAM (8x B300 or 16x H200)
  • Maximum raw coding leaderboard performance — SWE-bench Pro 54.3% and Terminal Bench 63.8% trail closed frontier

How it performs by task

Code generation

Very Good

Solid SWE-bench Verified 77.6%, competitive with open-weight leaders; trails closed frontier on harder repos

Reasoning

Very Good

AIME 2026 97.1% is near-saturated; HLE 29.7% shows limits on hardest reasoning vs. 53%+ from closed leaders

Multimodal understanding

Very Good

Strong vision (MMMU Pro 73.5%, Charxiv 82%) with rare native audio input; limited to text output

Agentic tool use

Very Good

MCP Atlas 74.1% and BrowseComp 77.1% show capable agentic behavior across diverse tool ecosystems

Safety and refusal

Excellent

Best-in-class among open-weight on FORTRESS; calibrated uncertainty instead of hallucination on ambiguous queries

Factuality

Fair

SimpleQA 43.9% is a clear weakness vs. closed models at 70%+; factual reliability needs fine-tuning

Pricing

Input

$1.87 / 1M tokens

Output

$4.68 / 1M tokens

Context

Up to 1M tokens native

View full pricing

Benchmarks

BenchmarkScoreSource
SWE-bench Verified77.6% Source
AIME 202697.1% Source
GPQA Diamond87.2% Source
HLE (text only)29.7% Source
Terminal Bench 2.163.8% Source
MCP Atlas74.1% Source
MMMU Pro (Standard 10)73.5% Source
IFBench79.8% Source
SimpleQA Verified43.9% Source

No verdict changes yet

The clock starts day one — changes land here as our verdict evolves.

Verification log

  • Pricing— No changes

    Automated agent

How we evaluate