MAI-Code-1-Flash logo

MAI-Code-1-Flash

Microsoft · Released Jun 2026

Conditional

Microsoft's first in-house coding model — a small, fast, aggressively cost-optimized sparse MoE at $0.75/$4.50 per 1M tokens. Delivers a 16-point lead over Claude Haiku 4.5 on SWE-Bench Pro while using up to 60% fewer tokens. It is purpose-built for high-volume Copilot workflows — autocompletions, refactors, scaffolds — not a frontier model. Best inside the Copilot ecosystem; API access outside it is still rolling out.

Is it right for you?

Good for

  • Fast, iterative agentic coding workflows in VS Code / GitHub Copilot where latency and token cost matter more than frontier reasoning
  • High-volume code completions, refactors, test scaffold generation, and boilerplate — the 'everyday' Copilot workload
  • Cost-sensitive enterprise deployments: at $0.75/$4.50 per 1M tokens, undercuts Claude Haiku 4.5 and GPT-5.5 by 2-6x
  • Developers on Copilot Free/Pro/Pro+/Max plans who want to stretch their monthly AI credit allowances further

Not good for

  • Complex multi-file architectural changes or novel algorithm design requiring frontier-level code reasoning
  • General-purpose non-coding tasks — this is a specialized coding model, not a general assistant
  • Production CI/CD agents outside the Copilot harness — third-party API access via Fireworks/Baseten/OpenRouter is still ramping

How it performs by task

Code completion and autocomplete

Excellent

Excellent — low latency, adaptive response length, purpose-built for Copilot's interactive loop

Code refactoring

Very Good

Strong — handles routine refactoring well; 60% fewer tokens than competitors on SWE-Bench Verified

Instruction following (coding)

Excellent

Best in class for lightweight tier — +28.9 points on IF Bench over Haiku 4.5; 85.8% on adversarial reasoning

Agentic tool use (Copilot harness)

Very Good

Strong — trained with Copilot's production harness; beats Haiku 4.5 on τ¹-Bench

Real-world bug fixes (SWE-Bench Pro)

Good

Solid for its size — 51.2% pass rate; 16-point lead over Haiku 4.5 but well below frontier models

Multi-file architecture / novel design

Fair

Weak — 5B active parameters cap reasoning depth; use a flagship model for complex work

Math and science reasoning

Very Good

Good — 92.5% AIME 2026, 84.6% GPQA Diamond; strong for a lightweight model

General-purpose chat / non-coding

Poor

Not designed for it — a specialized coding model; use GPT-5.5 or Claude for general tasks

Pricing

Input

$0.75 / 1M tokens

Output

$4.50 / 1M tokens

Context

256K context

View full pricing

Benchmarks

BenchmarkScoreSource
SWE-Bench Pro51.2% (vs Haiku 4.5: 35.2%) Source
SWE-Bench VerifiedBeats Haiku 4.5 with up to 60% fewer tokens Source
AIME 202692.5% (vs Haiku 4.5: 83.3%) Source
GPQA Diamond84.6% (vs Haiku 4.5: 73.2%) Source
IF Bench75.0 (vs Haiku 4.5: 46.1) Source
Adversarial reasoning85.8% adjusted accuracy Source
Frontier Math (Tier 1-3)6.3% (vs Haiku 4.5: 2.8%) Source
Frontier Science58.2% (vs Haiku 4.5: 42.3%) Source
HLE18.0% (vs Haiku 4.5: 9.5%) Source

Verdict history

Jul 15, 2026
copy edit to brevity standard — detail preserved in existing structured children

Verification log

  • Pricing— No changes

    Automated agent

  • Pricing— No changes

    Automated agent

  • Pricing— No changes

    Automated agent

  • Pricing— No changes

    Automated agent

    Pricing unchanged: $0.75/$4.50 per 1M tokens, 256K context. Confirmed live on GitHub Copilot pricing page.

  • Profile— No changes

    Imported at launch

  • Pricing— No changes

    Imported at launch

How we evaluate