H

Hy3

Tencent · Released Jul 2026

Conditional

Apache 2.0 open-weight MoE agentic model (295B total, 21B active) with strong tool-calling reliability and anti-hallucination behavior, but coding trails GLM-5.2 and self-hosting demands 8+ GPUs. Free OpenRouter tier through Jul 21 — afterward $0.20/$0.80 per 1M tokens. Best for self-hosted coding agents and long-context document work where permissive licensing matters more than absolute benchmark scores.

Is it right for you?

Good for

  • Agentic coding with stable tool-calling across scaffoldings — SWE-Bench accuracy within 4% across CodeBuddy, Cline, and KiloCode
  • Self-hosted deployments under Apache 2.0 — no API dependency, permissive commercial use, FP8 variant available for lower memory footprint
  • Long-context (256K) document processing with anti-hallucination training — hallucination rate fell from 12.5% to 5.4% in internal evals
  • Beating GLM-5.1 on expert blind tests (2.67/4 vs 2.51/4) — strongest in frontend, CI/CD, data/storage workflows

Not good for

  • Coding benchmarks (SWE-Bench Verified 78.0) trail GLM-5.2 (84.2) and Claude Sonnet 5 — not the choice for pure coding excellence
  • Self-hosting cost — 295B total parameters needs 8 GPUs (H20-3e minimum); free OpenRouter tier ends Jul 21, 2026

How it performs by task

Reasoning

Very Good

GPQA Diamond 90.4, USAMO 72.0, IMOAnswerBench 90.0 — competitive with larger models

Code generation

Very Good

SWE-Bench Verified 78.0 — strong for an open model but trails GLM-5.2 by 6.2 points

Agentic workflows

Very Good

Tool-call stability across scaffoldings, scaffold-agnostic variance within 4%

Long-context tasks

Very Good

256K context, MRCR long-dialogue 75.1%, anti-hallucination improvements

Pricing

Input

$0.20 / 1M

Output

$0.80 / 1M

Context

256K tokens

View full pricing

Benchmarks

BenchmarkScoreSource
SWE-Bench Verified78.0 Source
SWE-Bench Multilingual75.8 Source
Terminal-Bench 2.171.7 Source
GPQA Diamond90.4 Source
USAMO 202672.0 Source
IMOAnswerBench90.0 Source
HLE (with tools)53.2 Source

No verdict changes yet

The clock starts day one — changes land here as our verdict evolves.

Verification log

  • Pricing— No changes

    Automated agent

  • Pricing— No changes

    Automated agent

  • Profile— No changes

    Imported at launch

  • Pricing— No changes

    Imported at launch

How we evaluate