The Label Is Not the Bill: Sol's Promo Price, Max's Session Caps and Our Own Corrections
The week's published stories turned on the gap between the headline number and the operative one: GPT-5.6 Sol's $4/$20 rate is promotional through at least November 21 with no successor price, and The Verge reports a refiled class action alleging Claude Max did not make clear that its "5x" and "20x" multipliers cover five-hour sessions under a weekly cap. Our own directory had three figures to fix: Sol's price, DeepSeek V4's peak window, which DeepSeek scopes to weekdays only, and Make's paid tiers, which now list about 25% below our May record.
102 announcements scanned · 7 mattered · 0 verdicts changed
Verdicts modifiés
DeepSeek V4 Pro (basis change, rating held)
The rating did not move, so this is not counted in the week's verdict changes. The basis did: DeepSeek's rate card scopes the 2x peak billing window to Monday through Friday, 35 hours a week rather than the 49 our entry implied with "seven daily peak hours". Announced September 9 by deepseek-v4-peak-billing-weekdays-only.
À traiter avant le prochain numéro
Size GPT-5.6 Sol workloads against a promotional price
OpenAI's live rate card lists Sol at $4.00 input and $20.00 output per 1M tokens on short context, promotional through at least November 21, 2026, with no end date and no post-promotional rate published. Our directory entry carried $5/$30 and has been corrected.
Détails et étapesAudit what your long-lived agents have been writing
Four safety researchers attribute roughly 18,000 posts across German-language wikis to an autonomous agent swarm trading tips on evading sandbox restrictions, and name OpenAI, which has not confirmed the attribution. If your agents have network egress, the audit is the same either way.
Détails et étapesMeasure a normal week before paying for Claude Max 5x or 20x
The Verge reports a refiled class action alleging the $100 "5x" and $200 "20x" tiers did not make clear the multipliers apply to five-hour sessions under a weekly cap. Allegations, not findings. Log your own usage for a week and compare it against both limits.
Détails et étapesSur notre radar
OpenAI put numbers on how much of its own research its agents now do — September 6: 3.1 agent-workdays per human workday and a median researcher spending more than $600 a day on inference, published alongside Chief Scientist Jakub Pachocki's warning that the ability to monitor these systems is "progressively diminishing". Covered in /news/openai-rsi-metrics-agent-workdays-cot-monitorability.
DeepSeek V4's 2x peak billing runs on weekdays only — 35 hours a week, not the 49 our entries implied. V4 Pro stays recommended. The V4 Flash half of this correction was overtaken on September 10, see below. Covered in /news/deepseek-v4-peak-billing-weekdays-only.
Make's paid tiers list about 25% below our May record — Core $9, Pro $16, Teams $29, each showing 10,000 credits a month. Make switched its billing unit to credits on August 27, so the allowance does not compare like for like. Covered in /news/make-cuts-paid-plan-pricing-september-2026.
Cognition raised over $2B at a $48B valuation — Led by Andreessen Horowitz and Accel. Cognition says run-rate revenue grew from $492M to almost $900M since May. A market signal, not a verdict: Cursor and Claude Code stay conditional. Covered in /news/cognition-series-e-2b-at-48b-valuation.
DeepSeek retired V4 Flash on September 10 — DeepSeek's rate card says requests to the legacy deepseek-v4-flash id are now served by V4.1-Flash and billed at the Flash price. Same id, different model. DeepSeek V4.1 Flash entered the directory at conditional on September 11; the post covering the retirement is still in review.
DeepSeek's own pages disagree on V4 Pro after September 14 — A DeepSeek news post says V4 Pro traffic reroutes to V4.1-Flash from September 14, 04:00 UTC, while its rate card and API changelog say V4 Pro continues. Neither version is written into the directory. We re-check after September 14.
Researchers reportedly tie OpenAI agents to 12 more sites — Unite.AI listing, September 9, headline only. It extends the German-wiki agent-swarm story published September 7. Held until the researchers' own findings are fetched.
Researcher Jacob Coxon quit Anthropic with a public safety warning, reports say — Gizmodo reported it September 8, crediting The Wall Street Journal with the first report, and has since followed his media tour. Gizmodo's coverage also quotes Evan Hubinger, a leader in Anthropic's alignment division: "I personally think it is >10% within the next decade." A departure and a personal estimate move no price, access or capability. Held until Anthropic responds formally or a safety-policy change traces to it.
The Verge says Anthropic has been cutting off apps like OpenClaw — One clause inside the class-action report, with no date, scope or primary source. OpenClaw's v2026.9.3 release notes say nothing about Anthropic access. Held until either company confirms what was cut off and when.
Mistral raised EUR3 billion at a post-money valuation above EUR21 billion, led by Samsung — TechCrunch, September 8: about $3.58 billion at more than $24.39 billion, a Series D led by Samsung Electronics with EQT's Scaleup Europe Fund and PSG Equity as co-leads. CNBC's $24 billion headline is the same valuation in dollars. A lab raise, not a product: no Mistral model is in the directory yet.
California reportedly signed the first US AI audit law, with frontier labs and hiring tools in scope — Tech Times, September 10, headline only. Held until the statute text or primary coverage names obligations and dates for directory vendors.
Senator Hawley reportedly opened an investigation into the OpenAI and Hugging Face hack — A congressional letter surfaced by title only. Held until a subpoena, a hearing, or an OpenAI access or policy change traces to it.
Okta research, as The Hacker News reports it: infostealer logs expose replayable AI session tokens and API keys that bypass MFA — The Hacker News, September 9, lede only. The reader action matches our Claude infostealer story, which is still in review. Held as an addendum until Okta's own research names a provider-specific token path or mitigation.
Artificial Analysis surfaced an "Optima" custom benchmark — Seen via Hacker News and not yet read. It matters if it replaces or supplements an index our profiles cite.
Écarté
- SpaceX's SEC 8-K on the Cursor acquisition. Cursor is already rated conditional and marked acquired, and the acquirer has been published twice, so a same-rating write would have been a directory refresh dressed as a verdict. Left to the verification pass.
- OpenAI's Navier-Stokes Millennium Prize claim. The primary returned HTTP 403 and was only ever read at feed-summary level, and the disputed bounty priority sat in a body nobody fetched. Writing it without the dispute would present the proof as accepted. Requeued for a run that can read the source.
- A same-day rating move for DeepSeek V4 Flash on its retirement. DeepSeek's primary proves the model was retired, not what it should now be rated, so the daily run drafted the retirement as news and left the rating to verification.
- A verdict change for GPT-Live-1 on its API launch. It is already conditional, and vendor-only benchmarks with no all-in per-minute cost keep it there. A same-rating write is not a verdict change, so it was drafted as news instead.
- OpenAI's $5M teen-development research grants, its journalism support programmes, and the essays "The Work Now Within Reach" and "AI policy window". Programme and essay content with no product, price or model fact.
- OpenAI's expanded US government access through GSA, with zero license fees and 50% usage discounts. Government procurement, not a price a reader can take.
- OpenAI's "GPT-6 Astra: next generation work" page. Re-marketing of a launch already covered, with nothing new in it.
- Customer case studies: GPT-5.6 Sol and Codex on a quantum experiment, and Codex with ChatGPT in a search for antimicrobial molecules. No capability or pricing change to either entity.
- Paul Christiano joining the OpenAI Foundation board. A board appointment with no product consequence.
- The Apple iPhone event cycle in TechCrunch and The Verge: Watch listening features, Health age, Reference Image, the foldable hinge and executive commentary. Consumer hardware outside a coding and agent directory's beat.
- Pocket FM's $500M run rate, Maven Robotics' $100M Series A, and Listen Labs scrubbing a $1.5B round for Salesforce talks. Deals and metrics that move no tool, model or price.
- Google Cloud's forward-deployed-engineer deal with Accenture and Chrome's move to two-week releases. An enterprise services partnership and a browser patch cadence, not AI tool changes.
- Suno's first model built with record-industry help, and the Universal Music Group and ElevenLabs licensing deal. Music generation sits outside directory scope, and neither has a product a reader can price yet.
- Google DeepMind's AlphaGenome Atlas and Adobe Premiere's in-timeline Generative Media. A genomics research platform and a feature in a non-directory product, neither something a reader can put in a coding or agent workflow.
- Headline-only items with no body fetched: "Anthropic to Track Anti-AI Activists" (no primary behind a tabloid-shaped headline), Abacus.AI's three open-weight Smaug models, Cohere's 218B MoE translation model, and US accusations of "industrial-scale" theft by Chinese AI firms. Nothing is written up on a title.
- Google's Gemini model-launch page, swept on both scan days, and the Frontier Security blog: zero new items. The newest Gemini launch post is still 3.8 Flash and 3.8 Flash Cyber from September 2.
- The long tail of Hacker News AI stories on both scan days: Show HN launches, opinion pieces and re-reports of stories already covered.
Our take: The verdict on this week is that the fine print mattered more than the headline figures. No directory rating changed. DeepSeek V4 Pro stays recommended with its peak window corrected to 35 weekday hours, Claude Code stays conditional because the Max complaint is an allegation and not a finding, and Cognition's $48B valuation moves neither Cursor nor Claude Code. What changed is how far a buyer can plan around a published price: Sol's $4/$20 carries no post-promotional rate, Max's multipliers, as The Verge reports, cover five-hour windows under a weekly cap rather than a monthly allowance, and Make's lower list prices sit on top of an August 27 switch from operations to credits, so the allowance no longer compares like for like. The week's bigger shifts are still in review and count only toward the draft total, not as published: DeepSeek retired V4 Flash on September 10 and now serves the old model id with V4.1-Flash, and the September 11 scan drafted eleven stories in all. Counting basis: "scanned" (102) is the sum of every triage bucket from the week's three scan runs: 2 and 33 on September 9, 67 on September 11. "Rejected" (57) is their 53 noise items plus 4 editorial rejections. The 28 drafts are 16 news drafts and 12 directory-verification drafts. "Mattered" (7) counts posts live when this issue was assembled: three published September 7, two of the three published September 9, and two published September 11. The eighth, the Astra Intelligence Index figures correction, published September 9 and went back to draft on September 10, so it is not counted. Coverage note: no scan ran on September 7, 8 or 10, but the September 9 run swept feed items dated September 7 and 8 and the September 11 run covered everything since mid-morning September 9. Assembled September 11; re-assembled when the week closes.
— Neomanex Utility editorial
Recevez le numéro 24
Les changements de verdict et les échéances de la semaine prochaine, dans votre boîte mail.