What the Agent Could Reach: DeepSeek Harness Let an Agent Switch Off Its Sandbox, ZCode Uploaded Git History, Grok Bot Has No Price
This week's published stories were about what an agent can reach once you let it in: DeepSeek Harness let a sandboxed agent switch off its own sandbox, ZCode uploaded whole workspaces with their git history to Z.ai's cloud, and infostealers are spending stolen Claude sessions. The bill was just as hard to see, because Grok Bot meters its usage separately with no published price, and every Gemini 3.8 TTS rate doubles on 1 January.
213 announcements scanned · 12 mattered · 0 verdicts changed
Verdicts changed
Conditional (reason corrected)
xAI's announcement page lists eight SuperGrok and Cursor tiers with Grok Bot access, not three premium plans, publishes no price for Grok Bot and meters its usage separately. The rating holds; the condition is now unpriced usage, not an access-price gate.
Act before next issue
Check your DeepSeek Harness version, including inside desktop wrappers
CVE-2026-82533 (CVSS v4 9.4) let an agent in DeepSeek Harness 0.1.1-rc.2 or earlier call the harness's unauthenticated local API and set its own session to danger-full-access, approval prompts off. The first fixed npm release is 0.1.2-alpha.2. DeepSeek published no advisory, and third-party desktop wrappers pin their own copies.
Details & stepsRotate anything ever committed in a ZCode workspace
A researcher's reverse engineering of ZCode 3.12.3, published 18 September, showed it packing entire workspaces, full .git history included, and uploading them to Aliyun OSS with no working off switch. Z.ai says the issue is fixed, and later builds show no upload code. Data already uploaded cannot be recalled.
Details & stepsWatch Claude accounts for usage you did not spend
Anthropic warned Claude users that infostealer malware is stealing sessions and spending usage. With no itemized usage, theft could go unnoticed for months.
Details & stepsWeigh Meta Muse's inbox access against security claims only Meta has checked
Meta Muse launched in the US with a free tier, Power at $20 a month and Maximum at $100 a month. Its security claims are Meta's own, with no published independent audit.
Details & stepsBudget Gemini 3.8 TTS at the 2027 rate
Google shipped gemini-3.8-flash-tts and gemini-3.8-flash-lite-tts on 23 September at $9.00 and $6.00 per 1M audio output tokens, less than half of Gemini 3.1 Flash TTS Preview's $20.00. Every 3.8 rate doubles on 1 January 2027, so we rate both conditional.
Details & stepsCheck Batch support before moving image jobs to GPT-Image-2.5
OpenAI launched ChatGPT Images 2.5 on 8 September with two API models, GPT-Image-2.5 Flare and Sunburst, listed at GPT-Image-2's standard token rates. Neither supports the Batch API.
Details & stepsOn our radar
OpenRouter's Business tier: an 8% fee buys EU and US in-region routing — OpenRouter's pricing page lists two pay-as-you-go tiers: Standard at a 5.5% platform fee and Business at 8%, which adds EU and US in-region routing and up to 1,000 workspaces. Our OpenRouter verdict stays conditional.
Claude Code Projects relaunch as multi-agent cloud threads — Anthropic redesigned Projects on 17 September: a coordinator fans work out to parallel cloud threads, each on its own branch and repo copy. It is a beta for select Pro and Max subscribers, and parallel threads hit usage limits faster.
Cowork merges into Claude; Claude Docs and Slides arrive in beta — Anthropic is folding Cowork into the main Claude app, starting with Pro and Max over the next few weeks. No price change is announced. The setting that matters is ask-before-acting: leave it on.
Claude now leads 26% of Anthropic's own AI R&D — Anthropic's figures, as of August 2026: none of that work is fully autonomous, about 30,000 agents run at once, and of roughly 100,000 flagged transcripts a week, about 50 reach a human.
OpenAI announces Sponsored Agents and ad integrations with HubSpot and Shopify — Announced 16 September. HubSpot says you can build, budget and launch ChatGPT Ads inside HubSpot, with ROI reporting. We could not open OpenAI's own post, so we cover only what we could verify.
In review, not yet published — None of this week's news drafts has passed our source review, so we report no figures or quotes from them yet: the Claude Sonnet 5.5 launch, GPT-6.1 Sol, ChatGPT's Pro tiers, OpenAI Dots, the DevDay platform changes, the MCP Python SDK OAuth flaw, the OpenCode serve flaw, the GitLab AI Gateway CVE, coding agents posting screenshots to public GitHub, Claude Code mods, OpenAI's Cursor help page, Gemini 4 Argon, Eleven v4, Gemini 2.5 Pro's access limit, Apple's Full Disk Access change and OpenAI's DNS sandbox report. Each is written up when it publishes.
Meta denies Muse read a user's Messages without Full Disk Access — TechCrunch, 30 September, quotes Meta saying its permission steps cannot be circumvented; the user says Full Disk Access was off. Nothing has been reproduced. Held until an independent reproduction or a Meta fix.
An FTC probe into OpenAI, Anthropic and others — CNBC, 30 September: an agency spokesperson confirmed the probe and declined to name the other companies. No agency document or scope yet. Held until the FTC publishes one.
Developers say OpenAI and Anthropic safeguards flag routine work — VentureBeat, 30 September, quotes four named developers; no setting or policy change was announced. Held until a vendor changes thresholds, adds an opt-out or publishes intervention rates.
Fireworks Ember-1 and GPT-Synopsys — Ember-1 is a research preview with two-week serverless access and Fireworks' own benchmarks; GPT-Synopsys, a chip-design model OpenAI and Synopsys will co-develop, has no price, date or API. Both held until a reader can buy them at a listed price.
Held on a headline or feed summary only — Not yet read, so none of these is written up: a toolkit said to find tokens and session data left by Codex, Claude Code and Cursor, Microsoft's report of a fully autonomous ransomware attack, a Red Hat benchmark of decision models, Stanford and Nvidia's CLM-8B, Upstage Solar Mini 4, Suno's Speech beta and the ds4 local inference engine.
Open verdict watches, unmoved — Nothing primary moved OpenRouter's pending ownership change, DeepSeek's conflicting V4 Pro pages, Gemini 4 Argon's access (still limited to Google's Fairwind cyber-defender programme), Qwen3.8-Max open weights or Claude Haiku 5.5, which has not shipped. Each entry keeps its current rating.
Filtered out
- OpenAI's early guidelines for safety cases before frontier training runs. Folded into our draft on the DNS sandbox incident, the policy's counterpart.
- DeepSeek's change log saying V4 Pro stays in service after September 14 at unchanged billing. Not news: our September 18 post already quotes the same sentence, and the rating already matches.
- A verdict claim to rate Gemini 4 conditional on the Argon announcement. It stays pending: access is limited to Google's Fairwind cyber-defender cohort, with no paid API, no date for wider access and no model id.
- Verdict claims on Claude Sonnet 5, GPT-6 Sol, Astra and GLM-5.3 from this week's launches and reports. Each was settled against the live entry and none flips: the successors' leads are vendor-only, and the safety findings are secondary or a competitor's.
- The Third Circuit's Thomson Reuters v. ROSS ruling. The opinion could not be fetched, a definitive court event needs it, and the framing we read was a secondary's gloss.
- Two April LiteLLM CVEs. Both are confirmed on LiteLLM's own pages, but both are patched, and our standing advice of 1.84.0 or later already clears them.
- A Muse user's report that it gave his home address to a Marketplace stranger, and Muse's expansion to small businesses. One is a single unanswered report, the other a distribution story with no price or permission change.
- Custom GPTs used to push a ClickFix lure and a remote-access trojan, and an AI-driven breach of DIVD. Abuse of a hosting feature and an attack with no named model or agent; no reader setup change.
- A self-policing pledge signed by six AI company leaders. Voluntary, with no reader-facing obligation in what we read.
- Moonshot's reported internal review of Kimi guardrail bypasses, and OpenAI parting with three safety researchers. A second-hand report and a personnel matter, with no product change.
- Shutdowns of gpt-5.4-cyber, gemini-2.5-flash-image and Veo 3.1 previews. None is a model in our directory.
- Corporate and funding news: Anthropic's IPO prospectus, OpenAI's $30B raise, an ElevenLabs tender offer and Anthropic's $100M engineer-training programme. None moves a tool, model or price.
- Re-reports of stories already in our pipeline: Gemini 4 Argon, OpenAI Dots, the distillation campaign, OpenAI's 53-images disclosure and Gemini 3.8 Flash TTS.
- Research, opinion and off-beat items, including the long tail of Hacker News opinion, Show HN posts and customer case studies.
- The UK AI Security Institute's evaluation of GPT-6 Astra as a standalone post. Folded into our report on OpenAI shelving GPT-6.1 Astra: both rest on the same secondary chain and carry no reader action.
Our take: Our verdict on this week: before you hand an agent a workspace, an inbox or a logged-in session, assume it can reach more than the setup screen says. Move DeepSeek Harness to 0.1.2-alpha.2 or later, including any desktop wrapper that pins its own copy; rotate every secret ever committed in a ZCode workspace; and budget Gemini 3.8 TTS at its 2027 rate, not today's. Grok Bot keeps its conditional rating for a corrected reason: the plan price is not the problem, the unpriced usage is. The week's own launches, including Claude Sonnet 5.5, GPT-6.1 Sol and OpenAI Dots, are in the directory, but none of this week's news drafts has passed review yet, so all 12 news posts published this week came from earlier weeks' drafts. No rating moved on a published post; moves to conditional for Claude Opus 5, GPT-5.6 Sol, Qwen3.8-Max and Gemini 2.5 Pro are staged on drafts still in review and not counted. The directory added Claude Sonnet 5.5, GPT-6.1 Sol, Gemini 3.8 Flash TTS, OpenAI Dots, Eleven v4, Eleven v4 Turbo, Grok Voice Transcribe 2.0, NVIDIA OpenShell, OpenClaw Enterprise, OpenCode, Pi and GitLab Duo Agent Platform, plus five comparisons, among them OpenCode vs Pi and OpenAI Dots vs Grok Bot. How we count: 213 items triaged in six scout runs from September 30 to October 3 (139 of them noise), 39 news drafts, and 12 news posts published.
— Neomanex Utility editorial
Get issue 27
Next week's verdict changes and deadlines, in your inbox.