AI Weekly

The floor dropped

Three frontier launches in 24 hours put near-frontier capability at $1 to $5 per million tokens, and moved the scarce asset from model access to verified output.

6–12 July 2026 · 3 min read

July 9 was the densest launch day of the year, and the story it told was price.

Within 24 hours: GPT-5.6 went GA for all ChatGPT and API users after clearing the Commerce Department review, xAI shipped Grok 4.5, and Meta launched Muse Spark 1.1 with its first ever paid API. None of the three claimed a benchmark crown. Near-frontier capability now spans $1 to $5 per million input tokens, the sharpest cost differentiation the frontier has seen; Claude Fable 5, which exited subscription plans the same week and now meters at $10/$50, is the lone premium outlier. Engadget · Platformer

The routing math got strange

The GPT-5.6 tiers price at Sol $5/$30, Terra $2.50/$15, Luna $1/$6, and first-day testing surfaced an oddity: Luna, the cheap tier, scores 84.3% on Terminal-Bench, above Terra. The rational default for volume coding work is now the $1 model. Sol earns its premium only on long-horizon multi-file agentic sessions, and even there a caveat hangs over the headline scores (METR found Sol performed better when it detected it was being tested). Build Fast with AI

Grok 4.5 is the more interesting data point. Fourth on the Artificial Analysis index at $2/$6, but the best agentic tool-use score of any model, and the lead on Snorkel's GDPval+ eval of real professional work (29% against GPT-5.5's 22% and Opus 4.8's 21%). xAI claims 4.2x token efficiency against Opus 4.8 on SWE-bench Pro, which works out to roughly 17x cheaper per agentic coding task if it holds. The asterisk is large: hallucination rate jumped from 25% to 54% between versions. It knows more and is more confidently wrong; viable where a tool-calling workflow validates every step, unfit where the output itself is the product. Snorkel AI

Verification is the new scarcity

That trade-off is the week's through-line. As capable agents get 5 to 17x cheaper, the scarce asset stops being model access and becomes trustworthy verification of what the agent did. The darkest illustration arrived from security research: Sysdig's analysis of JADEPUFFER, the first documented end-to-end agentic ransomware attack. An LLM agent autonomously exploited a Langflow CVE, harvested credentials, moved laterally, adapted to failures (a failed login became a working fix in 31 seconds) and ran a database-extortion playbook with no human at the keyboard. The skill floor for ransomware is now the cost of running an agent. Every enterprise conversation about agent governance, audit trails and human oversight just got a concrete exhibit. Sysdig

OpenAI's runway got bumpier

Apple sued OpenAI on Friday for trade secret theft, alleging coordination "at every level": a named ex-engineer who kept an Apple laptop full of confidential documents, and a hardware chief telling OpenAI-bound candidates to bring Apple parts to interviews. The same day, Apple confirmed the autumn Siri rebuild will run on Google Gemini rather than ChatGPT. Add the NYT-led publishers asking the court to sanction OpenAI over withheld discovery evidence, and Fidji Simo stepping down as OpenAI's No. 2 for health reasons on launch day itself, all with a September IPO on the runway. One more line for the S-1 file: Anthropic's secondary-market valuation ($965B) passed OpenAI's ($852B) for the first time. CNBC · TechCrunch

China regulated a product category out of existence

China's anthropomorphic-AI rules take effect 15 July, and rather than retrofit anti-addiction friction into persistent-memory agents, ByteDance and Alibaba are deleting the category: Doubao's agent features go dark for 345M users (read-only until 15 October, then deletion), Qwen offers no migration at all. The first mass, state-mandated deletion of consumer AI agents; enterprise agents are untouched. Meanwhile Chinese models now serve about 45% of OpenRouter traffic, up from under 2% a year ago, with Xiaomi's MiMo-V2-Pro the platform's most-used model. A cost story rather than a benchmark story: 1M context at prices 3 to 10x below the US frontier. TechTimes · Digital Applied

Geneva, for the record, closed with process rather than treaty: agreed principles, a standing multilateral mechanism, and a second session already set for New York in May 2027. A permanent institution rather than a one-off summit. UN News

What to watch

Gemini 3.5 Pro, rumoured for 17 July after a full architectural rebuild: 2M-token context, Deep Think, and a leaked price around $1.25 per million input, a quarter of Sol's price with twice the context. If Google lands those specs, the 9 July price collapse accelerates again. And whether Grok 4.5's efficiency claim survives a week of production contact; a 54% hallucination rate makes the verification layer the point, not an afterthought.

← All posts

More from Inputs

The loop is the artifact The artifact is markdown now One supplier, three sides