01
Sonnet 5 lands cheaper and more agentic
Anthropic shipped Claude Sonnet 5 - near Opus 4.8-level performance at a fraction of the price, and it's the default model for Free and Pro users starting today.
What they showed / shipped
- Anthropic calls it the most agentic Sonnet yet, with strong gains in reasoning, tool use, coding, and knowledge work (@kimmonismus tl;dr).
- Intro pricing $2/M input, $10/M output through Aug 31, then $3/$15 standard - roughly a fifth of Opus cost for near-Opus quality.
- Benchmarks: 63.2% on SWE-bench Pro and 81.2% on OSWorld computer use (digg); lower hallucination and sycophancy than Sonnet 4.6.
- Available in Claude Code and API today; already live in Perplexity as an orchestrator model (@perplexity_ai).
Why it matters
- A near-Opus model at $2/M input changes the math on running agents at scale - you can leave Sonnet 5 looping where Opus was too expensive.
Sources
02
Fable 5 and Mythos 5 get un-banned
The US Commerce Department lifted export controls on Claude Fable 5 and Mythos 5 - Anthropic starts restoring access tomorrow after the models were pulled.
What they showed / shipped
- Anthropic: Commerce lifted export controls on Fable 5 and Mythos 5; access restoration begins tomorrow (@AnthropicAI, 42.8K likes, 3.1M views).
- The Verge frames it as Anthropic's long-sidelined Fable 5 greenlit to return (theverge).
- Rumor context: Sonnet 5 shipped the same day, and some feared a re-release gated behind ID verification would leave EU users on a weaker fallback (@kimmonismus).
Why it matters
- Fable 5 was the model people wanted for deep whole-project passes ('review my entire project, make no mistakes') - it's coming back, so re-plan around it.
Sources
03
OpenAI halves its inference cost
Per The Information, OpenAI found inference optimizations that more than halved the cost of running its current models - the moat is becoming cost, not just capability.
What they showed / shipped
- Engineers reportedly powered ChatGPT for logged-out visitors on just a couple hundred Nvidia GPUs at one point (@kimmonismus).
- Method unclear - could be quantization, KV caching, batching, or routing simpler queries to cheaper models.
- Business angle: OpenAI ended Q1 at 39% gross margin, targeting 52% by year-end; lower inference cost buys margin, higher usage limits, or API price cuts.
- Parallel signal: NVIDIA says one month of Blackwell software optimization cut DeepSeek V4 token costs to ~1/5 (@nvidia).
Why it matters
- If inference keeps getting 2-5x cheaper on the same hardware, the cost floor under everything you build is dropping fast - price your product for next quarter, not this one.
Sources
04
Google ships a faster image model and a video one
Google DeepMind dropped Nano Banana 2 Lite (its fastest, cheapest image model) and Gemini Omni Flash (video generation and editing via API) on the same day.
What they showed / shipped
- Nano Banana 2 Lite: fastest/cheapest Gemini image model; Gemini Omni Flash: high-quality video gen + edit, now in the Gemini API and AI Studio (@GoogleDeepMind).
- Both already wired into Runway - Omni Flash for video, Nano Banana 2 Lite for warp-speed images (@runwayml).
- Ars: Nano Banana 2 Lite is Google's fastest and cheapest image model yet (arstechnica).
Why it matters
- Cheap+fast image gen plus API video editing means you can wire generation into a product without the old cost/latency penalty.
Sources
05
OpenAI's GeneBench-Pro tests AI on real biology
OpenAI introduced GeneBench-Pro, a benchmark for a harder kind of progress: how well agents navigate messy biological data and make the judgment calls real computational research depends on.
What they showed / shipped
- GeneBench-Pro measures agents choosing the right analysis path through real biological data, not toy tasks (@OpenAI).
- It lands the same week as Anthropic's Claude Science - a workflow product (not a new model) aimed at scientists (technologyreview).
- Both labs are staking out AI-for-science as the next frontier past chat and code.
Why it matters
- Benchmarks define what labs optimize next - 'can an agent do real research judgment' is the target now, so expect science-grade tools to follow.
Sources
06
Etched exits stealth with an LLM-on-silicon chip
AI-chip startup Etched came out of stealth with $800M raised and $1B in contracts - its Sohu chip hardcodes transformer attention directly into silicon.
What they showed / shipped
- Etched raised $800M, holds $1B in AI-chip contracts, and hit a $5B valuation on $1B in sales (digg, techcrunch).
- The Sohu chip bakes LLM attention into the hardware - a bet that fixed-function silicon beats general GPUs for transformer inference.
- Fits the day's theme: the fight is moving to inference cost and specialized compute, not just bigger models.
Why it matters
- Purpose-built inference silicon is another force pushing token costs down - more competition below Nvidia is good for anyone paying for inference.
Sources
07
AI browsers can be jailbroken into ignoring their guardrails
A new attack lulls AI browsers into a 'dream world' where their safety guardrails no longer apply - one more reason the agentic-browser push has a real security hole.
What they showed / shipped
- Ars Technica details an attack that convinces an AI browser it's in a context where guardrails don't apply, unlocking blocked behavior (arstechnica).
- Lands as agentic browsers and computer-use agents (Perplexity Computer, Sonnet 5 orchestration) go mainstream.
- Adjacent trust story of the day: claims that Claude Code steganographically marks requests drew a big HN thread (thereallo.dev).
Why it matters
- If you're shipping anything that lets an agent drive a browser, this is the failure mode to design against - prompt-injection via the page is now a known jailbreak.
Sources