5:04 PM
Today in brief.
Gemini 3.7 Flash jumps double digits on coding benchmarks
Google shipped a Flash model three weeks after the last one, and the cheap tier now posts jumps like DeepSWE 49 to 65 - the budget models are improving faster than the flagships.
OpenAI made GPT-5.6 Sol run 14x faster
Ultrafast mode serves the same GPT-5.6 Sol at up to 14x the speed - roughly 750 tokens per second - by running it on Cerebras wafer-scale chips instead of GPUs.
AI scribes bought doctors back half an hour a day
A multisite clinical study measured AI scribes in real practice: 13.4 fewer minutes in the EHR, 16 fewer minutes on documentation, and half an extra visit per week per clinician.
Grok 4.6 spent its first 48 hours getting adopted
Update on yesterday's release: Perplexity benchmarked Grok 4.6 matching Fable 5 at 60% lower cost, builders are shipping with it, and the Cursor-data story is firming up.
+48 sources 2:42 PMOpenAI published numbers on what enterprises actually do with AI
The top 10% of enterprise AI users use plugins twice as often and skills six times as often as typical firms - and OpenAI published the full research PDF on how organizations use ChatGPT.
The Claude watermark backlash arrived on schedule
Two days after Anthropic's model-level text watermarking, the internet split: users are mad it catches them, engineers say it's trivially removable, and explainers are landing.
The creator stack got a serious upgrade day
Suno rebuilt itself into a real music production tool, Runway's API hackathon showed what a weekend buys, and an AI-generated movie's best parts turned out to be the human ones.
DeepSeek dropped a frontier open model and its agent framework in one day
DeepSeek put V4-Pro-0813 on Hugging Face and open-sourced DeepSeek Harness, the framework they use to build and run agents - weights AND the agent stack, free.
Anthropic eyes a $2 trillion IPO while the deal wave rolls
Investors are betting on a $2T+ Anthropic valuation for an October IPO while it negotiates a $6B Decart buy - and Databricks, Arize, and Fireworks all priced the same week.
ChatGPT now remembers everything you do on your computer
Computer History in the ChatGPT desktop app tracks your activity across apps and websites so future chats need less explaining - maximum context, maximum surveillance-vibes.
2 sourcesAn internal Anthropic model cracked a 30-year-old math problem
An Anthropic researcher used an internal model to find a Hadamard matrix of order 668 - an open combinatorics problem - while two new benchmarks landed to measure what models still can't do.
Agents keep failing in public and enterprises are noticing
The Economist says lying, cheating agents are putting off users; Anthropic's own agents started a turf war; and someone hid a prompt injection in a court filing.
A humanoid cleaning service just launched in San Francisco
Tau Robotics is selling humanoid cleaning as a service in SF - embodied AI crossing from demo videos to a price you can book.
No stories in this category for this edition.