10:45 PM
Today in brief.
Claude Code Auto Mode fell to an 80% RCE chain after Anthropic billed 0.00%
A third-party eval Anthropic commissioned put Opus 5 Auto Mode at zero prompt-injection success. Embrace The Red published a zip-and-module-shadowing chain that hit 60-80%, and Anthropic closed it as working as designed.
The Pentagon put ChatGPT and Grok on GenAI.mil for 1.7 million users
ChatGPT Mil and Starshield AI's Grok for Government went live on the Pentagon's IL5 portal, joining Gemini, with 1.7 million unique users already on the platform out of 3 million personnel.
Runway shipped Solaris, an OS that generates the interface as you use it
Instead of compiling a UI to code, Solaris generates every frame of the interface in real time from clicks, drags and language - a world model that is the app.
Apple got blindsided by AI demand for Mac Mini and Mac Studio
Enterprises are buying Mac minis and Studios to run frontier models locally, OpenAI is reportedly taking them by the tens of thousands, and a memory shortage is emptying the high-end configs.
A French director is shooting an AI movie, and China is already replacing actors
The artifact to watch is 'La Nona Gigante,' a French AI-video series about a boy and the last giant on her island - while a separate report says generated video is slowly displacing actors and livestreamers in China.
Agent memory as a zip of markdown files, and it actually works
Cal Paterson argues the agent-memory products are overbuilt, and ships a portable format that is just markdown pages plus an optional SQLite vector index inside a zip.
PhoneLLM matches GPT-5.6 Terra on voice agents at about 1/14th the cost
Pipecat's open-weights PhoneLLM Alpha 1, a 30B Nemotron fine-tune, scores 72.3% on PhoneBench next to GPT-5.6 Terra's 72.4%, at $0.0025/min versus $0.0347 and a much faster first token.
The EU sent its first AI Act RFIs to frontier labs, four weeks after the rules went live
GPAI obligations became enforceable on August 2. By August 29 the AI Office had mailed legally binding requests on model security, external evals, market monitoring and training-data summaries, with 15M-euro or 3% turnover fines for bad answers.
No stories in this category for this edition.