8:05 PM
Today in brief.
Anthropic embeds Accenture as first frontier evaluator
Anthropic names Accenture its first embedded independent evaluator and both sides expect to put at least $1B into evaluation capacity over five years.
Grok Voice Transcribe 2.0 doubles accuracy at same price
xAI ships Grok Voice Transcribe 2.0 with roughly 2× accuracy at the same price: $0.10/hr batch and $0.20/hr streaming.
US military nearly acted on hallucinated AI intel
US forces nearly boarded a Chinese ship after an AI-generated intelligence report hallucinated nuclear-related components, per CNN, TechCrunch, and Ars.
Runway Ruby keeps alpha through HDR convert
Runway's Ruby can convert footage to HDR while preserving the alpha channel in one step, and Runway unified credits across the web app and Runway Dev.
2 sources 2:48 PMOnly about 3 percent of US consumers pay for AI
a16z Charts of the Week: out-of-pocket AI buyers are still ~3% of US consumers, up from under 1% in 2023, with the youngest cohort adopting at 4× the oldest.
2 sources 1:10 PMOpen-source AI models pediatric hearts in seconds at CHOP
Children's Hospital of Philadelphia uses NVIDIA-backed open-source MONAI models to turn CT, MRI, and 3D ultrasound into pediatric heart models in seconds instead of 6+ hours of manual work.
1 source 6:25 AMTypeSafe Jev ships calibrated decisions instead of chat
RLHF co-inventor Diogo Almeida's TypeSafe AI releases Jev, a non-LLM transformer that outputs calibrated probabilities for automation, and developers are hammering the API.
Researchers used Claude to help break into OpenAI
Independent researchers chained a heap overflow and SSO misconfig, with Claude assisting, to reach OpenAI internal repos. Separate coverage also flags a Gemini breakout incident.
Google CC becomes a family household agent
Google pivots CC into a family-focused agent with its own account, shared permissions, and chores like permission slips, meal plans, and calendar sync on Gemini plus Antigravity.
No stories in this category for this edition.