01
Nvidia's agent scored 100% on ARC-AGI-3 and the harness got the credit
Nvidia AVO cleared every level of an interactive reasoning benchmark with no instructions, and the takeaway everyone landed on is that the scaffolding around the model did the work.
What they showed / shipped
- Nvidia AVO completed all 183 levels across all 25 public ARC-AGI-3 environments, figuring out goals with no instructions, no explicit rules, and no stated objective (NVIDIA AI).
- ARC-AGI-3 is the interactive version of the benchmark - the agent has to act in an environment and infer what winning means, not answer a static puzzle.
- TechCrunch's read: the harness, not the model, is now the real hero - the loop, tools and memory around the model are where the capability came from.
- Same week, open harnesses kept shipping: Seed, a minimal self-modifying agent harness, and Proliferate, a self-hostable Codex for any coding agent.
Why it matters
- If the harness is the differentiator, your leverage is not picking a smarter model - it's the loop you wrap around it. That's the part you can actually control and iterate on daily.
Sources
02
Codex can send your iMessages now
OpenAI wired Codex and ChatGPT into Apple Messages, so an agent can read a thread, draft a reply and send it after you approve.
What they showed / shipped
- Typing
@messages in Codex exposes an Apple Messages workflow - the agent finds the contact, composes the text, and shows an approval card before anything sends (Riley Brown demo).
- Approval is required by default; the agent will not send without the confirm step, and the draft is editable before it goes.
- It also reads back: he pointed it at a group thread and asked it to analyze the conversation and suggest how to communicate better.
- Same week in the same lane: Anthropic is merging Claude Design into Claude Code, shipped Co-work on iOS, and Slack now takes Codex and Claude directly inside a channel.
Why it matters
- The agent surface moved off the terminal and into the messaging app you already live in. Approval-gated write access to a personal channel is a real primitive.
Sources
03
Someone fingerprinted Ox Alpha and it looks like GLM
A stealth model got taken apart by tokenizer forensics, and the evidence points straight at z.ai.
What they showed / shipped
Why it matters
- Temp-0 output matching plus identical error strings is a repeatable method. You can run this on any anonymous endpoint you're evaluating.
Sources
04
a16z's data center charts and the power mandate that landed with them
a16z published county-level data saying data centers make local economies better, in the same week Trump moved to make those data centers generate their own power.
What they showed / shipped
Why it matters
- The constraint on AI capacity is now electricity permitting, not chips. A self-generation mandate changes who can build and how fast.
Sources
05
AI raised homework scores and then exam scores fell
A study found the measurable win and the measurable loss in the same students - grades up where AI helped, down where it couldn't.
What they showed / shipped
Why it matters
- This is the offloading problem with a number on it. If your tool does the reps for the user, measure them on the thing your tool can't do.
Sources
06
Runway shipped SDR to 16-bit HDR conversion
Runway Ruby converts any video, generated or uploaded, into 16-bit HDR ProRes and EXR sequences - which puts AI video into a real post pipeline.
What they showed / shipped
Why it matters
- BT.2020 with PQ or HLG and 12-bit ProRes is the spec sheet of a finishing tool. The bar moved from 'looks good on a phone' to 'survives a grade'.
Sources
07
Thinking Machines is paying for agent data with free access
Inkling is free on OpenRouter for a few weeks, agentic harnesses only, and the price is your usage data.
What they showed / shipped
Why it matters
- Free frontier-ish inference for agent work, right now, with a clearly stated data cost. Read the trade and decide - don't route client work through it blind.
Sources
08
A million people clicked LinkedIn's AI slop button
LinkedIn shipped a way to flag AI-generated posts and over a million users used it, while the anti-AI campaign it feeds is being called well-run by the people it targets.
What they showed / shipped
Why it matters
- 'reads as AI' is now a measured, product-level signal on a major platform. Detectability is a distribution problem, not a vibes problem.
Sources
09
DeepMind is teaching agents to play games they've never seen
After Atari and StarCraft, DeepMind is back on games as the testbed - this time for agents that have to navigate 3D worlds they weren't trained on.
What they showed / shipped
- DeepMind framed 15+ years of games research as the through-line from Atari to StarCraft II Grandmaster to SIMA teaching agents to understand 3D worlds.
- The stated frontier is navigation and generalization in 3D environments, not mastering one title.
- This sits directly alongside the ARC-AGI-3 result - both are 'act in an unfamiliar environment and work out the goal'.
- Robotics kept pace: Galbot showed a new agile humanoid at WRC'26 and humanoids are now playing autonomous tennis.
Why it matters
- Games are cheap infinite environments with clear reward. The pattern transfers to any agent that has to operate a UI it has never seen.
Sources
10
Anna's Archive says publishers are destroying the books after scanning
The pitch is a race: rare physical books are being destructively scanned for AI training, and the archive wants them captured before the originals are gone.
What they showed / shipped
- Anna's Archive argues AI companies are destroying physical books in the scanning process and is calling for rare books to be scanned before it's too late.
- Destructive scanning - cutting the spine to sheet-feed pages - is faster and cheaper than non-destructive imaging, which is why it wins at volume.
- The post hit HN twice from different mirrors, drawing 833 and 2 comments respectively - a signal of how hard this one landed.
- It follows the Amazon rare-books story from earlier this week, but with a specific ask rather than a report.
Why it matters
- Training data provenance is becoming a physical-world question, not just a licensing one. The artifact gets consumed to make the dataset.
Sources