01
Nvidia's agent scored 100% on ARC-AGI-3 and the harness got the credit
Nvidia AVO cleared every level of an interactive reasoning benchmark with no instructions, and the takeaway everyone landed on is that the scaffolding around the model did the work.
What they showed / shipped
- Nvidia AVO completed all 183 levels across all 25 public ARC-AGI-3 environments, figuring out goals with no instructions, no explicit rules, and no stated objective (NVIDIA AI).
- ARC-AGI-3 is the interactive version of the benchmark - the agent has to act in an environment and infer what winning means, not answer a static puzzle.
- TechCrunch's read: the harness, not the model, is now the real hero - the loop, tools and memory around the model are where the capability came from.
- Same week, open harnesses kept shipping: Seed, a minimal self-modifying agent harness, and Proliferate, a self-hostable Codex for any coding agent.
Why it matters
- Builder lens: if the harness is the differentiator, your leverage is not picking a smarter model - it's the loop you wrap around it. That's the part you can actually control and iterate on daily.
- Creator lens: this is the cleanest explainer of the year for why two people using the same model get wildly different results. The scaffolding is the skill.
Sources
02
Codex can send your iMessages now
OpenAI wired Codex and ChatGPT into Apple Messages, so an agent can read a thread, draft a reply and send it after you approve.
What they showed / shipped
- Typing
@messages in Codex exposes an Apple Messages workflow - the agent finds the contact, composes the text, and shows an approval card before anything sends (Riley Brown demo).
- Approval is required by default; the agent will not send without the confirm step, and the draft is editable before it goes.
- It also reads back: he pointed it at a group thread and asked it to analyze the conversation and suggest how to communicate better.
- Same week in the same lane: Anthropic is merging Claude Design into Claude Code, shipped Co-work on iOS, and Slack now takes Codex and Claude directly inside a channel.
Why it matters
- Builder lens: the agent surface moved off the terminal and into the messaging app you already live in. Approval-gated write access to a personal channel is a real primitive.
- Creator lens: 'my AI texts for me' is a demo that lands with a non-technical audience instantly - and the approval card is the part that makes it not creepy.
Sources
03
Someone fingerprinted Ox Alpha and it looks like GLM
A stealth model got taken apart by tokenizer forensics, and the evidence points straight at z.ai.
What they showed / shipped
Why it matters
- Builder lens: temp-0 output matching plus identical error strings is a repeatable method. You can run this on any anonymous endpoint you're evaluating.
- Creator lens: 'here's how you unmask a secret AI model' is a genuinely fun teach - it's detective work with three pieces of hard evidence.
Sources
04
a16z's data center charts and the power mandate that landed with them
a16z published county-level data saying data centers make local economies better, in the same week Trump moved to make those data centers generate their own power.
What they showed / shipped
Why it matters
- Builder lens: the constraint on AI capacity is now electricity permitting, not chips. A self-generation mandate changes who can build and how fast.
- Creator lens: these are real charts with real county-level numbers - the exact thing to put on screen when the 'data centers ruin towns' argument comes up.
Sources
05
AI raised homework scores and then exam scores fell
A study found the measurable win and the measurable loss in the same students - grades up where AI helped, down where it couldn't.
What they showed / shipped
Why it matters
- Builder lens: this is the offloading problem with a number on it. If your tool does the reps for the user, measure them on the thing your tool can't do.
- Creator lens: two directional numbers from one study is the strongest possible shape for this argument. It kills both the hype take and the panic take.
Sources
06
Runway shipped SDR to 16-bit HDR conversion
Runway Ruby converts any video, generated or uploaded, into 16-bit HDR ProRes and EXR sequences - which puts AI video into a real post pipeline.
What they showed / shipped
Why it matters
- Builder lens: BT.2020 with PQ or HLG and 12-bit ProRes is the spec sheet of a finishing tool. The bar moved from 'looks good on a phone' to 'survives a grade'.
- Creator lens: your existing footage is in scope, not just AI generations. That makes it a utility for real projects today.
Sources
07
Thinking Machines is paying for agent data with free access
Inkling is free on OpenRouter for a few weeks, agentic harnesses only, and the price is your usage data.
What they showed / shipped
Why it matters
- Builder lens: free frontier-ish inference for agent work, right now, with a clearly stated data cost. Read the trade and decide - don't route client work through it blind.
- Creator lens: this is the clearest recent example of the real business model. The model is free because the traces are the product.
Sources
08
A million people clicked LinkedIn's AI slop button
LinkedIn shipped a way to flag AI-generated posts and over a million users used it, while the anti-AI campaign it feeds is being called well-run by the people it targets.
What they showed / shipped
Why it matters
- Builder lens: 'reads as AI' is now a measured, product-level signal on a major platform. Detectability is a distribution problem, not a vibes problem.
- Creator lens: a million clicks is the number to use when explaining why default model voice is a liability. The audience is actively flagging it.
Sources
09
DeepMind is teaching agents to play games they've never seen
After Atari and StarCraft, DeepMind is back on games as the testbed - this time for agents that have to navigate 3D worlds they weren't trained on.
What they showed / shipped
- DeepMind framed 15+ years of games research as the through-line from Atari to StarCraft II Grandmaster to SIMA teaching agents to understand 3D worlds.
- The stated frontier is navigation and generalization in 3D environments, not mastering one title.
- This sits directly alongside the ARC-AGI-3 result - both are 'act in an unfamiliar environment and work out the goal'.
- Robotics kept pace: Galbot showed a new agile humanoid at WRC'26 and humanoids are now playing autonomous tennis.
Why it matters
- Builder lens: games are cheap infinite environments with clear reward. The pattern transfers to any agent that has to operate a UI it has never seen.
- Creator lens: the Atari-to-StarCraft-to-3D-worlds arc is a ready-made 60 second history that explains where agents came from.
Sources
10
Anna's Archive says publishers are destroying the books after scanning
The pitch is a race: rare physical books are being destructively scanned for AI training, and the archive wants them captured before the originals are gone.
What they showed / shipped
- Anna's Archive argues AI companies are destroying physical books in the scanning process and is calling for rare books to be scanned before it's too late.
- Destructive scanning - cutting the spine to sheet-feed pages - is faster and cheaper than non-destructive imaging, which is why it wins at volume.
- The post hit HN twice from different mirrors, drawing 833 and 2 comments respectively - a signal of how hard this one landed.
- It follows the Amazon rare-books story from earlier this week, but with a specific ask rather than a report.
Why it matters
- Builder lens: training data provenance is becoming a physical-world question, not just a licensing one. The artifact gets consumed to make the dataset.
- Creator lens: 'the book is destroyed to feed the model' is a concrete, visual, non-abstract version of the training data debate.
Sources