01
Anthropic's Fable reproduced half of OpenAI's math results in 24 hours
Yesterday OpenAI's unreleased Astra model got the headlines for cracking ten open math problems. Today an Anthropic researcher reproduced five of them with a model you can already buy, and a formal paper says one of the ten proofs is simply wrong.
What they showed / shipped
- Anthropic's Levent Alpöge reproduced 5 of the 10 results using Claude Fable - autonomously, generic prompt, no internet access, with safeguards so OpenAI's published solutions couldn't leak into context.
- Only one of the five used essentially the same argument as OpenAI's, so these are largely independent derivations rather than recall.
- A separate paper argues OpenAI's claimed disproof of Connes' Rigidity Conjecture is invalid - one of the ten headline results does not hold up.
- Gary Marcus published "amazing but vastly oversold", and the same discourse ran on r/singularity and Digg noting the model solves hard math but still ignores simple instructions.
Why it matters
- Builder lens: the gap between "unreleased frontier model" and "the API you already pay for" was about 24 hours here. Don't rearchitect around a model you can't call yet.
- Creator lens: this is the whole news cycle in one beat - day one is the announcement, day two is the reproduction and the retraction. Covering day two is where the trust is.
Sources
02
94.8% of websites are never cited in an AI answer
Almost nobody blocks the AI crawlers, and almost nobody gets cited by them either - a visibility index just put a hard number on the new referral cliff.
What they showed / shipped
- Only 8.9% of sites block AI crawlers, but 94.8% are never cited in AI answers - from the AI Visibility Index.
- That means the overwhelming majority of the web is being read for training and retrieval while receiving no traffic and no attribution back.
- It lands next to an arXiv paper finding generative AI floods and dilutes the market for books - the same oversupply dynamic, one layer up.
- Reddit's own version of this fight is still live, with the CEO publicly questioning Google's AI Overviews.
Why it matters
- Builder lens: if you ship content as a distribution strategy, the citation rate is now the metric that matters, not the crawl rate. Blocking is not the lever - being citable is.
- Creator lens: this is the number behind the vibe every creator already feels. Search sends less, AI answers send almost nothing.
Sources
03
DeepSeek V4-Flash now runs locally with llama.cpp speculative decoding
The cheap frontier-class model everyone benchmarked last week is now actually runnable on your own machine, fast, with the standard local stack.
What they showed / shipped
Why it matters
- Builder lens: this is the moment a model moves from "cheap API" to "runs on my hardware, no vendor." Different risk profile entirely.
- Creator lens: a genuinely runnable frontier-class local model is the most demo-able thing in AI right now - it's tangible on camera.
Sources
04
Seedance 2.5 held product consistency well enough to fake an ad
One photo of a real product plus a prompt produced an influencer video where the product stayed itself the whole way through - the failure mode that kept AI out of paid ads.
What they showed / shipped
Why it matters
- Builder lens: the unlock isn't fidelity, it's consistency across frames. That's the constraint to test any video model against.
- Creator lens: brand deals were the safe income floor because AI couldn't hold a product. That floor just got shorter - and also became a service you could sell.
Sources
05
Three coding-agent things worth opening today
A sub-1MB Codex clone in C++, a live API cost calculator, and a cheap-model trick that stops you burning premium tokens on deploys.
What they showed / shipped
- MicroCodex - OpenAI's codex agent reimplemented in C++ as a binary under 1MB. Useful as a readable reference for what a coding agent actually needs.
- CostPerPrompt - live AI API pricing with real-workload cost calculators, so you can price an agent loop before you run it.
- Matthew Berman's cheap-deploy trick: route deploys to Luna Max in a fresh thread via an agents.md rule, and keep premium tokens for the work that needs them.
- Context for where this is heading: Boris Cherny on getting Claude Code to rewrite the Claude app in Swift.
Why it matters
- Builder lens: the cost calculator is the one to open first. Agent economics are decided by which model handles the boring steps.
- Creator lens: "here's how to cut your AI bill without cutting quality" is an evergreen, highly shareable format.
Sources
06
China's DFSX claims double the memory bandwidth of Nvidia's GB200
Memory bandwidth is the real bottleneck for inference, and a Chinese chip is claiming 2x Nvidia's flagship on exactly that number.
What they showed / shipped
Why it matters
- Builder lens: if bandwidth-per-dollar genuinely doubles outside Nvidia, inference pricing follows within a year or two.
- Creator lens: the compute story is usually told as chips and export bans. Bandwidth is the more honest framing and almost nobody explains it.
Sources
07
The benchmarks people actually trust are now private jokes
The pelican-on-a-bicycle test stopped discriminating between models, so people are quietly inventing weirder personal benchmarks - which says more about evaluation than any leaderboard does.
What they showed / shipped
Why it matters
- Builder lens: if a public benchmark is saturated, it's marketing. Build a small private eval on your actual workload - that's the only number that moves your decisions.
- Creator lens: "here's the silly test I use to judge every new model" is a format people copy and share. Own one.
Sources
08
Mozilla published a State of Open Source AI report
A neutral party finally put numbers on what "open source AI" actually means in practice, at the moment the term is most contested.
What they showed / shipped
Why it matters
- Builder lens: "open weights" and "open source" are not the same licence, and the difference decides what you're allowed to ship commercially.
- Creator lens: a neutral, citable report is rare in this space. It's a reference you can point an audience at without it looking like vendor marketing.
Sources
09
An AI poster won a state fair art contest
The Ohio State Fair poster contest was won by an AI image, three years after the last time this happened made national news - and this time the argument is smaller.
What they showed / shipped
Why it matters
- Builder lens: the interesting question isn't detection, it's disclosure - contests without a stated policy will keep producing this exact story.
- Creator lens: this is the local, human-scale version of the AI art fight. A state fair is far more relatable than a copyright lawsuit.
Sources
10
The decel debate got a name and a venue
After last week's slow-down letter, the argument about whether slowing AI is even legal is now being had out loud by the people who'd be slowed.
What they showed / shipped
Why it matters
- Builder lens: the legal question - can a slowdown even be coordinated without collusion - decides far more than the ethical one.
- Creator lens: r/singularity softening on regulation is a genuine sentiment shift, and sentiment shifts make better content than position papers.
Sources