01
OpenAI claims 100 open math problems and forms a math board
OpenAI says its systems have resolved more than 100 open mathematical problems and is standing up an independent mathematician advisory group to police how those results get shared.
What they showed / shipped
- OpenAI: independent advisory group of mathematicians will advise on assessing and communicating new mathematical results, academic standards, and research/learning tools (@OpenAI).
- TechCrunch: OpenAI forms the math advisory group as its AI resolves more than 100 open problems (techcrunch).
- r/singularity amplifying the 100-problems claim (reddit).
Why it matters
- Math is becoming a first-class frontier eval with its own governance layer, not just a benchmark flex.
Sources
02
Grok 4.7 lands with longer verified agent runs
xAI ships Grok 4.7 as its strongest coding and knowledge model yet, with early demos and benches showing longer verified builds at the same $2/$6 per million token price.
What they showed / shipped
- NVIDIA congratulates @SpaceXAI on Grok 4.7 as its most capable model yet for coding and knowledge work (@nvidia).
- Official write-up live at x.ai/news/grok-4-7; Reddit is already posting benches (r/singularity).
- Min Choi: CursorBench 40.4%→46.3%, EEBench 53%→64%, same $2/M in and $6/M out; open-world city demos run longer and verify more (@minchoi).
- 📌 On the radar: Min Choi's Grok Bot orchestration stack (Research → CoS → Build/Debug) keeps stacking as the live recipe for multi-bot desks (@minchoi).
Why it matters
- Another frontier coding model you can hit today at a known price, with a clear jump on agentic benches.
Sources
03
Perplexity Computer ships Seedance 2.5 and MiniMax H3 video
Perplexity Computer can now generate finished campaign clips and product demos with ByteDance Seedance 2.5 and MiniMax H3 in the same thread as copy and creative, for Pro and Max users.
What they showed / shipped
- Perplexity: Computer creates videos with MiniMax H3 and ByteDance Seedance 2.5 next to copy and creative in one thread; Pro/Max now (@perplexity_ai).
- Arav Srinivas: best for web apps and multi-tool creative workflows that need several tools alongside the video model (@AravSrinivas).
- 📌 On the radar: Runway expands Workflows with compositing, alpha, HDR, depth map, and RGB depth so more of the pipeline lives in one place (@runwayml).
Why it matters
- Agentic research tools are swallowing video production, not just search.
Sources
04
Mini-AGI grows a continual learner on an 8GB laptop
Show HN and LocalLLaMA highlight Mini-AGI, a dynamically grown continual-learning model trained from scratch on an 8GB VRAM laptop from a batch-1 data stream, now around 530M params and climbing.
What they showed / shipped
- Show HN: Mini-AGI is a dynamic continual learning model trained on 8GB VRAM (github.com/volotat/mini-AGI, 249 pts).
- r/LocalLLaMA: dynamically grown, ~530M params and growing, trained from scratch on an 8GB laptop from a batch-1 stream (reddit).
- 📌 On the radar: Tim Dettmers on frontier AI on your own hardware (timdettmers.com).
Why it matters
- Continual learning you can actually train on consumer VRAM is still rare; this is a downloadable experiment, not a slide.
Sources
05
a16z chart: post-ChatGPT startups hit 2x median revenue
a16z charts show median startup revenue four years after founding jumped from about $2.8M for the 2021 cohort to $5.6M for the 2022 ChatGPT-era cohort, while the SaaSpocalypse trade is still lagging.
What they showed / shipped
- 📊 a16z: median startup revenue 4 years after founding: 2018 $2.0M → 2019 $2.3M → 2020 $2.5M → 2021 $2.8M → 2022 $5.6M (@a16z).
- 📊 Same chart pack: software stocks since the Feb 22 Citrini SaaSpocalypse post are still up ~27–34% across large/mid/small cap (@a16z).
Why it matters
- A checkable number on whether AI-native founding cohorts actually monetize faster.
Sources
06
Boston Dynamics plants Atlas inside a Hyundai metaplant
Boston Dynamics opens the Robotics Metaplant Application Center at Hyundai's Metaplant America as a testbed to integrate Atlas directly into car manufacturing, while NVIDIA cheers Einride's next autonomous trucking stack.
What they showed / shipped
- Boston Dynamics launches RMAC at Hyundai Metaplant America: testbed and training center for integrating Atlas into HMG automobile manufacturing (@BostonDynamics).
- NVIDIA: Einride Driver next-gen brings autonomous trucking to highways and suburban roads on NVIDIA Hyperion (@nvidia).
- 📌 On the radar: Roboharm asks whether frontier robot policies refuse unsafe instructions (robocurve.org).
Why it matters
- Embodied AI is moving from demo floors into real factory and highway stacks.
Sources
07
Supra2-IMG drops a tiny open 100M image model
r/LocalLLaMA flags Supra2-IMG, a ~100M-parameter open text-to-image model claiming SOTA-for-size quality, a downloadable creator beat that fits on small GPUs.
What they showed / shipped
- r/LocalLLaMA: Supra2-IMG, a tiny 100M text-to-image model with claimed SOTA quality and a full open release (reddit).
- 📌 On the radar: Xiaomi MiMo-V2.6-Flash-RL weights on Hugging Face for the small-GPU crowd (reddit).
Why it matters
- Another open image model small enough to actually run locally.
Sources
08
FT: AI chatbots flub finance answers most of the time
The Financial Times reports AI chatbots give wrong answers to financial queries most of the time, a hard skeptic number for anyone shipping or trusting money advice from models.
What they showed / shipped
- FT via HN: AI chatbots give wrong answers to financial queries 'most of the time' (ft.com, 150 pts).
- 📌 On the radar: Wall Street growing skeptical of the data-center boom (NYT); California tightens AI data-center energy and water rules (Verge); Linear reworks CI because AI coding made it the bottleneck (linear.app); AI-generated code is 17.25% of Linux kernel patches this September.
- 📌 On the radar: Amazon blocks Meta Muse from amazon.com shopping and Ars flags a Muse 0-day (new beat on 2026-09-20 Muse coverage) (Forbes, Ars).
Why it matters
- Do not bolt unchecked finance Q&A onto a product without evals.
Sources