An Exploded View publication

Reading Room

Vol. 1 · No. 41 Wednesday, September 9, 2026 aikansh.com

This week's theme

A research breakthrough, and a one-in-four fix rate

AI did real open-ended research this week, yet its everyday output still depends on who checks it and how it was set up.

A 7-minute read · 4 stories

In this issue

01 Front Page

AI reportedly solves a $1 million fluid dynamics problem

A long-standing math problem may have just fallen to an AI model, and now two labs are fighting over who gets credit.

This is the first real sign that AI can do open-ended scientific research and not just fast pattern matching, and it changes what you should assume AI still cannot do.

A team appears to have cracked one of the six "million dollar" problems in math. The Clay Mathematics Institute has offered $1 million to anyone who proves whether the Navier-Stokes equations (the equations that describe how fluids like water and air move) always behave, or whether they can produce an impossible result under some conditions. Mathematician Tristan Buckmaster says a proof now exists showing the equations do break down, what mathematicians call "blowing up." He called it "a Deep Blue-Kasparov moment," referring to the 1990s chess match where a computer first beat world champion Garry Kasparov.

The path took close to a year. Two mathematicians, Diego Cordoba and Luis Martinez-Zoroa, found a trick called "forcing" that uses one part of the equations most researchers normally leave out, assuming it does not matter. Buckmaster and a co-author, Levent Alpoge, who works at Anthropic, spent months running that trick through AI models to search the possibilities. On August 15 they proved that a simpler cousin of the equations, the Euler equations, does blow up, a major result on its own.

Buckmaster alleges that rumors of the unpublished work reached OpenAI, and that an OpenAI team then used the same "forcing" idea and one of its own internal models, over a single weekend, to extend the result and blow up the full Navier-Stokes equations, the harder problem with the $1 million prize attached. Buckmaster says a call with an OpenAI researcher, Bubeck, turned into a dispute over credit: he alleges OpenAI offered to name him sole author of a paper crediting its own model, but wanted to leave off Alpoge because Alpoge works for a rival company. He posted his and Alpoge's results, along with this account, just before midnight on Monday.

One question sits underneath the credit fight: whether OpenAI's model had access, directly or through training, to Buckmaster and Alpoge's own private prompts to a rival AI tool. He says he asked and was told no, but got no answer on whether user prompts feed into training more broadly. There is also a technical catch: the Clay Institute's problem, as formally written, may not even include the "forcing" term the whole proof leans on, since most mathematicians write the problem without it. The prize committee still has to decide whether this counts.

The win here was not an AI inventing a new idea from nothing. Two human mathematicians spent close to a year finding the "forcing" trick; the AI's job was to grind through the resulting search space once it had the right lever to pull. That is still a real jump: a foundational problem that had stood unsolved for generations, decided in days. It also shows the same race dynamics already familiar from AI products now playing out in research, two labs chasing the same unpublished idea, and a credit fight over whose model gets the win.

via Scientific American →
02 Also on the Front Page

Study: AI fixes security bugs correctly only 1 in 4 times

You lean on Claude Code for real engineering work, so this tells you exactly where to keep a human checking the AI's patches instead of trusting them.

Researchers at Off-by-1 Labs tested how well today's best AI models fix known security bugs in code, with the paper published through 1Password. The models produced a genuinely correct fix only about one time in four. The rest of the time they did one of three things: left the bug half fixed, made unrelated changes to code nobody asked them to touch, or introduced a brand new security hole while claiming to fix the original one.

The researchers also tested the obvious fallback: have a human review the AI's proposed patch before it ships. Their finding was that this does not really solve the problem either. Properly checking whether an AI-written patch is safe takes so much careful thought that it comes close to being as much work as writing the patch yourself. Their conclusion: for security fixes specifically, write the patch by hand rather than asking an AI to draft one and checking its work afterward.

The full paper is said to include practical advice on what kind of context helps a model do better at this particular task, though those specifics were not part of what got saved here.

via reddit r/ArtificialInteligence →
03 Learnings

A product exec's real playbook for building loops in Claude Code

This maps directly onto the loop patterns you already use and write about, so it's worth sharpening your thinking against someone else's real production experience.

The framing from this interview: 2023 was the year of writing good prompts, 2024 was about giving the AI the right background material, 2025 was about the "harness" (the surrounding software that lets an AI actually take actions), and the claim is that 2026 is becoming the year of "loop engineering," building small automated cycles that run themselves. The guest is Tyler Folkman, Chief AI and Product Officer at JobNimbus, a company that just closed Utah's largest-ever Series B (a later-stage startup funding round).

Folkman's definition of a real loop, as opposed to just telling an agent to run forever: it fetches its own input, does the work, has a gate (a checkpoint that decides if the result is good enough to continue), writes out a named result, then goes back and fetches its next input on its own.

His most useful, and most counterintuitive, advice: write your first loop by hand, not with AI. His reasoning: he compares an AI-written loop to an e-bike, once the motor does the pedaling, you stop pedaling yourself, meaning you stop doing the thinking that used to sharpen your own judgment. His method: open an empty file, write a short plain-language instruction by hand, test it, improve it, and only then hand the AI a one-line command to wrap it into a loop, then close the loop by hand once to see what it's actually doing and ask it how to improve itself for next time.

Folkman recommends product managers build at least five loops. Customer outreach is the one he personally gets the most value from, since the real cost was never writing the message, it was finding time to figure out who to contact. A thinking-partner loop does the opposite of most AI tools: instead of producing work, it pushes back on the person's own thinking rather than doing it for them. He is blunt that skipping this and forwarding AI output straight up the chain "is not doing product work."

A divergent-prototypes loop builds several product variants from one thin prompt, in his example a crew-scheduling tool for a contractor, then having a simulated customer review the options before any real person sees them. Tyler runs several of these loops side by side using a terminal tool called herdr, which launches multiple AI coding assistants at once whenever a prompt includes the word "agent."

via news.aakashg.com →

Not all compressed AI models are equal, format matters

Compressing an AI model down to a smaller size, called quantization, is not one standard thing: three different quantization methods on the same model weights can swing code-writing accuracy by close to 10 points and speed by 20 to 30 percent on identical hardware. The better methods, like AWQ, protect the small set of values that carry the most meaning; plain rounding does not, so it's worth checking the method, not just the file size, before running or evaluating a local model.

via Business Analytics Review →
The Last Word
The ceiling moved. The floor still needs someone reading the diff.
The Desk Report

How this edition came together — from bookmarks and feeds to the page.

223links gathered
40read by the desk
4made the edition

Where they came from

On the cutting-room floor — 36 links read but not run this week

Quality over volume: most links get a second look and a pass. The ones that made it earned their place.

Reading Room — every Sunday

The week's AI signal in 7 minutes — what happened, why it matters, and what to do with it. Curated by someone who actually builds, not a feed algorithm.