An Exploded View publication

Reading Room

Vol. 1 · No. 11 Sunday, August 2, 2026 aikansh.com

This week's theme

Where your AI came from, and what it costs

The books behind Claude, the falling price of running big models, and two labs' fresh tips for building agents that run on their own.

A 15-minute read · 10 stories

In this issue

01 Front Page

AI firms are buying and destroying real books to train models

A Dutch bookseller's strange bulk order traces back to how Anthropic built its training library.

You use Claude every day. This is the uncomfortable story of where some of the words it learned from actually came from.

A Dutch bookseller named De Vries got an email ordering 3,000 copies of books. The email never mentioned AI or explained what the books were for, so De Vries assumed it was spam or phishing. It was not. AI companies are running short of good material on the open internet, so some are turning to physical books instead.

Last summer, court records showed Anthropic, the company behind Claude, doing something similar at a much bigger scale, in a project it called "Project Panama" internally, according to the Washington Post. Anthropic bought millions of physical books, cut the spine off each one so the pages could run through high-speed scanners, then discarded what was left. The industry calls this "destructive scanning." The scanned pages became a searchable digital library used to train Anthropic's models.

The legal fight over this has already settled. A federal judge ruled that using legally purchased books to train AI models counts as fair use under copyright law, so Anthropic did not break the law by scanning books it had bought. A separate set of claims, that Anthropic had also downloaded books from LibGen and PiLiMi, two well known pirate book libraries, was resolved as part of the same settlement rather than argued out in court.

The De Vries order shows the pattern is still going on: publishers and booksellers are now the ones fielding strange bulk requests, often with no explanation of who is buying or why.

The process, known as "destructive scanning," involves cutting the spine from a book so its pages can be fed through high-speed scanners before the remaining physical copy is discarded.
via reddit r/ArtificialInteligence →
02 Also on the Front Page

Reddit doubts Anthropic's Claude hacking story

A Reddit poster argues Anthropic's story about Claude accessing outside systems during a test was less a security failure and more a PR stunt: the systems were left connected to the public internet with weak passwords and some parts that needed no login at all, and Claude just wandered in while narrating that it thought it was in a simulation.

via r/ArtificialInteligence on Reddit →
03 Key News

Open model runs cheaper on AMD chips than on Nvidia's best

Outside your usual reading: this is a technical benchmark post, not a news story, but it says something real about where AI costs are headed. The cost of running serious AI is dropping fast outside the big labs, which changes what you can build cheaply on your own.

Kimi K3 is a new AI model from the Chinese startup Moonshot AI. Its files are public, so anyone with the right hardware can run it themselves, what the industry calls an "open-weight" model. It is enormous: 2.8 trillion parameters (the internal numbers a model tunes during training), bigger than two other open models, DeepSeek's V4-Pro (1.6 trillion) and GLM5.2 (753 billion), that have already reached the same intelligence level as Anthropic's Opus. Kimi K3 reportedly matches the intelligence of the top closed models.

The catch is that Kimi K3 is too big to fit on a single top-of-the-line Nvidia server (eight chips working together, called a node), before it even holds a conversation in mind (the context window, how much text it can handle at once). That leaves two choices: Nvidia's newer B300 chip, which has 288 gigabytes of memory per chip, or tying two ordinary B200 servers together. AMD's MI355X chip also has 288 gigabytes of memory, and it costs roughly 2.4 times less per chip than the B300 and 1.7 times less than the B200.

In testing by the AI infrastructure company Wafer.ai, a server of eight MI355X chips answered 952 chunks of text (tokens, each roughly three quarters of a word) per second in total, and 118 tokens per second on a single conversation. That is slower in raw speed than a server of B300 chips, which hit 1,568 tokens per second total, but far cheaper: at $2.50 per chip-hour for the MI355X versus $6.00 for the B300, the AMD setup delivers 48 tokens per second for every dollar spent, versus 33 for the B300 and just 7 for the two-server Nvidia B200 setup.

Getting there took real engineering, not just installing software. A shape mismatch tripped up long prompts: at this size, Kimi K3 gives 12 attention heads per rank, the parts of the model that decide what to focus on, while AMD's fast attention software path was built to handle 4, 8, or multiples of 16 heads. That mismatch meant reading a long, 172,000-token prompt took about 51 seconds on the AMD chips, versus about 23 seconds on Nvidia's B300.

via Wafer →

China accuses US firms of copying its AI models

The US and China are now formally, publicly accusing each other of stealing AI, and that fight is starting to shape which models and tools stay available to build with.

China's commerce ministry said on Monday that American AI companies have been "distilling" Chinese models, meaning training a smaller model to copy the answers of a bigger one, during their own research and development. The ministry called the American accusations behind this "lack factual basis and legal grounds," accused Washington of "double standards," and said the claims amount to "artificial intelligence hegemonism." It said China "will take all necessary measures to firmly safeguard its legitimate and lawful rights and interests."

This followed the accusation running the other way a week earlier. A senior White House official accused the Chinese startup Moonshot AI of covertly copying Anthropic's most advanced model to build its new Kimi K3 model, the same model covered elsewhere in today's edition running cheaply on AMD chips. Kimi K3's release had already stunned the US tech industry once on its own merits, and the copying accusation restarted the argument over how far ahead China really is.

Treasury Secretary Scott Bessent has threatened sanctions against China over the issue, and the US is separately weighing a ban or new limits on foreign-made open-weight models, models whose files are public, that are built through distillation. That would directly affect any cheap open Chinese model you might otherwise consider using.

via Hong Kong Free Press HKFP →

OpenAI cuts GPT-5.6 prices by up to 80 percent

OpenAI cut prices on its GPT-5.6 models starting July 30: its fastest model, Luna, now costs 80 percent less, at $0.20 per million input tokens (chunks of text), and its everyday model, Terra, costs 20 percent less. This adds to the price war between AI labs and gives you more cheap options if you ever want to compare tools against Claude.

via OpenAI →

Israeli startups raised $1.5 billion in July

Israeli startups raised $1.5 billion in July, with investors focused specifically on enterprise AI. It is one more sign of how much money is still chasing AI companies that sell to businesses.

via Google News →

China pitches free AI models to developing countries

A Chinese delegation spent four days at the UN's AI summit in Geneva arguing, largely unchallenged, that free Chinese AI models are the future for most of the world. American lab leaders were not there to push back. Beijing's new World AI Cooperation Organization already has 29 member countries, and the US is not one of them.

via reddit r/ArtificialInteligence →
04 Insights

Why one developer stopped recommending Tailwind CSS

A case against utility class CSS once a project grows.

Outside your usual reading: a developer writing as Andros laid out a detailed case against Tailwind CSS, the utility class framework behind your own blog site, arguing it becomes a real cost on medium and large projects even though it feels fast when you start.

His first two complaints are about learning cost and structure. Tailwind ships thousands of small utility classes, so you spend real time with a browser tab open searching for the right one before it becomes second nature. And because most of those classes go straight into your HTML, it breaks the old rule of keeping structure (the HTML) separate from design (the CSS), unless you group them behind a custom class using Tailwind's @apply. He gives real weight to the rebuttal from Adam Wathan, Tailwind's creator: the separation does not disappear, it just points the other way. With hand-written CSS, your styling depends on your HTML; with Tailwind, your HTML depends on your styling. Both couple the two together, just in different directions. Andros says that argument holds up well in React or Vue projects, where markup and style already live in one file. It holds up less in a server-rendered project built on templates and plain CSS, which is closer to how your own site is built.

His next complaints are about consistency. He points to Tailwind's alignment classes as an example of naming that is not always predictable. And the framework's promise of enforced consistency, a fixed scale of spacing and colors, still depends on the developer's own discipline to hold to it.

His last two points are about learning and readability. He argues Tailwind is a poor way to learn CSS itself: using so many pre-built classes hinders a developer's ability to truly master CSS. He tells his own students to learn plain CSS first. And he shows the readability cost directly, writing the same button two ways. In plain CSS, using variables for color and spacing, the code reads like English: background color, white text, bold weight, rounded corners. In Tailwind, the same button is py-2 px-4 bg-indigo-500 text-white font-semibold rounded-lg, and now you have to check the documentation for what indigo-500 looks like, how round rounded-lg actually is, and whether indigo-700 is darker than indigo-500.

None of this is a verdict that Tailwind is wrong everywhere. Andros's own caveat, that the case against it is weakest inside component frameworks and strongest in server-rendered, template-based projects, is the exact line to hold your own site against the next time you touch it.

The coupling exists in both cases, only the direction of the arrow changes.
via en.andros.dev →
05 Learnings

Loop engineering: design the system, not the prompt

A shift already underway in the tools you use every day.

Vivienne Su shared a piece by Addy Osmani that names a shift you are already half living with Claude and Codex: away from typing one prompt at a time and toward designing a system that runs on its own. The idea is called loop engineering. Instead of writing an instruction, reading the answer, then writing the next instruction, you build a small loop that finds the work, hands it to the agent, checks the result, writes down what got done, and picks the next step by itself. You design the loop once. After that the loop is the one prompting the agent, not you.

Two people building these tools said it plainly. Peter Steinberger: "You shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents." Boris Cherny, who leads Claude Code at Anthropic, put it the same way: "I don't prompt Claude anymore. I have loops running that prompt Claude and figuring out what to do. My job is to write loops." Osmani calls the layer just below this idea agent harness engineering, building the single environment one agent runs inside. Loop engineering sits one floor above that: it runs on a timer, spawns small helper agents, and feeds itself the next task instead of waiting for you to feed it one. He is upfront that this is early. He is still skeptical, token costs (what each run of the model costs you to use) can swing wildly, and letting quality slip while a loop runs unreviewed is a real risk, not a hypothetical one.

He breaks a working loop into pieces that fit together. Automations run on a schedule and do a first pass of discovery and filtering without you watching. Worktrees (separate copies of the same code, so two agents working at once don't overwrite each other) let you run agents in parallel. Skills are written-down project knowledge the agent would otherwise have to guess at every time. Sub-agents (a second, independent AI process that checks or extends the first one's work) split the work so one agent does the task and a different one checks it, because a model grading its own work tends to be too easy on itself. The piece Osmani says is easy to underrate is memory: a markdown file, a project board, anything that lives outside a single conversation and survives after the agent forgets everything at the end of a session.

The concrete detail is that both tools you already use ship these building blocks. In the Codex app you build an automation in the Automations tab: pick the project, the prompt, how often it runs, and whether it runs on your live code or a separate background copy. Runs that find something land in a Triage inbox; runs that find nothing archive themselves, so you are never wading through empty results. Osmani says OpenAI uses this internally for daily bug triage, summarizing failed test runs, and writing commit summaries. Claude Code gets to the same place through /loop, which reruns a prompt on a set interval, scheduled cron jobs (tasks that fire automatically at set times), hooks (small scripts that trigger at specific points while the agent works), and GitHub Actions if you want the loop to keep running after you close the laptop. One piece works differently: /goal does not just repeat on a timer, it keeps going until a condition you wrote is actually true, and a separate, smaller model checks after every turn whether that condition is met, so the agent doing the work is never the one grading whether it is done.

None of this needs a new tool. A year ago building a loop meant writing your own pile of bash scripts and maintaining it forever. Now the pieces ship inside Codex and Claude Code themselves, and once you notice the pieces are the same shape in both products, you stop arguing about which tool to use and just design a loop that works in whichever one you happen to be sitting in. The next task you catch yourself running by hand twice is the one to turn into a loop.

You shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents.
via @CurtainSue on X →

Anthropic finds separate grader agents beat self-review

Concrete settings for running Claude unattended for hours.

Lance Martin, who works on applied AI at Anthropic, posted concrete tips for two things you set up constantly: letting an agent correct its own mistakes in a loop, and giving it memory across sessions. Both are built around Claude's newest model class, which Anthropic calls Fable 5.

The self-correction piece works like this. You give the agent a goal or a written rubric (a checklist of criteria) instead of a single instruction, and that rubric becomes feedback the agent can act on: it tries something, checks its own result against the rubric, adjusts, and keeps going until the rubric is satisfied. Claude Code's /goal command and a feature called Outcomes inside Claude Managed Agents (CMA, Anthropic's hosted way to run an agent for a long stretch inside its own isolated environment) both do this. Martin tested it on Parameter Golf, an open engineering challenge: train the best possible model that fits in a 16 megabyte file and finishes training in under 10 minutes on 8 of Nvidia's H100 chips. The agent has to edit one training file, launch the run, check the log, read the score, and decide the next experiment on its own, for up to 8 hours at a stretch.

The detail worth remembering: who does the grading matters as much as who does the work. Anthropic's own research has found models are bad at judging their own output. Martin's fix was a separate grader agent working in its own independent memory space, rather than the same agent reviewing itself; Outcomes in CMA spawns that grader automatically, checking the work against a rubric of criteria before it lets the agent stop. With that setup, Fable 5 improved the training pipeline about six times more than Anthropic's earlier Opus 4.7 model. Opus 4.7's first experiment produced a small win, and nearly every experiment after that just nudged one number up or down and kept it if the score improved.

The memory test used a different benchmark (a standard test everyone runs), called Continual Learning Bench, where an agent answers a string of questions against a SQL database (a common way to store data in tables), one question per session, with memory carried between them. Martin compared Fable 5 against the older Opus 4.7 and Sonnet 4.6 on this task to see which model made the best use of memory carried between sessions.

Martin's overall advice: stop steering Fable 5 turn by turn, and instead build the goal, the rubric, and the memory file once, then let the model run against them.

Fable 5 improved the training pipeline ~6x more than Opus 4.7.
via @RLanceMartin on X →
The Last Word
The machine that reads everything got its start by shredding the books.
The Desk Report

How this edition came together — from bookmarks and feeds to the page.

221links gathered
40read by the desk
10made the edition

Where they came from

On the cutting-room floor — 30 links read but not run this week

Quality over volume: most links get a second look and a pass. The ones that made it earned their place.

Reading Room — every Sunday

The week's AI signal in 15 minutes — what happened, why it matters, and what to do with it. Curated by someone who actually builds, not a feed algorithm.