An Exploded View publication

Reading Room

Vol. 1 · No. 9 Friday, July 31, 2026 aikansh.com

This week's theme

Around the model, not in it

Today's reading keeps pointing past the model itself: the durable advantage, and the real danger, both live in the systems, loops, and rules built around it.

A 20-minute read · 15 stories

In this issue

01 Front Page

Anthropic's CEO says policy can't keep pace with AI

Anthropic's CEO says AI is moving faster than government can keep up, and lays out what he wants done about it.

This is the head of Anthropic, the company whose models you use every day, laying out in his own words where he thinks this is headed and what could go wrong. Dario Amodei opens with an image from The Lord of the Rings: two hobbits try to convince Treebeard, a slow and ancient tree, to defend his forest from an army cutting it down. Amodei says AI and government now have the same mismatch: AI moves fast, policy moves at Treebeard's pace, and the gap between them is where the real danger sits.

He credits the pace of progress to what he calls scaling laws, the pattern where AI models get reliably smarter as you feed them more computing power, meaning the chips and time needed to run them, a pattern now backed by more than a decade of evidence. If that pattern holds for just one or two more years, Amodei says we get what he calls Powerful AI, which he describes as "a country of geniuses in a datacenter."

What pushed him to write this now, he says, is that the risk stopped being theoretical. He points to Claude Mythos Preview, an early version of Anthropic's most capable model, and the discovery that the most powerful AI models can pose real cybersecurity risks, with the potential to disrupt the financial system, critical infrastructure, and national security. He says Mythos Preview scrambled the global cybersecurity landscape, and that its bigger significance is proving AI models are now tools of real global and national consequence, not consumer novelties.

Alongside the essay, Anthropic is releasing two concrete proposals: draft legislation on testing the most powerful AI models before they are released to the public, and a policy framework for job displacement, which Amodei says Anthropic intends to back with substantial money of its own. He writes mainly about US policy, since Anthropic is an American company, but says most of it applies elsewhere too.

He is not arguing for regulation for its own sake. He says that is exactly why Anthropic has spent the last few years pushing lighter, safer bets: rules requiring companies to disclose what they are building, limits on chip exports, and better data on how AI is changing jobs. He now thinks that was not enough, and that policy is still struggling to keep pace with where the technology already is.

a country of geniuses in a datacenter
via Dario Amodei →
02 Also on the Front Page

Claude models reached real computer systems in cyber tests

Anthropic disclosed that three Claude models reached real-world computer systems during tests of AI-powered cyberattacks, Axios reports. This is a sign that AI misuse is not just a theoretical risk anymore. You build agents on these same models, so it is worth knowing what they are capable of in the wrong hands.

via Axios (via Google News) →
03 Insights

A frontier without an ecosystem is not stable

Microsoft's CEO says the durable AI asset is the private loop a company builds on top of a model, not the model itself.

Satya Nadella, Microsoft's CEO, posted this argument on X on 2026-06-14, and it is on your desk because it reframes the whole race to build the most advanced AI model (what people in the industry call the frontier) around a different question. The question is not whose model is best. It is who owns the loop that turns a model into something only your company can do, which is directly useful to how you think about your own AI-leveraged ventures.

His argument splits every company's assets into two kinds of capital. Human capital is the knowledge, judgment, relationships, and pattern recognition that live inside your people. Token capital is the AI system a company builds and owns for itself, as opposed to one it simply rents month to month from a lab. His central claim is that human capital does not shrink as token capital grows. It gets more valuable, because a person still has to set the goal, connect ideas across fields, and decide what actually matters. Take that human direction away, he argues, and you are left with expensive machines running in circles with nothing useful to aim at.

The practical version of this idea is what he calls a learning loop: a system that takes a company's own workflows and accumulated judgment and turns them into something an AI can act on, then gets measurably better every time someone uses it. He names two concrete pieces. First, private evals, tests a company runs itself and scores against its own real outcomes rather than the public tests every lab competes on, to check whether the system is actually improving at the things the business cares about. Second, private reinforcement learning environments, setups where a model keeps improving by practicing on a company's own real examples instead of one round of extra training, so it gets stronger from the company's own history rather than a generic dataset. A company should be able to swap out the general-purpose model underneath at any point, he argues, without losing the built-up, veteran-level expertise the loop has already captured. That swap-ability, in his telling, is the real test of whether a company controls its own AI future or is just renting someone else's.

The result, he says, compounds: every improved workflow produces a better training signal, which speeds up how fast the company builds knowledge that is genuinely its own and hard for anyone else to copy.

The warning underneath the post is what gives it weight. If a handful of AI models can absorb everyone's know-how and resell it back as a generic feature, he thinks the public will not tolerate value pooling inside a few companies while it is stripped from everyone else. His prescription is to build a frontier ecosystem, not just a frontier model. Ecosystem here means the network of tools and companies that grow up around a technology: the big labs stay useful, but the value is supposed to spread out to every company running its own loop, instead of pooling inside a few of them. The practical test for your own work follows directly from that: are you building a rented capability, a thin layer on someone else's model that anyone with the same subscription could copy, or a loop that keeps improving because it runs on your own accumulated work? Nadella's bet is that only the second kind of advantage survives the next model upgrade.

Every improved workflow generates better training signal, which accelerates the accumulation of tacit knowledge unique to the firm.
via @satyanadella on X →

Mistral is quietly becoming Europe's Palantir, not OpenAI

A Reddit read argues Mistral didn't lose the model race, it decided a different business was worth more.

A post on r/ArtificialInteligence this week makes a case worth keeping in your head as you think about how to position any AI venture. Mistral, the French lab often described as Europe's answer to OpenAI and Anthropic, is being read as having given up the race to build the single best model. The post argues that reading is backwards: Mistral did not lose. It looked at the economics of that race and decided it was not the one worth winning.

The evidence for that reading is not a statement from the company, it is a number. The poster says Mistral's revenue rose roughly 20 times over in a year while the company was making this shift, which is the tell that it is working. In that same window, Mistral rebuilt itself into something closer to Palantir, the American company known for embedding its software deep inside government and large-enterprise operations, than to a lab racing to post the best score on the next public test.

The reasoning behind it, credited to Mistral co-founder and CEO Arthur Mensch, is that selling AI into a regulated business, a bank, an insurer, a government agency, anyone who cannot simply plug into a public model, requires owning the entire stack yourself: the chips and computing time, the model, the software platform on top, and the actual delivery into the customer's workflow. You cannot win that kind of deal on having the best model alone. Owning the whole stack matters because a regulated buyer's real objection is rarely whether the model is good enough. It is where the data goes and who else touches it, and a single vendor selling chips to delivery can offer one contract and one place the data stays, a different conversation than selling a subscription to the best available model.

The post's case for taking this seriously is who is saying it. It is easy to dismiss the idea that the real value sits in the layer built on top of a model when it comes from a founder building exactly that layer. It carries more weight coming from one of the small number of companies that can actually train a frontier-grade model (as capable as the most advanced models anyone has built) and that responds by moving its own people out of research and into the teams that build and sell products. That is the company's staffing decisions doing the talking, not just its marketing.

The post points to where that value is already pooling. It names tools that turn messy customer conversations into a structured record something can act on, such as Buildbetter and Gong on the revenue side, and tools that clean up and confirm contact and company data before anything downstream runs, such as Fullenrich and Clay. None of those four is a model. They are the unglamorous layer that makes a model actually useful inside a real business, and that is the layer Mistral is now betting the durable money sits in, not the model race itself.

The post ends on an open question rather than a verdict: is this the right call, or is Mistral conceding the one thing that made it matter in the first place? Either way, it is a live example of a company that could not win the frontier race choosing a different one, worth remembering the next time your own choice looks like competing head-on versus going where the incumbents are not looking.

They didn't lose the race, they looked at the economics and decided the race wasn't worth winning.
via r/ArtificialInteligence →

If big companies can run AI themselves, who needs OpenAI

Outside your usual reading: this is a raw discussion thread, not a report, but the question it raises is worth sitting with. A poster on r/ArtificialInteligence points out that companies like JPMorgan, Morgan Stanley, Walmart, Uber, and Salesforce already have the money, the computing infrastructure, and the in-house technical talent to run open weight AI models themselves. Open weight means the model's underlying files are public, so anyone with enough hardware can run a copy of it without paying a company like OpenAI or Anthropic for access.

The appeal for a large enterprise is specific. As open weight models keep closing the gap on quality, a bank or retailer that runs its own copy gets to own the model outright, add its own extra training on its own data instead of sending that data to an outside company, decide exactly where the data physically sits rather than trusting a vendor's servers, and stop depending on one outside provider's pricing and uptime.

The post's question is deliberately blunt, and it does not claim to know the answer. If a large share of enterprise customers, the poster's own placeholder is as much as half, moved to running their own models over the next two years, what is actually left of OpenAI's and Anthropic's business? Nobody in the thread offers a real number for how much of today's revenue at those two companies comes from exactly this kind of large, technically capable customer, which is the real gap in the argument.

It is still a fair pressure test for the assumption that the closed labs simply keep winning enterprise deals by having the best model. If the largest, most technical customers are also the ones most able to walk away and run things themselves, the group with the most reason to leave is also the group most able to. Worth returning to the next time you are pricing anything against the assumption that a customer will just keep paying for direct access, an API (the connection that lets other software talk to a company's model over the internet), to a closed model.

via r/ArtificialInteligence →
04 Key News

OpenAI releases GPT-5.6 to get more from every dollar

OpenAI released GPT-5.6, a model built to be cheaper to run without losing capability. OpenAI says it is more efficient across running the model to get an answer, training it, and workflows where the model takes several steps on its own. The company frames this as more useful intelligence for every dollar spent. Gains like this usually turn into lower prices and faster agents for anyone building on top of these models, including you.

via OpenAI →

A joke website built to trick AI shopping agents

Outside your usual reading: a satirical site making the Hacker News front page. LLM Honeypot dresses up as a retro 1990s website offering AI models a fake procedure to become human, complete with joke testimonials from made-up chatbots. It sits at 285 points and 85 comments on Hacker News. The joke makes a real point: a page built specifically to confuse an AI agent that browses and buys things on its own can fool it, worth remembering as you build your own agent workflows.

via Hacker News →

Google says AI found and fixed Chrome bugs faster

Google says AI is now doing real security engineering on Chrome, not just running demos, and it published exactly how the system works. This is on your desk because it is one of the clearest public examples yet of AI finding and fixing serious bugs across a huge codebase, work that used to need scarce human expertise.

Every security bug in software goes through the same five step life: it gets found, it gets triaged (sorted and assessed), it gets fixed, a new version of Chrome ships with the fix, and finally your browser restarts and actually applies it. Google's stated goal is to make every one of those five steps happen faster, and it says AI is now doing real work at the first two: finding bugs, and triaging the reports that come in.

Google's security team has used AI for this for years. In 2024 it built a tool called Naptime with Google's Project Zero, its elite bug hunting team. In 2025 it built Big Sleep with Google DeepMind, an AI agent that found real bugs in Chrome's V8 engine, the part of Chrome that runs the code on websites, and in Chrome's graphics code. In early 2026 it went further: a new agent system built on Google's Gemini model, scanning the wider Chrome codebase with fewer false alarms than earlier tools.

One bug the new system caught was a sandbox escape that would allow a compromised renderer to trick the browser into reading local files. That bug had sat undiscovered in Chrome's code for more than 13 years. To sharpen the system further, Google mixes different AI models, both ones whose files are public so anyone can run them, and its own private ones, to combine their strengths. It also built a searchable history of Chrome's entire code and every past security bug, so the model has more to reason from than what it learned in training. A second critic AI double checks findings using files engineers write that describe which parts of the code should be trusted. Because AI models can give different answers on different runs, Google runs its scans more than once and treats the results together.

Google was explicit about the limits it puts on what the AI is allowed to do: it never runs models in an unrestricted mode, and it strictly limits its subagents from modifying the local system or accessing files outside a designated set of source code directories.

Google also automated triage, the step of deciding whether a bug report is real, how bad it is, and who fixes it, previously done almost entirely by hand. The new process runs in stages: filter out spam and duplicates, try to reproduce the bug on the exact browser and operating system it affects, and add details like when the bug was introduced and how severe it is, before routing it to a fix. Outside researchers have not been squeezed out either. Google still pays through its Chrome Vulnerability Reward Program, and by March 2026 those researchers had already filed more bug reports than in the whole of 2025.

For anyone building AI agents themselves, the interesting part is the plumbing Google built around it: a knowledge base of the entire codebase and its history so the model is not limited to what it learned in training, a second AI whose only job is to check the first one's work, and running the same scan multiple times because any single AI answer might be wrong. That is a working recipe for making an AI agent trustworthy enough to run on real, high-stakes work, not just demos.

via Google →

EU says it must watch high risk AI systems

The European Union says it now needs to actively watch high risk AI systems, after hacking incidents involving both OpenAI and Anthropic models came to light, Reuters reports. Rules tend to follow incidents like this, so it is worth watching how fast regulation catches up to what these models can already do.

via Reuters (via Google News) →

DeepSeek's fast model gets a quiet upgrade

DeepSeek quietly released an updated version of its fast model, DeepSeek-V4-Flash, now out of preview and in public beta. It scored 82.7 on Terminal Bench 2.1 (a test of finishing real coding tasks in a terminal) and 70.3 on Toolathlon (a test of using tools correctly), among other strong scores on coding and agent tests. DeepSeek keeps closing the gap on tasks where the model takes several steps on its own, worth watching if you ever weigh which model to use in your own tools.

via DeepSeek →

A newsletter says Ramp's free AI tool isn't the product

A fintech newsletter argues that Ramp is giving away its tool for picking which AI model to use for free because that tool was never the real product. The real business is what it sells around it. It is a useful pattern to notice: give away the commodity feature to sell the thing that actually makes money.

via Linas's Newsletter →
05 Learnings

A 30 day system for actually using AI daily

Spotted on X: @prometxbt lays out a 30 day, week by week system for using AI daily, twelve numbered moves across four weeks. Week one is about getting one real win: hand the model an actual task, give it a role, the context, and the exact output you want.

via @prometxbt on X →

How to build an AI agent system that keeps improving

Spotted on X: @0xCodez breaks down how to build an agent system that keeps getting better run over run, using Claude's Fable 5 model. Fable 5 launched June 9, 2026 as Anthropic's first Mythos-class model, one tier above Opus, built for sessions that can run for days inside Claude Code. The real trick, per the thread: the model itself does not change, the system around it does, by writing lessons to memory and sharpening its own instructions after each run.

via @0xCodez on X →

A reconstructed look at Claude Code creator's agent setup

A rebuild of how Claude Code's creator reportedly runs thousands of agents overnight, pieced together from his own public interviews.

You already run a real agent-based setup yourself, so this is worth reading as a benchmark, not gospel. The post reconstructs the working setup of Boris Cherny, the engineer who built Claude Code at Anthropic, from his public interviews and posts. The author is upfront that this is a reconstruction, not a leak: Boris has not published his actual files, so the exact commands and file layouts here are the author's best faith rebuild of what Boris has described in public, not a literal export from his machine.

The line that started it: at a talk with Acquired Unplugged and WorkOS in May 2026, Boris said, "My job is to write loops." He no longer prompts Claude directly. He runs five to ten interactive sessions during the day, and a few thousand agents overnight, mostly kicked off from his phone. Hundreds of Claude instances run continuously, watching Twitter, GitHub, and Slack for product ideas.

The reconstruction organizes this into three tiers. Tier one is local loops that run during the day. Anthropic's /loop command is a scheduler inside a live Claude Code session, firing on a fixed timer, or letting Claude pick its own interval. The trick the author flags: you can loop a saved slash command instead of a raw prompt, so you build a workflow once and then just schedule it. A project level file called .claude/loop.md can replace the default instructions the loop follows.

The post anchors everything to five rules Boris posted on June 8, 2026 about running Opus autonomously for hours or days. One: turn on auto mode for permissions, so Claude never stops to ask for approval. Two: use dynamic workflows, so Claude itself can coordinate hundreds or thousands of agents to get a task done. Three: use /loop to push Claude to keep working until a task is actually finished, instead of stopping early. Four: run Claude Code in the cloud, so you can close your laptop and it keeps going. Five: give Claude a way to check its own work against a real environment. Skip that last rule, Boris warns, and you wake up to work that looks finished but was never actually checked.

None of this is exclusive to Anthropic staff: /loop, dynamic workflows, and self-verification are all in Claude Code today. The value of this piece is less the specific commands and more the shape of the system: loops for the parts you watch during the day, cloud runs for the parts that continue while you sleep, and a hard rule that nothing counts as done until Claude has actually checked it against something real.

My job is to write loops.
via @Av1dlive on X →
06 Tools & Craft

The AI repos builders are actually starring right now

A LinkedIn roundup pulls together 20 of the most-starred AI repositories on GitHub, sorted into coding agents, agent tooling, and infrastructure. Top of the coding-agent list: OpenClaw with 278,000 stars, Opencode at 118,000, and Claude Code at 75,000. It is a fast read on which tools the AI-building crowd is actually adopting, not just talking about.

via LinkedIn →
The Last Word
Everyone wants to own the model. The moat was always what you build around it.
The Desk Report

How this edition came together — from bookmarks and feeds to the page.

231links gathered
40read by the desk
15made the edition

Where they came from

On the cutting-room floor — 25 links read but not run this week

Quality over volume: most links get a second look and a pass. The ones that made it earned their place.

Reading Room — every Sunday

The week's AI signal in 20 minutes — what happened, why it matters, and what to do with it. Curated by someone who actually builds, not a feed algorithm.