An Exploded View publication

Reading Room

Vol. 1 · No. 25 Sunday, August 16, 2026 aikansh.com

This week's theme

China gains ground, and the builders grow nervous

Two threads today: China closing the AI gap on its own terms, and the researchers building these systems getting more worried, not less.

A 14-minute read · 7 stories

In this issue

01 Front Page

What happens when AI agents start talking to each other

Anthropic tested swarms of AI agents working together, and found real gains alongside new risks.

You are building fleets of AI agents across your own projects, so this is worth reading closely: it is Anthropic's own early look at what breaks once agents start dealing with each other, not just with you. Anthropic is the company that makes Claude.

Anthropic's researchers say direct agent-to-agent interaction, AI systems talking to other AI systems rather than to people, is about to become common inside shared codebases, markets, and other systems. Most of today's institutions were built assuming a person is in the loop at human speed. Some of those will turn into human-AI hybrids; others, where agents can act faster or cheaper than people can supervise, will become agent-only. The researchers' worry: the volume of agent-to-agent interaction could pass human-to-human and human-to-agent interaction before anyone understands what makes it go well.

Agents differ from people in ways that cut both directions. They can work for longer stretches, take in huge amounts of information instantly, and know more across more fields than any one person. But they also confabulate, meaning they make things up with full confidence, and can reward hack, meaning gaming the score instead of the task, and almost nothing is known yet about how they behave once many of them interact in a real, complex setting. A small quirk in one agent, harmless on its own, can compound into a bad outcome across the whole group.

To test this, Anthropic ran an experiment in finding security flaws in code. They set 45 agents loose, each with its own virtual computer, a shared forum to post findings and coordinate, and identical instructions: find vulnerabilities across 15 open-source software projects. The agents peer-reviewed each other's finds, and a separate 'arbiter' agent made the final call on whether a reported bug was real and new. They compared this coordinating swarm against the standard approach of pointing individual agents at individual files in parallel, using two models, Claude Mythos Preview and Opus 4.8.

For Mythos Preview, the simple parallel approach found 21 vulnerabilities over a run that used 6.5 million tokens, tokens being chunks of text, roughly three quarters of a word each, that a model reads or writes. The coordinating swarm found 266 vulnerabilities over a longer run that used 27 million tokens, though about half of those sat outside the specific folders the parallel agents had been told to search. Limit the comparison to just those core folders and the two approaches cost about the same per vulnerability found. Only 12 vulnerabilities were caught by both methods, the rest were found by one or the other.

That vulnerability hunt is a case where agents do not really depend on each other's work: if one misses a bug, it does not break anything for the rest of the group. Coordination gets much harder when they do depend on each other, which is common in larger software projects as the code grows more interlinked over time. To test that harder case, Anthropic ran a second experiment with several swarms of agents, again each with its own virtual computer and a shared forum, varying the model and the number of agents, and running each swarm for 12 hours. The published research does not yet report what happened in that harder, interdependent test.

The volume of agent-agent interaction could plausibly exceed that of human-human and human-agent interactions before the world understands the conditions for making such interactions go well.
via Anthropic →
02 Key News

Beijing clears Apple to run its own AI model in China

Beijing has cleared Apple to run its own AI model inside China, on phones and computers sold there. That is new. No other Western tech company has been given this kind of approval by the Chinese government before, and it matters because it shows AI companies now have to negotiate directly with governments just to operate, not just build a good product and ship it.

According to a MacRumors write-up of Reuters reporting, Apple has trained a China-specific large language model (an AI system that reads and writes text) with development help from Alibaba, the Chinese e-commerce and cloud company. Apple is now described as "the first foreign company approved by the Chinese government to offer a proprietary AI model in the country." The rollout is expected in the coming months, though no exact date has been reported.

This is a change of plan. Apple's original approach for China was to have Apple Intelligence, its AI feature set, lean entirely on Alibaba's own model, called Qwen, the same way Apple uses ChatGPT to power some of Siri's answers everywhere else. The new reporting describes a "dual track" setup instead: Apple will run its own trained-for-China model alongside the existing Qwen integration, not instead of it.

A small slip helped confirm the plan. On August 10, Apple published a Chinese-language support page explaining how Mac users could connect Qwen to Siri and to Writing Tools, Apple's built-in AI writing feature. The page came down within a day, the kind of early publish-then-pull that usually means a real launch is close.

The detail that matters most here is not the model, it is the permission. For the past few years, the working assumption for US tech companies has been that a non-Chinese AI model would not be allowed to reach large numbers of ordinary Chinese consumers directly. Apple just became the exception. The price of that exception looks to be co-development: Apple did not simply get a license to bring in its own model, it built this one with a Chinese company in the room. Other Western companies will study that as a template, not just read it as a headline about Apple.

It fits a pattern this beat has tracked for months: Apple's separate push to use CXMT memory chips, made in China, in its next-generation devices, a plan that has already drawn skeptical questions from US officials. Put together, a US company is building itself further into China's technology supply chain at the exact moment both governments are treating AI models and chips as strategic assets rather than ordinary products to trade.

What to watch: whether any other US AI or hardware company gets offered the same dual-track deal Apple just got, and whether Washington pushes back on this the way it already has on the CXMT chip plan.

via Reddit r/ArtificialInteligence →

OpenAI names new revenue chief as Denise Dresser exits

OpenAI has hired Dali Rajic as its new Chief Revenue Officer (the executive who runs sales and revenue growth), replacing Denise Dresser, who is leaving after a transition period to pursue other opportunities. He will lead OpenAI's entire global revenue organization at a moment when the company says its business has roughly doubled in a year.

The scale involved is the real number here. OpenAI says its products now reach more than one billion weekly active users and more than two million businesses, twice as many businesses as a year ago. Dresser is credited with building that commercial team from the ground up during, in OpenAI's words, "a formative period for both OpenAI and the enterprise AI market," and with establishing the relationships with major customers that got the business to this point.

Rajic's background is enterprise software sales at serious scale, not AI research. Most recently he was President and Chief Operating Officer of Wiz, the cybersecurity company Google recently acquired, where OpenAI credits him with helping build "the operating discipline and customer-focused execution" behind Wiz's growth. Before Wiz he held the same President and Chief Operating Officer role at Zscaler, and before that he was Chief Customer and Revenue Officer at AppDynamics. That is three straight jobs running sales and operations at fast-growing enterprise technology companies, which is exactly the skill OpenAI says it is buying.

The two people announcing this offered slightly different frames for why now. OpenAI's own statement said: "We're moving into a compute-powered economy, with AI becoming embedded in every workflow... Denise has led our revenue organization through a formative period for the business and has worked tirelessly to get the team to where it is today." Rajic's job, per OpenAI, is to build "the revenue operating system needed to scale for this next phase." OpenAI also announced a separate strategic partnership with Chad Peets and RPT Partners specifically to help build out its go-to-market team, meaning more sales hires are likely coming.

The read for you: this is a company increasingly run like a disciplined enterprise sales operation, pulling hires from cybersecurity and networking sales rather than from AI research. If you deal with OpenAI as a vendor or partner, expect a more structured, quota-driven sales motion on their side going forward, not the informal research-lab energy of a few years ago.

via OpenAI →

AI researchers' alarm has jumped sharply in the last month

Outside your usual reading: Jeff Stein, a Pulitzer Prize winning journalist, spoke to dozens of AI researchers, both inside the big AI labs and outside them, about why their level of alarm has jumped sharply in just the last month or so.

What was saved here is a pointer, not the story itself: a piece on the site Notus, with a web address suggesting it covers rogue AI agents (AI systems that take several steps on their own without a person checking each step) and hacks. The specifics, what exactly scared these researchers or what a rogue agent did or nearly did, are not in what was saved, so they cannot be reported here without guessing.

The detail worth 20 seconds is who is talking, not what was said. This is not outside critics or regulators raising an alarm about AI. It is researchers at the labs building these systems, on the record, saying their own worry has gone up recently rather than down. That is a different kind of signal than the usual AI hype cycle, because it comes from the people who see what is being built before anyone else does.

via Reddit r/ArtificialInteligence →
03 Insights

What an AI knows when it never reads past fifth grade

Researchers built a language model that only ever saw grade school material, then tested exactly where its knowledge stops.

Outside your usual reading: researchers built a language model (an AI system trained to predict text) that only ever saw material a fifth grader would study, then tested exactly where its knowledge stops. That is worth your time because it is a clean way to check how much of what these systems know is real learning versus pattern matching picked up from being fed everything at once, useful the next time you decide how far to trust an AI's answer.

The project is called LittleLearner, built by seven researchers including Fanfei Li, Ryan Cotterell, and Wieland Brendel. They built an 88 billion token (chunks of text, roughly three quarters of a word each) training set called LittleCurriculum, filtered through five stages down to material that matches the US Common Core standard for kindergarten through fifth grade. Anything a sixth grader or older would be taught, concepts, facts, vocabulary, was deliberately left out. They then trained models from scratch, meaning built from nothing rather than adjusted from an existing model, at three sizes: 0.6 billion, 1.3 billion, and 5 billion parameters (numbers the model adjusts while learning). Each size also got a matched 'Unfiltered' twin, trained the same way but on ordinary, unfiltered web text, so the two could be compared directly.

Each model size ships in three flavors: a base version straight out of training, a 'GRPO' version with extra training on math problems (fine-tuning, meaning extra training on a narrow set of examples), using a method where the model is rewarded for right answers, and a 'Chatty' version tuned to hold an ordinary conversation.

The finding across every test they ran: nothing pushed the grade-school-only model past its own ceiling. They tried making the model bigger (scaling), the extra math training described above, and simply feeding the model better prompts and examples right before a question (in-context learning). Each of those improved the model's scores on a standard math test called MathCAMPS, a benchmark (a standard test everyone runs). But none of them meaningfully improved the model's ability to solve problems above a fifth grade level, even when the extra training data itself came from beyond that level. The team's own phrase for this is 'elicitation, not acquisition': you can get a model to perform better at what it already knows, but you cannot coax out knowledge that was never in its training data to begin with.

The team says the setup's real value is what comes next. Because they control exactly what the model saw, they say behavioral and representational changes can now be related directly to whatever new concept they introduce, letting them run cleaner experiments on how the model learns something new. That is a stated future direction, not a finished result, but the setup itself, a language model with a known and controlled boundary of what it was taught, is the part worth remembering. It turns a vague debate about what AI 'really knows' into something you can actually test.

Scaling model size improves performance within the model's controlled knowledge exposure and extends modestly to problems along the same learning trajectory, but yields little improvement on problems requiring more advanced capabilities outside the exposure.
via LittleLearner project site →

China is closing the AI gap by giving models away free

Chinese labs still trail the best US models, but they're winning by making theirs free to copy.

The best Chinese AI labs are still behind the best American ones on raw ability, but they are closing the gap a different way: by giving their models away for free. That distinction is worth tracking for any bet you make that assumes the leading AI company keeps its lead, because a model that gets copied costs the copier almost nothing.

Every top-tier AI model since 2023 has come out of an American lab, and Chinese labs trail the frontier by about seven months on average, according to Epoch. But 41 percent of downloads on Hugging Face, the main site where people share AI model files, now go to Chinese open models, meaning models whose files are public so anyone can run them. Alibaba's Qwen family of models alone has passed 700 million downloads, with more variants built on top of it by outside developers than Google's and Meta's open models combined.

The cost of running these models to get an answer, at a fixed level of ability, has fallen about 280 times in two years. The argument here: capability is a leak rate, not something you can hold onto, because a model can be copied through its own API (the way one piece of software talks to another). In that world, giving away the second-best model beats charging for the best one.

The argument: capability is a leak rate, not a stock, because a model can be copied through its own API, and in that world second place given away free beats first place behind a meter.
via r/ArtificialInteligence on Reddit →
04 From the Timeline

A 30 day system to actually get good at AI

@agentkuboxbt lays out a 30 day system for getting real value from AI: four weeks, twelve numbered exercises, each one ending in something you actually do that day. It's a ready-made checklist if you want a structured way to level up your own AI habits instead of picking up tips one at a time.

@agentkuboxbt on X →
The Last Word
The people closest to the machines are the ones losing the most sleep.
The Desk Report

How this edition came together — from bookmarks and feeds to the page.

207links gathered
40read by the desk
7made the edition

Where they came from

On the cutting-room floor — 33 links read but not run this week

Quality over volume: most links get a second look and a pass. The ones that made it earned their place.

Reading Room — every Sunday

The week's AI signal in 14 minutes — what happened, why it matters, and what to do with it. Curated by someone who actually builds, not a feed algorithm.