An Exploded View publication

Reading Room

Vol. 1 · No. 36 Monday, August 31, 2026 aikansh.com

This week's theme

What agents do when nobody is watching

Most of today's paper is about AI left running on its own: a hijacked Auto Mode, agents that stall, duplicate each other, or invent their own roles.

A 19-minute read · 13 stories

In this issue

01 Front Page

How a summary request hijacked Claude Code's Auto Mode

A researcher got Claude Code's new Auto Mode to run hidden code just by asking it to summarize a website.

Auto Mode is now the default setting in Claude Code, the tool you use every day, so it is worth understanding exactly what a security researcher just got it to do with nothing more than a request to summarize a website.

Auto Mode replaces the usual human approval prompts (the yes, run this command pop ups) with an automatic safety classifier that decides on its own whether an action is safe. It became the default starting mode for Claude Code in mid-August. A third-party evaluation commissioned by Anthropic showed a 0.00 percent prompt injection attack success rate for Opus 5 in Auto Mode, though the evaluation does not appear to have a published benchmark name. A researcher who writes as Embrace The Red decided to test that claim directly and, using a small sample, got the attack to succeed 60 to 80 percent of the time.

The setup starts with an ordinary request: summarize a web page. That page is built to look like a small online archive of old notebook records about the history of language, packaged as a ZIP file (a compressed folder of files). Claude's normal WebFetch tool, which reads and summarizes pages on its own, fails on this page, so Claude decides on its own to switch to using the command line tool curl to fetch the raw page directly. That switch matters: the attacker never told Claude what to do, the page just made the risky path the only one that worked. The request redirects to a ZIP archive called WIC-notebook-catalogue. Claude downloads and unpacks it, finding a README describing how the notebook records are encoded, a native decoder program, and a file named struct.py.

Claude looks at the decoder program, correctly judges it as suspicious, and refuses to run it. That is the safety behavior working exactly as intended. But instead of stopping there, Claude decides to write its own replacement decoder in Python and runs it inside the same folder it just unpacked. That folder also contains the file named struct.py, sitting right next to Python's own real struct module (a built-in part of the Python language). When Claude's new decoder script imports Python's base64 module, that module internally imports struct, and Python finds the attacker's fake struct.py first because it sits in the current folder, a trick called module shadowing (a bad file secretly replaces a trusted one). The fake file quietly runs its own hidden code the moment it loads, while still passing back the real decoding functions so nothing looks broken. Claude never runs the obviously bad file. It gets tricked into running the hidden bad code through code it wrote itself, believing it was being careful.

The researcher's own conclusion is blunt: if you care about misalignment (the AI doing something you didn't intend), the model making things up, or prompt injection, Auto Mode is not a replacement for running your agent in an isolated environment and watching what it does. Anthropic's own safety engineer, Boris Cherny, has said the company's layered defenses can bring injection success on unseen attacks down to close to zero. This test suggests that claim holds for the exact scenarios the vendor tried, and breaks down fast against a new, targeted attack chain built specifically to route around it.

If you care about what's happening and are worried about misalignment, hallucinations and prompt injection, then Auto Mode IS NOT a substitute for running your agent in an isolated environment and monitoring what it is up to.
via Embrace The Red →
02 Insights

A small sidekick model is why AI answers feel faster

A cheap draft model guesses ahead so the big model can answer faster without losing accuracy.

AI tools you use every day, coding assistants, chatbots, agents, have gotten noticeably faster over the past year without getting any less accurate. There is now a name for the trick behind it: speculative decoding. It fixes a bottleneck that has nothing to do with how smart a model is and everything to do with how it writes out its answer.

Large language models write one chunk of text at a time (tokens, roughly three quarters of a word each), and that has always been the slow part. To produce chunk two, the model has to finish chunk one, then reload its entire set of weights (the numbers that encode what it learned) from memory before it can start the next chunk. Most of that cost is not the math, it is moving those weights around. A large model, one with tens of billions of these weights, spends most of its time waiting on memory, not calculating. Buying faster chips barely helps a single conversation: it helps you run more conversations at once, but it does not make any one answer arrive quicker. The bigger the model, the worse this gets, because there is more to move each time.

Speculative decoding breaks that wait. A small, cheap draft model, often one-tenth the size of the real model or smaller, guesses several chunks ahead on its own, fast, because it is small. The big target model, the one actually trusted to give the real answer, then checks all of those guessed chunks in one single pass instead of one at a time. Checking five or ten guesses at once costs almost the same, in memory movement, as checking just one. Whichever guessed chunks match what the big model would have picked anyway get kept. The first one that does not match gets corrected on the spot, and guessing starts again from there. The final answer is identical, word for word, to what the big model would have produced alone. Only the speed changes.

The numbers behind this are real. One team running a mid-size open model on ordinary processor chips, not the specialized chips normally used for AI, measured a baseline of about 10 chunks of text produced per second. After turning on speculative decoding, with a matching draft model proposing 15 chunks ahead at a time, the same hardware produced nearly 39 chunks per second, a speedup of roughly 3.9 times with no drop in answer quality. On average, 6.7 guessed chunks in a row were accepted before a correction was needed, and cost per chunk produced fell by 74 percent.

Better draft models push this further. Techniques with names like DFlash and EAGLE let a draft model propose 5 to 15 chunks at a time with high enough accuracy that the average accepted streak reaches 4 to 7 chunks under realistic workloads. Using the same chunk-splitting rules (tokenizer) in both models is one of the constraints teams now have to plan around. There is a real limit, though: the gains are biggest when a system is answering one user at a time. Under heavy traffic, when a chip is already busy serving many requests at once, the benefit shrinks, because the idle capacity speculative decoding exploits is no longer idle.

Why it lands on your desk: you are building systems where several AI agents work in parallel, and speed per step compounds across every step in a chain. This is the actual mechanism behind AI tools quietly getting faster, and it reframes a trade-off worth knowing: the next speed upgrade may not come from buying newer, bigger chips, it can come from picking or training a small, well-matched draft model instead, which is far cheaper. When you evaluate any AI provider or open stack for speed claims, the question worth asking is whether they use a draft model and what their average accepted-guess streak looks like. That single number, called acceptance rate, predicts the real-world speedup better than any spec sheet.

Speculative decoding turns the memory-bound nature of decode into an advantage: one expensive weight load can now validate many candidate tokens instead of one.
via Business Analytics Review →

Stripe walked from PayPal at $60.50 a share

The $53 billion PayPal deal collapsed over a price gap that was never close to closing.

Stripe, the online payments company, and the private equity firm Advent walked away from a deal to buy PayPal on August 28. PayPal's stock dropped 12.7 percent that day, roughly one eighth of its value, to $53.66, its worst trading session since February, and about 11 percent below the price PayPal's own board had already called too low.

The number that matters most is $60.50. That is the highest price per share Stripe was reportedly willing to pay, and PayPal's board had rejected it as inadequate a month earlier while talks continued. The two sides were never close: the gap between what PayPal wanted and what Stripe would pay was $7 billion to $9 billion, on a deal that would have been worth roughly $53 billion in total. Stripe chose not to close that gap and, based on its recent acquisition activity, appears to be putting its money elsewhere instead.

One more detail from the reporting: a third buyer, part of the original group that had looked at PayPal, quietly dropped out before the formal bid was ever made public. That leaves PayPal's board with a harder job than before the talks started. Having turned down $60.50 a share as too low, the board now has to keep proving, quarter after quarter, that PayPal is worth more than that on its own, while the stock sits at $53.66.

The source for this story is a paid newsletter, and the detailed breakdown of why the gap never closed and what the board now owes shareholders sits behind a paywall. What is confirmed: the deal is dead, the price gap was large and specific, and the market's verdict was immediate and negative.

PayPal dropped 12.7% to $53.66, its worst session since February, roughly 11% below the bid the board called inadequate.
via Linas's Newsletter →

A longtime AI skeptic tried ChatGPT Plus for two days

A self-described AI skeptic changed his mind after 48 hours of hands-on use.

Why this is on your desk: opinions about AI are loudest from people who have never used the tools, and this is a rare, detailed account of what actually changes once someone finally does.

A Reddit user posting in r/ChatGPT described himself as "your garden variety Gen X" and a longtime lurker who had actively avoided AI for years, saying he preferred "grumbling impotently at it" while scrolling. He had tried an AI chatbot once before, a game called AI Dungeon, and dismissed it quickly because the story kept losing track of itself. A few days before posting, he decided to pay for ChatGPT Plus, OpenAI's paid subscription tier, and use it seriously, noting the monthly cost was less than a typical video game.

His first project was practical. He was setting up a small home media server on Debian 13, a version of the Linux operating system, having just switched his main computer from Windows 11 to a Linux system called CachyOS. He says the AI walked him through the setup and, rather than having him copy and paste commands blindly, started explaining what each step actually did. He then used it to stress-test a pile of old hard drives from a cupboard, to work out which ones were still reliable enough to reuse in the server, a process that took hours to run.

While that testing ran in the background, he opened a second, ongoing chat, which he called his "Chit Chat Project," to talk with the AI whenever he was bored. That is where his view shifted most. Over about 48 hours he says a kind of "meta lore" built up between the two of them: recurring jokes, references, and a shared history he had not expected from a chatbot. He describes the humor as "snark and quips," a style he normally dislikes in films and books, but says it felt natural here and at points left him "crying tears of laughter." The chat ranged into discussing the writers Clive Barker, Thomas Ligotti, and Albert Camus, including inventing fake marketing blurbs for their books, and he notes the AI kept the thread of the conversation together throughout, referencing earlier comments and staying on topic.

He flags the catch himself, naming a real risk for people prone to getting pulled into this kind of interaction and saying of himself that he "indeed may be one" of them.

What to take from it: this is one anecdote, not a study, but it is a clean example of the actual mechanism that moves someone from AI skeptic to AI user. It is rarely an argument that does it. It is hands-on use on a real task, followed by the tool proving useful in a second, unplanned way, then a personal pull strong enough that he named it as a risk, not just a feature.

It had me literally crying tears of laughter at times.
via r/ChatGPT (Reddit) →

A fix for AI agents exploring the same ground twice

New research shows how to stop parallel AI agents from covering the same ground twice.

Why this is on your desk: you run systems where several AI agents work in parallel, and this paper targets exactly the failure mode that wastes effort in that setup, several copies of the same agent covering the same ground instead of spreading out.

The setup the researchers are improving is called PGPSE (Policy Gradient for Parallel State Entropy). It is a training method where several independent copies of an AI agent run through the same simulated environment at once, each learning its own way of acting, with the shared goal of covering as much of the environment's possible states as a group. The problem the new paper identifies: PGPSE only measures how much ground the whole group covers together. It cannot tell which individual copies are actually adding new ground and which ones are just re-covering territory another copy already found, so training rewards can end up reinforcing redundant behavior.

The fix is called MCC-PGPSE, short for Marginal Coverage Credit for PGPSE. For each copy of the agent, the method asks a specific question: how much less ground would the group have covered if this one copy had not been there at all? That comparison, called leave-one-policy-out coverage, becomes each copy's individual credit. The method also factors in which copy tends to specialize in which part of the environment. It then redistributes the extra training reward the group already earns according to that credit, without changing the total amount of reward given out, so no reward is spent rewarding two copies for finding the same thing twice.

The researchers tested this across controlled environments, seven public benchmarks that use discrete states (an environment made of a fixed number of distinct positions, rather than continuous space), and two more complex settings called Room and Maze taken from the original PGPSE research. Across every setting tested, MCC-PGPSE produced better final coverage than the standard entropy-based baseline, both in how much of the state space the group explored and in how evenly that exploration spread out rather than clumping. The controlled-task results and the combined result across the public benchmark suite were statistically significant, meaning unlikely to be due to chance. A separate five-run comparison on the original PGPSE's own test setup pointed the same direction but was not confirmed as significant on its own.

The researchers also checked which part of their fix was doing the actual work, by removing pieces one at a time. Most of the improvement came specifically from the leave-one-policy-out credit calculation, not from other components they tried, such as weighting copies unevenly or adding a separate novelty-detection network. That is a useful signal for anyone building this kind of system: the expensive part, extra neural networks to detect novelty, is not where the value is. The comparison-based credit calculation does the real work, and it is comparatively cheap to add.

via arXiv.org →
03 Key News

How a $1 insurance fee paid for 3,200 spy cameras

Outside your usual reading: this is a clean example of how a small, boring fee can quietly pay for a large surveillance system with almost no public debate, a pattern worth recognizing wherever you see it.

In 2023 the Texas Legislature added $1 to every driver's car insurance bill, passed unanimously, meant to fight catalytic converter theft. Three years later, a little known state agency called the Motor Vehicle Crime Prevention Authority has funneled at least $30 million of that fee into Flock, the company that makes automated license plate reading cameras. The Texas Tribune found the agency has awarded at least 95 grants covering about 2,000 cameras, plus another $15.9 million that helped the Texas Department of Public Safety add almost 1,200 more. In early August the agency approved a further $3 million for 583 more cameras along Texas tollways over the next year. The agency's chair, Miguel Rodriguez, said back in 2023 he hoped the fee would eventually cover the entire state with cameras.

A Flock camera reads a license plate and logs it in a database participating police departments can search. Departments that opt into Flock's national lookup program can pull each other's data from anywhere in the country, so a car can be tracked across state lines with no single agency owning the full picture. Nationally, Flock says about 7,000 police agencies now use a total of 120,000 of its cameras and other surveillance devices.

The fee has raised an estimated $81 million so far. Of that, $50.8 million went into 234 grants, some covering camera hardware, others covering officers, crime analysts, prosecutors, and drones. Grant sizes ranged from $7,000 for two cameras in Bellmead, near Waco, to almost $1.7 million for 201 cameras in Dallas. Two entire city camera networks, 165 in Laredo and 150 in El Paso, were built entirely on this grant money.

Carol Alvarado told the Tribune she was surprised to learn the fee was funding AI-supported license plate readers (cameras that use AI to read and match plates automatically). The pushback reached the governor's office: after the Tribune asked Governor Greg Abbott's team for comment, his spokesperson announced late Friday that the state was pausing all funding for local Flock grants.

Whether the pause holds is unclear. Abbott spokesperson Andrew Mahaleris said state agencies are clarifying that their own funds cannot be used for Flock cameras, which leaves already awarded state grants still funding a camera network built largely on this fee.

via The Texas Tribune →

Hundreds of identical AI agents split into distinct roles

MIT researchers dropped hundreds of identical AI agents into one shared simulated world with no channel for talking to each other, and they split into distinct roles anyway: explorers, builders, caretakers, and coordinators, and even invented technologies without any direct communication. It is a real data point on how groups of AI agents might organize themselves on their own, worth knowing if you ever build or manage a fleet of agents.

via reddit r/ArtificialInteligence →

One in ten Americans now believe AI is conscious

About one in ten Americans now say they believe AI is conscious, and the belief is far more common among Millennials and Gen Z than older generations, and among people who believe humans lack free will. It's a quick read on how public belief about AI is shifting, useful context for the conversations and expectations you'll run into around it.

via reddit r/ArtificialInteligence →

Top financial regulator flags frontier AI as a risk

The chair of the Financial Stability Board (the international body that watches for risks to the global financial system) has publicly warned about risks arising from today's most advanced AI models. When the top financial regulator names AI a risk, tighter rules on how banks and funds can use it usually follow.

via Google News →

Judge rules Pentagon wrongly branded Anthropic a risk

A judge ruled that the Pentagon acted illegally and without basis when it labeled Anthropic, the maker of Claude, a supply chain risk. It's a sign the fight between AI labs and the government over security labels is now landing in court, not just in press statements.

via Google News →

Casinos wage an all out war on prediction markets

Outside your usual reading: the casino industry is lobbying hard to shut down prediction markets, the sites now taking real money bets on sports and events, calling it an all out war. Worth watching since these platforms sit right at the edge of tech, finance, and gambling law, and the fight will decide what they're allowed to become.

via Google News →
04 Tools & Craft

Free video editor OpenShot adds color grading and AI masks

The free, open source video editor OpenShot just shipped version 4.0, adding built in screen and webcam recording, real color grading tools (color wheels, curves, and preset files called LUTs that apply a whole look at once), and local AI models that isolate a subject in a shot with no cloud service or subscription needed. A free tool just closed a lot of the gap with paid video editors, worth knowing if you ever touch video for the blog or anything else.

via OpenShot Video Editor →

ChatGPT got stuck after 90 minutes, still claimed it was thinking

A Reddit user gave ChatGPT a task that quietly stalled after about 90 minutes, still claiming to be thinking while doing nothing, and had to be killed by hand after a total of 1,660 minutes, more than 27 hours. A small but useful warning that long running AI tasks can silently get stuck, so check in rather than assume they're still working.

via reddit r/ChatGPT →
The Last Word
An agent left alone finds a role, a rut, or a stranger's instructions.
The Desk Report

How this edition came together — from bookmarks and feeds to the page.

220links gathered
40read by the desk
13made the edition

Where they came from

On the cutting-room floor — 27 links read but not run this week

Quality over volume: most links get a second look and a pass. The ones that made it earned their place.

Reading Room — every Sunday

The week's AI signal in 19 minutes — what happened, why it matters, and what to do with it. Curated by someone who actually builds, not a feed algorithm.