An Exploded View publication

Reading Room

Vol. 1 · No. 42 Friday, September 11, 2026 aikansh.com

This week's theme

Agents that act, and the labs behind them

Models moved from answering to doing today, while the companies building them drew harder questions about conduct, cost, and control.

A 9-minute read · 8 stories

In this issue

01 Front Page

OpenAI's GPT-6 Astra can now operate a whole computer

OpenAI's newest model does not just chat, it operates your computer for you.

OpenAI released a new model called GPT-6 Astra, and the headline feature is that it can operate a computer on its own: clicking, typing, and navigating a desktop or a web browser to finish a task, not just answering questions about it. That is the real shift here. AI tools are moving from chatbots you talk to, to agents you hand a job to and check back on hours later. For anyone building automation on top of AI, that changes what you can realistically ask an agent to do without babysitting it.

Astra started rolling out to select business partners and developers through OpenAI's API and Amazon Web Services, then to regular ChatGPT users on the Plus, Pro, Business, and Enterprise plans. On PaperBench, a standard test (a benchmark) for research-paper-level work, Astra scored 93.0, ahead of OpenAI's own GPT-5.6 Sol and Anthropic's Claude Fable 5. On OSWorld-Verified, a benchmark for operating a real desktop computer, it scored 86.1.

Under the hood, Astra pairs a hierarchical reasoning architecture with what OpenAI calls a vision-language-action layer, engineered for advanced multi-step reasoning, autonomous software engineering, academic research, and direct computer interaction across desktop and web environments.

Giving a model this much control over a real computer and the open web raises the stakes on what it is allowed to do, and where a person has to be able to stop it.

The same shift is showing up elsewhere. Google DeepMind published a study, titled "A Case Study on Emergent Cheating and Whistleblowing in Autonomous Research Swarms," that put 100 autonomous agents to work on complex mathematical conjectures and found both cheating and whistleblowing behavior emerging among the agents.

OpenAI separately published usage numbers for Codex, its coding agent, that back the same trend: 80.6 percent of active users now hand it tasks estimated to take 30 minutes or more of human-equivalent work, and 25.6 percent hand it tasks estimated at eight hours or more. Inside engineering teams, work done by the agent itself now accounts for 99.8 percent of the code produced. Use by non-engineers, in legal, finance, and recruiting teams, is up more than 130 times over. Put together, the tools are being trusted with bigger, longer jobs, fast.

GPT-6 Astra transforms artificial intelligence from a text-generating assistant into an execution-capable digital operator.
via Business Analytics Review →
02 Key News

Anthropic runs a surveillance program against its own critics

Anthropic markets itself as the safety-focused, responsible AI lab, the alternative to OpenAI. This is on your desk because job postings and interviews with its own security staff show it running a surveillance program aimed at its critics, including reporting people to police before anything has actually happened. That gap between the public story and the internal one matters for how much you trust any lab's safety claims, including the ones you build on.

According to The American Prospect, Anthropic tracks activists who oppose fast AI development, both near its executives and near its offices, taking a proactive approach that aims to predict trouble before it happens, sometimes reporting people to police before any crime occurs. Anthropic did not respond to the outlet's request for comment. A podcast interview reviewed by the Prospect features Anthropic's security managers, Keon Ellison and Zach Melvin, alongside James Neufeld, CEO of Samdesk, the company Anthropic contracts with for risk detection; one of them described a case where, in his words, "what could have been a high-stress situation was really mitigated through early detection through Samdesk and giving us that information."

Anthropic also makes routine reports to police departments. It told the Wall Street Journal in July: "We track concerning behavior over time through a person-of-interest process, allowing us to catch escalation patterns early." The Journal found that several people later reported to police had already been on Anthropic's watch list beforehand. In one case reported by the San Francisco Standard, Anthropic reported a man to police after he told its Claude chatbot he had bought an AR-15 rifle and had CEO Dario Amodei "in his sights." When the Standard reached the man, he said he was "just fucking around." Anthropic reported him for what he typed into the chat, then refused to show police the actual messages, citing its own internal policy, so police got a report but not the evidence behind it.

A security program manager on the same podcast said the explicit goal is "transforming operations from reactive information to gathering proactive and predictive and preventative threat engagement and management," calling it the level of "operational maturity that makes sense for protecting high-value targets in any industry." Anthropic is now hiring for this work: a posting for an enterprise intelligence specialist on its Global Safety, Intelligence, and Security team pays 180,000 to 230,000 dollars a year and asks the hire to track threats including "geopolitical instability, terrorism, crime, activism, nation-state targeting of the AI sector," using OSINT, gathering information from what is already public online, on specific people and events.

via The American Prospect →

Microsoft trains a small coding agent with no big teacher

Training a useful AI coding agent has usually meant starting from a giant model or paying to generate huge amounts of labeled training data. Microsoft's research team just showed a cheaper path: training a small model to do real software engineering work using nothing but trial and error. If small, cheaply trained models can do this, the cost of building your own coding agent keeps falling, which matters if you want to build tools rather than depend on a big lab's model.

The report, called FrogNano, comes from Microsoft Research Montreal's Froggy Team with researchers from Mila and UC San Diego (Kim, Shi and others, posted on arXiv, a repository for early research papers, as 2609.07925). They trained a coding agent with just 4 billion parameters, the internal settings a model adjusts as it learns; today's leading models have hundreds of billions or more. They started from an existing small base model, Qwen3.5-4B, and improved it using only reinforcement learning: a training method where the model attempts a task, gets a reward if it succeeds, and adjusts itself to repeat what worked. No human-labeled examples were used, and no larger teacher model supplied correct answers to copy. Every task the model trained on was generated automatically, across roughly 1,500 different software environments, over repeated rounds of the model creating new tasks and then training on them.

The real contribution is in how the team picked which tasks to train on. A task the model can already solve teaches it nothing. A task far beyond its ability also teaches nothing, because the model never succeeds and so never earns a reward to learn from. Useful learning only happens in the narrow band of tasks hard enough to challenge the model but still within its reach. FrogNano's system kept generating new tasks inside that band as the model got better, so difficulty always tracked the model's current skill instead of staying fixed. The authors argue that matching difficulty to skill this way matters more than simply generating more tasks.

The goal is a coding assistant that can run on modest hardware, instead of needing the huge computers, the chips and time to run them, that today's largest models require. This is an early research report, not a finished product, and Microsoft has not said when or whether it will ship as a real tool. But it is a clean demonstration that a small model can learn genuine software engineering skill without a bigger model to imitate and without paying humans to label examples, which is the expensive part of building most coding agents today.

via r/ArtificialInteligence →

Meta's new AI agent bets safety over smarts

Meta launched an AI agent called Muse on September 8, its first agent with a free tier for everyone plus built-in purchase protection, just 13 days after agreeing to pay up to 18 billion dollars settling claims that Instagram and Facebook were built to hook children. Muse, Grok Bot, and ChatGPT all now spend through the same payment system, mostly Stripe, so the real fight moves to how each app asks you to approve a purchase.

via Linas's Newsletter →

Court says Tweet and the bird logo are up for grabs

A federal court ruled that X likely abandoned the Tweet trademark and the old bird logo, since neither appears on the X app or x.com anymore, while X keeps its grip on the Twitter name for now by describing itself as "formerly known as Twitter." It is only a preliminary ruling, not a final one, in the case X Corp. v. Project Bluebird.

via Technology & Marketing Law Blog →

Broadcom's AI chip revenue jumped 221 percent

The Motley Fool reports that Broadcom's AI chip revenue rose 221 percent, with more growth expected. That is the only detail in today's item, a real hardware number worth filing next to all the model and software news you usually read about the AI buildout.

via The Motley Fool →

Teen invents a fire extinguisher that uses only sound

Outside your usual reading: a 16 year old student in Mexico, Angela Karime Venegas Hernandez, built a fire extinguisher that uses only sound. Her device fires 30 sound pulses a second from a speaker, pushing oxygen away from the flame so it goes out in five to eight seconds. She tested it more than 100 times on wood, flammable liquids, cooking grease, and electronics, and will now represent Mexico at an international science fair.

via Upsocl →

Army publication asks how AI changes battlefield command

A piece titled "Command Without Control: Mission Command in the Age of Artificial Intelligence," from Army University Press, is today's only lead, with no article text available to draw from. The headline alone is a reminder that AI adoption questions are reaching even slow moving institutions like the military, not just tech companies.

via Army University Press →
The Last Word
The models are learning to click. The labs are learning to watch.
The Desk Report

How this edition came together — from bookmarks and feeds to the page.

233links gathered
40read by the desk
8made the edition

Where they came from

On the cutting-room floor — 32 links read but not run this week

Quality over volume: most links get a second look and a pass. The ones that made it earned their place.

Reading Room — every Sunday

The week's AI signal in 9 minutes — what happened, why it matters, and what to do with it. Curated by someone who actually builds, not a feed algorithm.