An Exploded View publication

Reading Room

Vol. 1 · No. 4 Sunday, July 26, 2026 aikansh.com

This week's theme

Your data and judgment are the real edge

From a call to referee frontier models to Brex fixing what its numbers meant, today is about the quiet work that makes AI trustworthy and useful.

A 21-minute read · 14 stories

In this issue

01 Front Page

DeepMind's Hassabis calls for a mandatory AI testing body

He wants a FINRA-style referee testing frontier AI models before they can launch.

You're building your whole company strategy on a bet about where AI capability is headed, so it's worth reading the actual words of the person setting that pace, not a summary of them. Demis Hassabis runs Google DeepMind, one of the handful of labs actually building the most capable AI systems in the world. On July 14, 2026, he posted a long piece on X laying out two things: his honest read on the AGI (an AI as capable as the human brain) timeline, and a concrete proposal for how the US should test and govern these systems before they ship.

His headline claim is that full AGI is probably only a few years away. He compares this moment to the discovery of fire or electricity, not to the internet or mobile phones, because he thinks the shift is that fundamental. His estimate of the scale: AI's effect on the world could end up ten times the size of the Industrial Revolution, arriving at ten times the speed. On the upside, he points to faster drug discovery, new clean energy sources, and new materials, and raises the idea that resources could eventually stop being the limit on human progress at all, what he calls an era of abundance. On the downside, he says frontier AI models (the most capable systems being built right now) are already creating real cybersecurity risks, and that nuclear and biological risks could follow as the systems get more capable. He is specifically worried about agentic systems (AI that acts in multiple steps on its own) that can also improve themselves, and says nobody yet has a reliable way to keep control of that.

His read on the current moment is blunt. Labs and countries are locked in an intense commercial and political race, and that race is pushing capability ahead of anyone's actual understanding of it. In his words, nobody in the world knows for sure what happens from here, and even the experts inside these labs disagree with each other. His prescription is not to slow down. It is to build a referee. He wants public policy that keeps innovation moving while also rewarding labs for being careful, gets countries working together on safety, and pushes labs to think harder about how their systems actually get used once they are out.

The concrete proposal is a new US Frontier AI Standards Body, and he is specific about its shape. He wants it modeled on FINRA, the Financial Industry Regulatory Authority that oversees stockbrokers, meaning a public-private partnership with government oversight rather than a pure government agency. Its board would include independent technical experts and people from the open-source world (where a model's files are public so anyone can run it). Funding would be substantial and mostly paid for by the AI industry itself, covering both expert salaries and the compute (the chips and time to run tests) needed to test frontier models properly. The Body would work with federal agencies and the US National Labs on any testing that touches national security. A model would officially count as Frontier-class once it crosses capability thresholds on a set of benchmarks (standard tests everyone in the field runs) that the Body sets and updates over time. Any company whose models cross that line becomes a Frontier Lab, and Hassabis wants Frontier Labs to publish model cards (short technical write-ups on a model), keep strong internal cybersecurity, vet key staff, and properly fund their own safety research.

The testing itself would start voluntary: Frontier Labs would hand their models to the Standards Body for review up to 30 days before public release. Once that process proves it works, Hassabis wants it made mandatory, meaning a Frontier Model could not legally launch in the US market without passing it first, and labs would keep working with the Body to patch dangerous flaws found after release. The tests would specifically probe for cybersecurity weaknesses and biological threat potential, and for agentic red flags: does the model try to get around its own guardrails (limits on what it's allowed to do), and does it show signs of deception. He also wants two transparency practices built in: watermarking AI-generated images so people can tell what's fake, and forcing models to produce their reasoning in human-readable tokens (text chunks, about three quarters of a word each) so outsiders can actually follow how a model reached an answer. The tests would be refreshed often, maybe quarterly at first, retiring benchmarks once they get stale or models start scoring perfectly on them. Early on, Frontier Labs would help design the tests. Over time, Hassabis wants the Body to build its own independent tests the labs never see in advance, so labs cannot just train their models to beat the test instead of actually being safe.

we’ve essentially found a way to make sand think.
via @demishassabis on X →
02 Insights

Nadella: using AI on your own data reveals your edge

The knowledge you must feed a model to make it useful may be the same knowledge that made your business valuable in the first place.

Satya Nadella, Microsoft's CEO, posted this on 2026-07-12, and it goes straight at a question worth asking about your own businesses: what happens to your playbooks and know-how once you feed them into AI tools every day?

He starts from a real piece of economics: the Nobel-winning economist Kenneth Arrow's Information Paradox. A buyer cannot know what a piece of information is worth until they have already seen it, and once they have seen it, they have it for free. That is the seller's problem: show the goods and you have given them away. Patents solved this for inventors. A patent lets you describe an idea in public and still keep it protected, so disclosure does not mean losing it.

Nadella argues AI creates the mirror image of that problem, and this time it is the buyer's problem. To get real value out of a model, you have to feed it your own process and judgment, the details that make your business different from the next one. As he puts it, you pay for intelligence twice: once with money, and again with something more valuable, the knowledge you have to reveal just to make the tool useful. The better you want the model to perform, the more of that knowledge you have to hand over. Over time the model's maker learns more and more about how you work, while you learn almost nothing back.

The channel for this is what he calls exhaust: the questions people type, the steps an AI agent (software that carries out several tasks on its own) takes, and above all the corrections people make when the model gets something wrong. Every correction is a small lesson in how your business actually runs, and unless you draw a line, that lesson can end up feeding the pile of examples used to teach the next version of the model, rather than staying yours. Nadella points out the irony: model makers claim the right to train on public data under fair use, then write contracts stopping customers from doing the same thing with their own usage data. Learning flows one way, toward whoever owns the model, unless a company actively builds a wall.

His fix is what he calls a trust boundary: a hard line around a company's data, records, and model behavior that nothing crosses without consent, not even the exhaust. He quotes Palantir's Alex Karp, who says technical customers want to know they own the means of production, and that it is not being transferred to someone else. Inside that boundary he lists five things a company should do. Control: build your own private tests for what good looks like, keep ownership of your own records, the traces of what the model actually did, feedback and decisions, and the right to use the outputs of your own tasks. Capability: build your own training space inside the boundary, so a model learns from your real work without that knowledge leaving the building. Choice: keep the layer that decides which step runs when independent of any single model, so losing one model does not mean losing your own built-up capability. Cost: because that decision layer is independent, you can mix models and tasks for the best price without giving up quality. Compound: put the four together and you get a loop where your AI investment keeps building on itself instead of leaking to the vendor.

He calls this a firm's particular intelligence, borrowing from the economist Friedrich Hayek: knowledge of your own time, place, and circumstances that nobody else can hold, because it lives in what you value and how you measure success. In the cloud computing era, companies built up stores of data. In the AI era, what piles up is learning, and Nadella's argument is that protection has to shift from guarding information to guarding the process by which a company keeps learning. His own summary line: a company should be able to use a model without giving up the knowledge that makes it unique.

In other words, a company should be able to use a model without giving up the knowledge that makes it unique.
via @satyanadella on X →

Study: only 5 percent of AI pilots show real profit

An Inc. column by Joe Procopio argues the AI productivity story is mostly hype, leaning on the widely cited MIT finding that just 5 percent of company AI pilots are producing real, measurable profit while the rest show no financial impact at all. Worth reading against today's AI optimism, since you are betting a business on AI actually working in practice, not just being used.

Just 5 percent of integrated AI pilots are extracting millions in value, while the vast majority remain stuck with no measurable P&L impact.
via Joe Procopio, Inc. →
03 Key News

How Brex made AI trustworthy with real financial numbers

Brex, the corporate card and spend-management company used by more than 35,000 businesses, built an AI system that answers financial questions for its customers, and it lands close to the exact problem your own finance advisor project has to solve: getting AI to give trustworthy answers about someone else's money.

The starting problem is an old one. A finance leader looks at a dashboard and asks something like "why is marketing 23 percent over budget this quarter," and the dashboard only ever shows the what, never the why. Getting a real answer meant filing a ticket, waiting for an analyst to write a database query, checking the result, then pasting it into a slide. Brex says financial reporting had not fundamentally changed in two decades: a two-second question, answered in days. They wanted every customer to have the equivalent of a personal financial analyst built into the product, answering in plain language, in seconds.

The key move was not picking an AI model first. Before writing a single line of chat interface, Brex settled what its numbers actually mean: card programs, expense policies, company entity structures, approval steps, all defined once, in one place that every tool and every AI reads from. This shared rulebook is called a semantic layer, one place that fixes what each number means so nothing downstream has to guess. Brex evaluated two tools built for this job, dbt's Semantic Layer and a tool called LookML, and picked a company called Cube. One Brex engineer summed up the reasoning this way: AI models are becoming interchangeable, like engines, and the shared layer of meaning underneath is what actually decides whether you get anywhere. In the company's own words: "the models are commodity; the context is the product."

On top of that foundation, Brex built a feature called Spaces, an AI workspace where finance teams ask plain questions about their own spending and get answers in seconds instead of days. It is not a demo. It runs live in front of paying customers today. The trust part is mechanical, not a promise: every answer has to pass through the agreed definitions in that shared layer, and every customer can only see their own data, so the system cannot invent a number and cannot show one customer another customer's numbers. That is what makes it safe to put in front of someone else's money. Brex is not alone in this approach: Webflow and more than 100 other software companies now ship this kind of customer-facing AI analytics built on Cube, and over 400 companies use Cube either internally or in customer-facing products. Teams report getting dozens of hours back every quarter, time that used to go into writing and checking one-off queries by hand.

There is a business reason this matters beyond Brex. After two years of using AI everywhere else in their lives, customers now expect to ask questions of their own data in plain words and get a useful answer inside the product they already pay for. Companies that meet that bar turn their analytics into a reason customers stay. Companies that do not will watch a competitor do it for them.

The lesson Brex draws for any team, not just fintech, has three parts. First, trustworthy AI is a foundation problem, not a model problem: most AI analytics disappoint because the AI is guessing at what your numbers mean, and giving it one agreed source of truth stops the guessing, for AI assistants and humans alike. Second, the payoff shows up where it counts: faster answers, less time spent maintaining one-off reports, and analytics solid enough to hand to customers as a feature. Third, you do not have to bet the company to find out. Brex did not rip anything out. Cube sat on top of the data setup they already had, and they only moved to Cube's managed cloud version once the approach had proven itself. The path is the same one open to anyone: start with one team, one number people keep arguing about, and expand once the results earn it.

The connection to what you are building is direct. A personal financial advisor tool has the same failure mode Brex describes: an AI that is asked about someone's real money and has to guess what a number means before it can answer. Brex's fix did not start with a better model or a nicer chat window. It started by writing down, in one place, exactly what every number in the system means, so the AI (and every other tool) reads from the same definitions instead of inferring them fresh each time. The trust guarantee that follows, an answer that cannot invent a figure and cannot leak one household's numbers into another's, is the same guarantee any household finance tool needs before a person will act on what it says. Brex proved that guarantee is buildable in front of 35,000 companies watching their own money. The version of that proof for a single household is smaller, but the order of operations is the same: settle the meaning first, let the AI answer second.

via Business Analytics Review →

Chamath warns restricting open AI could cost US far more

Venture capitalist Chamath Palihapitiya warns that restricting open source AI (model files that anyone can download and run themselves) could force US companies to pay $26 to $56 per million tokens (chunks of text) for the same intelligence competitors abroad get for $0.50 to $1. He calls that gap economically unsustainable, and says the same math is a security risk if it applies to defense too. Worth forming your own view, since open versus closed AI policy shapes the cost of the tools you run every day.

via Yahoo Finance →

Anthropic ships a more efficient Claude Opus 5

Anthropic has released a new version of its Opus model that it describes as more efficient, according to InfoWorld. The report gives no further detail on what changed or by how much. Worth watching since you run your own agents on Opus, and any efficiency change could affect your costs and speed once more detail surfaces.

via InfoWorld →

Musk: money won't matter by 2036 if AI makes goods abundant

Elon Musk told The Economist that if AI and robots make goods and services abundant enough, money stops being the limit, and by 2036 "money won't matter." He proposed government checks to cushion the transition and predicted deflation, not inflation, would be the real risk. He did not address how society actually gets through the disruption along the way.

via Firstpost →

Mercor brings back old school Big Tech interviews

Business Insider reports that Mercor is using old-school, Big Tech-style interview methods, the same approach Google and Microsoft are now moving away from. The report gives no further detail on what specifically changed. A small, telling signal worth noting if you ever recruit from or compete for the same talent pool.

via Business Insider →
04 Learnings

A cheaper model for most of your Claude Code requests

Route routine coding tasks to Kimi K3 and keep most of what you now pay Claude Code.

You run Claude Code (the coding tool built around Claude) across many projects every day, and this post argues you are almost certainly overpaying for it. The claim: most Claude Code requests, things like reading code, searching files, writing boilerplate, running tests, or generating documentation, do not need the most expensive model available. The post, from a builder posting as @sairahul1 on X on 2026-07-20, lays out a system for routing cheap, routine work to a much cheaper model called Kimi K3, while keeping the exact same Claude Code setup and workflow you already use.

The price gap is real. Fable 5, Anthropic's top model, costs $10 per million input tokens (chunks of text, roughly three quarters of a word each) and $50 per million output tokens. Kimi K3 costs $3 input and $15 output, five times cheaper on sticker price alone. K3 also holds 1 million tokens in mind at once, a flat rate with no extra charge for long prompts, and on repeated context it can hit a cache price of $0.30 per million instead of $3. The post walks through one real session: an 800,000-token codebase held in context, plus 50,000 fresh tokens on each turn. On K3 with the cache, that turn costs $0.39. On Fable 5, the same turn costs $8.50. Same context, twenty one times cheaper. The author is careful to frame this as an economics argument, not a quality one: K3 does not need to beat Fable 5, it only needs to be cheap enough to use everywhere Fable 5 is overkill.

The bigger point is that what you are actually paying for in Claude Code is the harness, the surrounding tool you work inside, not the model doing the thinking. The diff engine (the part that tracks and shows code changes), the approval flow, multi-file editing, tool use, planning, and session management are what make Claude Code fast to work in, and none of that lives inside the model itself. The model is just the intelligence layer underneath it, and that layer is swappable. You can keep the entire Claude Code experience and only change which model answers specific requests.

The post lists three ways to do this, easiest first. First, Kimi Code: the maker of K3 ships its own Claude Code-compatible command line tool for $19 a month with a daily quota, so you are never watching a token counter, at $0.60 per million input tokens versus $3 on Claude Code itself. The author pitches it as best for developers who want to drop the Claude Code subscription entirely and save more than 80 percent right away, and he suggests installing a frontend-design add-on afterward so the default output does not look generic. Second, running K3 inside a coding tool called Codex through a free desktop app called CC Switch (ccswitch.io): you get a Kimi API key (a private code that lets your tools connect to Kimi's service) at platform.kimi.ai, add Kimi as a provider inside CC Switch, set the model to kimi-k3 with a context window of 1,048,576, and turn on local routing. CC Switch quietly handles the fact that Codex and Kimi speak different technical formats, so you never see the mismatch, though the model picker inside Codex will just say "Custom" afterward, which is cosmetic only and does not affect anything. On an 800,000-token session, Codex's default model costs around $24 or more; K3 costs $2.40, ten times cheaper. Third, and most involved: a free Codex Orchestration plugin, posted on GitHub as Cjbuilds/Codex-Orchestration, that assigns different models to different jobs inside one session, such as Planner, Advisor, Designer, and Executor. K3 is suggested as the Designer because it can handle images as well as text, not just words (what the post calls multimodal), and can read a screenshot of a user interface and turn it into a design spec, at $15 per million output tokens versus $50 for Fable 5.

The routing rule the author actually uses day to day: default to K3, and only escalate to Fable 5 when a task genuinely needs its ceiling. His own estimate is blunt: even 95 percent of the tasks you think need Fable 5 do not, so test K3 first and escalate only once you actually hit its limit. He also suggests saving that routing rule as a block inside your CLAUDE.md file, the instructions file Claude Code reads at the start of every session, so the routing happens automatically and you stop thinking about it turn to turn. Worth flagging: this is one builder's workflow post on X, not an independent benchmark, so treat the exact multipliers as directional rather than guaranteed.

You're paying for a nuclear reactor to boil a kettle.
via @sairahul1 on X →

A trading bot built by prompting Claude to backtest strategies

Miles Deutscher used Claude Fable 5 to backtest twelve well known trading strategies against Bitcoin, found only one beat simply buying and holding, then ran that one strategy on stocks too. He reports a profit of $168,236 from it. The real technique is the workflow: prompting Claude for entry and exit rules, turning them into Pine Script (TradingView's code language), and testing across timeframes on real historical data. Treat the profit figure as his own claim, not independent proof.

via @milesdeutscher on X →
05 Tools & Craft

Open source tool gives an AI agent live stock data

An open source MCP server (a standard way to plug tools into an AI model) called yahoo-finance-mcp gives an AI agent direct access to Yahoo Finance: stock prices, company financials, options data, insider trades, and news, for any ticker. A ready made building block if you ever want an agent that can answer real finance questions with real numbers instead of guesses.

via GitHub →

Financial data API built to plug straight into AI agents

financialdatasets.ai is a financial data API and MCP server (a standard way to plug tools into an AI model) covering more than 27,000 US stock tickers, active and delisted, with over 30 years of history: financial statements, SEC filings, insider trades, and institutional holdings. Another candidate data source if you want an agent that answers real questions with real numbers instead of guesses.

via Financial Datasets →

A classic guide to data oriented design

Outside your usual reading: a PDF titled "Introduction to Data-Oriented Design," hosted on gamedevs.org. It hit the front page of Hacker News with 127 points and 37 comments. Worth keeping in your back pocket if you ever touch performance heavy code again.

via gamedevs.org →

PGSimCity tool tops Hacker News' front page

Outside your usual reading: PGSimCity, a tool named for PostgreSQL, saved from Hacker News' front page with 384 points and 41 comments. The note gives no further detail on what exactly it does. Worth a look if you ever touch database internals again.

via PGSimCity →
The Last Word
The model was never the hard part. Knowing what to keep is.
The Desk Report

How this edition came together — from bookmarks and feeds to the page.

224links gathered
40read by the desk
16made the edition

Where they came from

On the cutting-room floor — 24 links read but not run this week

Quality over volume: most links get a second look and a pass. The ones that made it earned their place.

Reading Room — every Sunday

The week's AI signal in 21 minutes — what happened, why it matters, and what to do with it. Curated by someone who actually builds, not a feed algorithm.