An Exploded View publication

Reading Room

Vol. 1 · No. 31 Wednesday, August 26, 2026 aikansh.com

This week's theme

Trust is running on the honor system

Today is about verification: who swapped your model, who actually wrote the op-ed, and why a plain system you can check beats a clever one.

A 16-minute read · 15 stories

In this issue

01 Front Page

Skip the fancy AI search stack, most teams don't need it

A simple keyword search often beats an expensive embeddings pipeline.

You are designing the search and memory system for memory-os right now, so this is a direct check against building something fancier than the job needs.

The piece argues that most teams jump straight to the complicated version of retrieval, called RAG (retrieval-augmented generation, meaning the AI looks things up in your own documents before it answers). They reach for embeddings (turning text into numbers so a computer can compare meaning), vector databases, and reranking pipelines (a second pass that reorders results by relevance), when their users are really just typing something like "how do I reset my password."

The author says the right method partly depends on how people search: keyword-style questions should start with plain search, conversational questions benefit from embeddings, and a mix of both needs a blended approach.

The starting point, called the MVP (minimum useful version), is plain keyword search: the BM25 method, or tools like Elasticsearch and Postgres full-text search, the technology that predates embeddings. It is built for cases like a user typing "pandas merge dataframe" or searching for an exact string like "invoice #12345." It costs nothing per search, answers in under 10 milliseconds, and is easy to debug because you can see exactly why a document matched. It needs no decision about how to split documents into chunks, no complicated way to test whether it is working, and no risk of a model being retired out from under you, since BM25 itself never changes. Its weakness is that it misses synonyms and struggles with a conversational question like "how do I fix this." The author says it still handles a large share of real use cases on its own and should not be skipped.

Here is why that first step is underrated: jumping straight to embeddings forces you to answer questions you don't need yet, like what size to cut documents into, how much those chunks should overlap, whether to split by meaning or by fixed size, and how to even tell if your split was any good. Plain keyword search skips all of that. Your documents stay whole, and search just works on them directly.

The next step up uses an AI model to clean up a messy question into good keywords, at a cost of roughly a tenth of a cent per search (the source's example uses GPT-4o-mini, a small, cheap OpenAI model, priced near $0.001 per search). The insight behind it: most "semantic search" problems are really just a badly phrased question, not a limitation of keyword search itself. The AI model can strip filler words ("how do I" disappears). You can also give it your own glossary of terms up front. The example given: a company has an internal tool called "Atlas." A general-purpose embedding model reads "Atlas" and thinks Greek mythology or maps, scoring the real match at just 0.15 out of 1, essentially useless. Telling the AI model in advance, through a short instruction, which words are exact domain terms to preserve (Atlas: the internal data processing framework, Mercury: the messaging system, Zeus: the auth service) gets a perfect keyword match instantly, with no retraining needed.

This approach is also more forgiving to fix. With embeddings, a bad result means adjusting how you split documents, redoing the embeddings for the whole collection, and rerunning tests to see if it actually improved. With query rewriting, a bad result just means editing the instruction and testing again immediately. The source also describes a multi-turn version: the AI model rewrites the question, searches, checks whether the results look good enough, and if not, tries again with what it learned, up to a handful of times, all without touching the underlying documents.

The most advanced version blends both: keyword search pulls the top 50 to 100 candidates, then embeddings rerank them down to the best 10, because keyword search is fast and precise on exact terms while embeddings are better at catching meaning.

The lesson for memory-os: start with the plain keyword version, and only add the expensive layers once real usage data proves you need them, not before.

via Lighthouse Newsletter →
02 Also on the Front Page

Anthropic's CEO admits people don't trust AI firms

Dario Amodei broke his usual silence on social media to explain why.

You build on top of Claude every day, so how much people trust the company behind it is not background noise. It is part of the ground your own work stands on.

Anthropic's cofounder and chief executive, Dario Amodei, posted a long thread on X this past Saturday. That is unusual for him, since he mostly stays off social media. He was responding to the idea that he personally is responsible for the public's overall sense of doom about AI. He pushed back on that framing, but admitted there is a real trust problem.

He said people are told there are only two choices: regulate AI and let a few companies capture the market and concentrate power, or spread AI everywhere, including releasing the model files publicly (open weight models, meaning anyone can download and run them) to keep the technology in check on its own. Amodei called this a false choice. He pointed to institutions like the court system as proof that power can be spread out without giving up on rules. He also noted Anthropic has backed policies that slow down the companies building the most advanced models (frontier AI companies) while giving smaller rivals a leg up.

Still, he admitted AI as a technology tends to concentrate power in fewer hands. He said this is not really about regulation. It comes from AI scaling laws, the pattern where a model gets better the more computing power and data you pour into building it. Open weight models help a little, but they do not fix the underlying problem. They just move the concentration of power to whoever owns the most computing hardware and chips.

For you specifically, this matters because trust is not abstract when your own tools and workflows sit on top of Anthropic's models. If big customers or governments start to see Anthropic as just another company chasing power, that shapes the rules, prices, and access you get downstream. Amodei naming the problem in public, instead of only talking about capability and safety, suggests the label is landing hard enough that even the labs feel they have to answer it.

Amodei pointed out that institutions like the court system can decentralize power, while noting Anthropic has been in favor of policies that slow down frontier AI companies and also give smaller rivals an advantage.
via Reddit (r/ArtificialInteligence) →
03 Tools & Craft

Apple's new Mac chip can run giant AI models at home

The M5 Ultra packs 512GB of memory, enough to hold models that used to need a rack of servers.

This lands on your desk because you keep circling back to running AI models locally and privately, off someone else's servers. Apple's new M5 Ultra chip, the top option in the redesigned Mac Studio, can hold up to 512GB of unified memory (memory built directly into the chip, shared with the processor) moving data at 1.2 terabytes per second. That is enough room to load models that used to need a room full of server-grade chips.

Real examples from the source: DeepSeek R1, a 671 billion parameter model (parameters are the internal settings a model tunes during training), Kimi K2.6 at 1 trillion parameters, and DeepSeek V4 Flash all fit and run at usable quantizations (a technique that shrinks a model's numbers to save memory, at some cost to precision). Until now, models of that size meant racks of GPUs (graphics chips, the hardware that does the heavy number crunching for AI) sitting in a data center, not a box under a desk.

This will not change how most people use a chatbot app. But it matters for researchers, developers, and companies that want serious AI power without sending private data to someone else's servers or paying per use. Self-hosting a huge model no longer requires building and maintaining a complicated multi-chip server rig.

Not every new Apple chip is going this direction. The M6 Mac mini tops out at 32GB of memory, keeping it firmly in the small-model category despite faster AI hardware. Local AI machines seem to be splitting into two lanes: cheap boxes that run small models well, and high-memory workstations, like this Mac Studio, that now run models once reserved for server rooms.

via r/ArtificialInteligence (Reddit) →

Travel firm lets non-engineers ship code with Codex

loveholidays turned product managers and marketers into builders, and AI-assisted code jumped from 7% to 79% in a year.

This is close to the model you are trying to build across your own projects, where AI agents do the work and the founder just reviews and decides. loveholidays, an online travel agency running across eight European countries, processes 60 trillion possible holiday package combinations every day, and it just gave people outside engineering the power to write and ship real product changes using Codex, OpenAI's tool that writes and runs code from plain instructions. "Everybody is a builder," says Dmitri Lerko, the company's Head of Engineering. "Making changes to our applications, infrastructure, and deploying code is no longer an engineering-only activity." The company says AI-assisted code changes have grown from 7% of all changes a year ago to 79% today, an eleven-fold jump, while engineering headcount stayed roughly flat.

The clearest example is something called Search Playground. Before, if someone outside engineering had an idea for a new way to search for a holiday, they had to convince an engineering team to spend time building a prototype, and every experiment competed with other engineering work for that time. loveholidays' engineers built Search Playground once, using the company's existing design system and Codex, so anyone could turn an idea into a working page themselves. More than ten new search experiences have been built through it, most by people who are not engineers, and at least three are now live on the loveholidays website. One, called Inspire Me, helps travelers browse trip types like beach breaks or food-focused trips. Another came from the marketing team: for a promotion called Crisps from Abroad, they built an interactive microsite themselves in hours using Codex, work that would previously have gone to an outside agency and taken far longer.

The same approach reaches into parts of the business normally locked behind specialist knowledge. loveholidays' data platform and infrastructure, the underlying systems that store and move the company's data, used to require knowing specific internal tools and processes, so any change needed a specialist engineer. Now, engineers write their best practices and safety checks into workflows that Codex walks other employees through step by step, so people can make the change themselves while the checks catch mistakes. Successful AI-assisted changes to the data platform rose from 58% to 93% over the year, and the team now handles four times as many data platform changes for every support request. Across the company's broader self-service systems, the success rate rose from 63% to 90%.

Overall deployment frequency, how often the company ships changes to production, is up 73% over the same period, without adding engineers. CTO Mike Jones frames the goal as building "the general intelligence for travel," meaning combining the company's existing systems with its people's know-how and making both easier to reach through AI. He is explicit that the company tracks this by business outcomes, not by how many people have tool access: "Technology is just a means to an end," he says.

Everybody is a builder
via OpenAI →
04 Insights

Revolut's real AI story is who owns the data

Revolut's new research lab shows proprietary data, not the model, is fintech's real advantage.

This is a useful lens for any of your own ventures: the model you plug in matters less than the data only you own. Revolut, the fintech app, just launched a research unit called Revolut Research, built around a model it calls PRAGMA. PRAGMA is a foundation model, a large general-purpose AI model trained once and then adapted to specific jobs, trained on behavioral data (records of what customers actually did, not just what they said) from 80 million Revolut customers across more than 100 countries. Revolut says the results are strong: PRAGMA is 2.3 times better at spotting customers heading toward credit default (failing to repay a loan), and it catches 65% more fraud cases than before.

The more interesting part is what is not unique about it. Nubank, the Brazilian digital bank, has already put a near-identical setup into production. Visa, Mastercard, Stripe, and Plaid have each independently built their own version of the same idea. Nvidia, the company that makes the chips these models run on, has even published a public blueprint so any other company can build the same thing. None of this was coordinated between the companies. When several of the biggest names in payments arrive at the same design on their own, it says something about where the real value sits.

The lesson: none of these companies are winning because of a better AI model. They are winning because each holds transaction data, records of real customer purchases and behavior, that nobody else has access to, and that data is what actually improves fraud and credit decisions. A rival could rent the same kind of model tomorrow. None of them can rent 80 million customers' worth of real transaction history.

via Linas's Newsletter →

Half of US adults say they don't understand statistics

Outside your usual reading: a nationwide survey of 1,000 US adults found a quarter say they have no understanding of statistics at all, and another 37% call their grasp limited. Nine in ten said they would use statistics more often if they understood the topic better. Worth remembering whenever you or your agents explain a data-backed decision to someone outside a technical room.

While statistics are not hard to understand, they are even easier to misunderstand
via Penn State News →
05 Key News

Druckenmiller admits AI wrote his Wall Street Journal op-ed

This is a real sign of how normal AI-written content is becoming, even in a place as high-stakes as the Wall Street Journal's opinion page, which is exactly the ground your own writing work sits on.

Stanley Druckenmiller, the billionaire hedge fund investor, admitted that a recent Wall Street Journal opinion piece under his name was written entirely by AI. The piece attacked U.S. Treasury Secretary Scott Bessent's plan to expand a bond buyback program (the government buying back its own older debt) by $1 trillion, calling it a "doomed price control."

The Wall Street Journal defended its decision to publish the piece once the AI authorship came out, saying the core arguments in it belonged to Druckenmiller even though he did not write the words himself.

The admission set off a debate in the tech world over major public figures using AI models to ghostwrite opinion pieces that can move markets and shape policy debates, and whether that quietly changes what a byline is supposed to mean.

For you, it is a preview of the norm your own AI writing projects already operate inside: the words can come from a model as long as a real person stands behind the argument. The open question the debate raises is whether readers should be told when that happens, something neither Druckenmiller nor the Journal appear to have done up front.

via Reddit (r/ArtificialInteligence) →

New trick makes AI models answer 4.5 times faster

Researchers built KVBoost, a way to reuse an AI model's cached computations even when repeated text shows up in the middle of a request, not just at the start. Tested on a mid-size model, it cut the wait for the first word of an answer from 639 milliseconds to 142 milliseconds, a 4.49 times speedup, with no drop in accuracy.

via arXiv →

A Claude user says the model swapped without telling him

Outside your usual reading: a Reddit user says his Claude conversation quietly switched from the model he expected to a different one, and the interface only admitted the swap when he asked directly. You run a lot of your own work through Claude daily, so this kind of user complaint is worth more than official release notes.

via Reddit (r/claude) →

A tamper-proof paper trail for AI decisions

A new protocol called AIREP proposes a signed record, checkable offline by any party, every time an automated AI system releases, blocks, defers, redacts, or escalates something it produced. Worth remembering as you give AI more autonomous decisions inside your own systems.

via arXiv →

Left and right team up against AI data centers

Communities on the political left and right are joining forces to oppose new AI data centers being built near them, according to a report from NBC Chicago.

via Google News (NBC Chicago) →

Chinese lab's stealth model reportedly rivals DeepSeek

Z.ai's Ox Alpha, described as a stealth model, reportedly rivals DeepSeek, according to a report flagged on Hacker News.

via Bloomberg →
06 Learnings

Odds ratios are not the same as probability

Most people read an odds ratio (a way of comparing how likely two outcomes are) as if it were a straight percentage change, and that mistake skews decisions built on models used for credit, health, or customer churn. An odds ratio of 2 does not mean "twice as likely": the same multiplicative change in odds produces a different shift in probability depending on where you start. Convert to plain probability before trusting the number.

Always convert back to the probability scale at the relevant baseline before claiming practical impact.
via Business Analytics Review →
07 From the Timeline

Founder says his AI agent nearly pays for itself

@sairahul1 (Rahul) posted on X on August 24, 2026: "My Grok Bot just paid its own salary." He gave the AI agent, Grok Bot, a paid AI assistant that runs tasks on its own cloud computer, one job: publish one SEO tool page (a small web page built to rank in search results) every day. It finds keywords, builds the page, publishes it, logs it, then queues tomorrow's run overnight, almost covering its own cost.

@sairahul1 on X →
The Last Word
The model swapped, the op-ed was ghostwritten, and nobody thought to mention it.
The Desk Report

How this edition came together — from bookmarks and feeds to the page.

226links gathered
40read by the desk
15made the edition

Where they came from

On the cutting-room floor — 25 links read but not run this week

Quality over volume: most links get a second look and a pass. The ones that made it earned their place.

Reading Room — every Sunday

The week's AI signal in 16 minutes — what happened, why it matters, and what to do with it. Curated by someone who actually builds, not a feed algorithm.