An Exploded View publication

Reading Room

Vol. 1 · No. 6 Tuesday, July 28, 2026 aikansh.com

This week's theme

From prompting AI to building your own systems

Today leans into building AI you control: small models trained for one job, agent loops that run themselves, and why filling up their context makes them sloppy.

A 23-minute read · 16 stories

In this issue

01 Front Page

Companies beat top AI models by training their own smaller ones

Three real companies proved a smaller, custom-trained model can beat giant general AI models at one specific job.

Here is the direct takeaway for how you build your own AI agents: for a narrow job you do over and over, a small model trained specifically for that job can beat an expensive general purpose model, and cost much less to run each time you use it. Three real companies just proved this, each in a completely different business, and over the past two years the approach has hardened into a repeatable playbook. Step one, take a model whose files are public, called an open weights model, meaning anyone can download it and retrain it themselves. Step two, teach it your own real examples using reinforcement learning, which means training a model by rewarding it when it gets your specific task right and correcting it when it does not. Step three, score it against your own version of the actual job, not a generic public test that has nothing to do with what you actually need done.

Bridgewater Associates, one of the largest hedge funds in the world, has analysts who sift a constant stream of news articles, regulatory filings, and emails, judging which ones matter to the firm's investment views and where the useful content stops and boilerplate starts. The catch is that relevant means relevant by Bridgewater's own internal judgment, built up over decades, and no amount of clever prompting got the big frontier models, meaning the best general purpose AI models from the top labs, to reliably match that judgment. So Bridgewater trained an open weights model on labels created by its own expert investors, essentially teaching the model to think like their analysts already think. The trained model makes roughly 30 percent fewer mistakes than the best frontier model, and it costs a fraction as much to run each time, what the piece calls its inference cost, meaning the cost of getting one answer out of the model once it is trained and running.

Harvey, which builds AI agents for law firms, hit the same wall on its hardest work: due diligence on business transactions and drafting legal memos. These jobs are long horizon, meaning the agent has to take many steps in a row through large sets of documents, and small errors stack up and compound as the steps pile up, the way one wrong assumption early in a memo can poison everything that follows. Even the best frontier models, run at their maximum reasoning effort, the setting that makes a model think longer before it answers, kept falling short of the quality bar that law firms actually need to trust the output. Harvey's fix was to run reinforcement learning on an open weight model trained specifically on legal work, rather than continuing to push a general purpose model harder. The result outperforms both GPT-5.5 and Claude Opus 4.8, Anthropic's own top model, on Harvey's internal scoring rubrics, the standards it grades answers against, on the exact tasks its lawyers actually do.

Intercom's AI support agent, called Fin, resolves close to two million customer issues a week, and at that volume the real problem stops being accuracy alone and becomes unit economics, meaning what is left over once you subtract the cost of each answer from what it is worth to the business. Frontier model pricing per call adds up fast at that scale, and every extra point of issues resolved without a human matters directly to the bottom line. So Intercom's AI team post-trained its own vertical model, named Fin Apex, on billions of real customer service conversations gathered from its own platform, rather than continuing to lean on a general purpose model built for every task at once. Intercom reports that Fin Apex resolves more issues than the best frontier models, while costing less to run per conversation, a combination a general model could not match.

The article says the same shape shows up in eight more deployments beyond these three, collected in an appendix, each running its own version of the same recipe: pick an open weights model, gather your own real task data, and train against a scored version of the actual job instead of a generic benchmark that was never built for your use case. None of these three companies replaced their frontier model everywhere, and none of them are claiming a small model beats a big one in general. They replaced it for one narrow, high volume, well defined job, where they already had enough of their own labeled examples to teach a smaller model the one specific judgment call that mattered most to their business. That is the part worth carrying into your own agent work: before reaching for the biggest, most expensive model on a job you do over and over, ask whether you already have enough real examples of the right answer sitting around to train something smaller and cheaper to do it instead.

The trained model makes roughly 30% fewer mistakes than the best frontier model, at a fraction of the inference cost.
via Fermisense →
02 Also on the Front Page

Private Claude chats turned up in Google search results

Conversations people had with Claude that felt private turned up searchable on Google, Fortune reports. You use Claude for deeply personal work daily, so treat any shared conversation link as public from the moment you create it, and check your own sharing settings.

via Fortune →
03 Insights

Coding's best builders stopped prompting and started writing loops

A viral tweet fight over one word says the job just moved up a level.

One tweet split AI coding builders into two camps this month, and it is worth your time because the fight is really about where your job goes next: from typing instructions to a coding agent (an AI system that writes and runs code on its own) to writing the small program that decides what to tell that agent, and when. The recap here comes from @mvanhorn (Matt Van Horn) on X, posted June 8, 2026. He ran a quick search across the last 30 days of posts to trace the argument properly, and found it everywhere: fifteen Reddit threads and twenty one X posts, most of them repeating a phrase almost nobody using it could actually define.

The tweet came from Peter Steinberger, an independent developer well known in AI coding circles, posted June 7, 2026: "Here's your monthly reminder that you shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents." It reached 2.2 million views. In the replies, Varadh Jain asked the only question that mattered: what does that look like in practice? Matthew Berman's answer set the mood for the whole thread: "nobody knows but him and boris." One reply, from a poster named Trash Panda Emoji, got closest to a real answer: this is not the old, simple "ralph" loop, it is "some kind of continuous orchestration loop that oversees other threads/agents," meaning a program that decides which step runs when.

The "boris" in question is Boris Cherny, who built Claude Code as a side project in September 2024. It is now reportedly behind close to four percent of all public commits on GitHub. Speaking at the WorkOS Acquired Unplugged event on June 2, 2026, he gave the plainest version of the idea: "I don't prompt Claude anymore. I have loops that are running. They're the ones that are prompting Claude and figuring out what to do. My job is to write loops." He describes three stages of his own work: a year ago, writing code by hand with autocomplete; then running five to ten Claude sessions in parallel and prompting each one himself; now, writing no prompts at all, only loops, while a couple hundred agents read his GitHub activity, Slack messages, and Twitter feed and decide on their own what to build next. He has a number behind it: in the 30 days before December 27, 2025, reported by Simon Willison, 100 percent of his contributions to Claude Code were themselves written by Claude Code, across 259 landed pull requests (proposed code changes waiting to be merged in). He deleted his code editor in November and has not reopened it since. The part both camps tend to skip: he is not saying engineers stop mattering. Someone still has to decide what to build, talk to customers, and coordinate teams. The job moved up a level, from writing the code to writing the thing that writes the code.

The argument got messy because "loop" quietly means five different things depending on who says it, and a five year ladder makes it clearer than any one definition. First: the 2022 academic pattern called ReAct, where a model reasons, calls a tool, reads the result, and repeats, with one human watching one model. Second: AutoGPT in 2023, which handed a model a goal and let it prompt itself, and became infamous for spinning forever without finishing anything, a failure that fed years of "agents are a toy" opinion. Third: the "ralph" loop, published by developer Geoffrey Huntley in July 2025, close to a one line script that feeds the same prompt file into the agent over and over. Its real innovation was discipline, not cleverness: every pass wipes the conversation and resets to a fixed set of reference files, so the agent never drifts off track. Huntley used it to build an entire programming language for about 297 dollars in running costs. Fourth, in spring 2026, both Codex and Claude Code turned that into a built in "/goal" command that repeats the ralph loop until a small checking program confirms the task is actually finished.

Fifth is what Cherny and Steinberger are actually describing, and it is genuinely new, not a rename of the ralph loop. Four things changed: the loop itself, not the individual task, becomes the unit of work. Loops start supervising other loops, running at the same time and on a schedule rather than only when a person kicks one off by hand, so the work runs on the clock instead of on your attention. And durability becomes an explicit requirement: progress is saved to git (the tool that stores and tracks changes to code) so a loop can crash and pick back up, because unlike the ralph loop, nobody can assume your terminal window stays open. The best skeptical line in the whole thread still lands, even cut off mid sentence: it is just a scheduled background job with a fancier name on it.

Here's your monthly reminder that you shouldn't be prompting coding agents anymore. You should be designing loops that prompt your agents.
via @mvanhorn on X →
04 Learnings

How to set up Claude loops that run without you

You already run scheduled background jobs, called launchd jobs, for your own personal systems, so this post is a direct comparison: could a simpler setup replace part of what you built by hand? It walks through how one builder replaced typing prompts one at a time with agents that run themselves on a schedule.

The post opens with a claim about whoever built Claude Code, Anthropic's coding tool: they say this person has not written a line of code this year, works mostly from a phone, and runs a few thousand agents overnight while they sleep. The shift credited for that is not a better model, it is loops replacing one-off prompts.

A loop, in this post's terms, is Claude running on a repeating schedule instead of answering one prompt and stopping. The setup it describes is plain: you have Claude use cron, a scheduling tool built into every computer, to create a job, then tell it how often to repeat, anywhere from every minute to once a night. No framework, no extra software, just a schedule. The trick is matching the interval to the task. A loop watching your automated tests might fire every few minutes, since problems there need a fast response. A loop that summarizes the day's work runs once at night, so the summary is waiting for you when you wake up.

The mental habit the post pushes is a switch from asking what to prompt next, to asking what job should run on its own from now on. Its rule of thumb: anything you do more than twice, anything you keep checking by hand, anything that breaks at 3am, is a loop waiting to exist.

The real gain, it argues, comes from running several loops side by side, each with one narrow job. Its examples: one loop watches open pull requests, proposed code changes waiting for review, and fixes failing tests on its own. One keeps flaky tests, tests that pass sometimes and fail other times for no code reason, patched. One pulls reader feedback from a feed and groups it into themes every 30 minutes. None of the three need a person to start them. The catch with running loops on your own laptop is obvious: close the lid and they all stop. The fix, which the post calls a routine, is the same idea moved to a server. You set the job up once and it fires on a schedule, an automatic trigger, or a simple request to another program, whether your laptop is open or not.

The post is blunt about how to start: not with ten loops and a dashboard on day one. That collapses by the weekend because you lose track of which loop did what. Start with the single task you check out of habit every day and turn only that one into a loop. A good first loop has three traits: it runs on a clear schedule, it has a job narrow enough that it cannot misread the instructions, and its output is something you can check in a few seconds, like a test-suite watcher, a pull-request updater, or a daily summary. The loops that fail are the vague ones. "Improve the codebase" is not a loop, it is a wish. "Find functions over 50 lines and open an issue for each" is a loop, because it is specific enough to trust without watching it.

Risk gets handled by keeping a person in the decision, not in every single step. A pull-request loop can rebase code and fix failing tests on its own, but merging into the main codebase still waits for your yes. A loop built to touch a hundred files across a codebase can open all hundred pull requests, but a person still approves the first one before the rest go out. The stated goal is zero human involvement in the boring 95 percent of the work, and full attention reserved for the risky 5 percent. The post's claim for what changes after a week: the pull-request loop saved around 40 interruptions where the builder would otherwise have switched tasks, the nightly summary was waiting every morning without a single typed prompt, and a feed that used to get ignored turned into a short, readable list of themes.

You're paying for a fleet of agents and using one chat window.
via @hanakoxbt on X →

Your AI agent gets sloppy because its context fills up

You run several agent workflows that take many steps across your projects, and this explains a pattern you have probably seen: they start sharp and get sloppy by step 15 or 20. The post's argument is that the fix is almost never picking a better model. It is managing what the model can actually see at each step, a discipline it calls context engineering.

A chatbot answers one question and stops, so writing a good prompt is most of the job. An agent takes actions instead: browsing, calling tools, writing code, running commands, one after another, sometimes for dozens of steps. Every step's output gets added back into what the model is holding in mind. That holding space is the context window, how much text the model can hold in mind at once, and it has a hard limit measured in tokens, chunks of text of roughly three quarters of a word each. The post's analogy: the model is the processor, the context window is the computer's working memory, and just as a computer slows when memory fills up, an agent's reasoning gets worse as its context window fills.

The decline has a name in the post, context rot, and it is not a cliff at the hard limit, it starts early and builds gradually. Chroma, a company that ran the study, tested 18 leading AI models, including GPT-4.1, Claude 4, Gemini 2.5, and Qwen3, and found every one of them got worse as the amount of input text grew, well before hitting its stated maximum. A model advertised as holding 200,000 tokens can show a real drop in quality by 50,000. Part of the cause is a pattern the post calls "Lost in the Middle": models remember the start and the end of their context well, but tend to ignore the middle. Researchers measured over 30 percentage points of accuracy lost when a fact was moved from the start of the context to the middle of it. In practice, that means your original instructions, buried under tens of thousands of tokens of tool output, can effectively vanish. Claude Code users, the post says, have noticed output quality drop once context use hits 40 to 60 percent of capacity, well short of the actual limit.

The post lists seven things competing for that same limited space: the system prompt (the agent's standing instructions), the list of tools it could call and their descriptions, the results of every tool call already made (a single web page fetch can add 5,000 to 10,000 tokens), anything pulled from your own documents to inform an answer (retrieval, sometimes called RAG), the full back-and-forth conversation so far, short and long-term memory of past sessions, and the agent's own working notes on where it is in the task. All seven fight for the same window, and context engineering, in the post's framing, is deciding which ones win.

It groups every fix into four buckets: write, select, compress, and isolate. The one it details most is write, giving the agent ways to save information outside the context window so it is not lost when the window fills or gets trimmed. That takes three forms: a scratchpad, a tool that lets the agent jot down findings and decisions as it works; a rules file, the same role your CLAUDE.md files play, instructions read fresh at the start of every session so the agent never forgets the basics; and memory extraction, where the agent saves facts and preferences that outlive the current session entirely. The post cites one concrete result for the scratchpad idea: Anthropic built a dedicated "think" tool for Claude to reason through problems, and it improved performance by up to 54 percent on tau-bench, a standard test for how well an agent handles multi-step tasks. The next fix in the framework, select, starts from the same problem stated but not resolved in what was saved: an agent given 40 tools, a large knowledge base, and several sessions of history cannot load all of it at once, so choosing what to hand it moment to moment is the next piece of the playbook.

Context is the set of tokens included when you sample from an LLM. Context engineering is optimizing the utility of those tokens to consistently achieve a desired outcome.
via @sairahul1 on X →

Free guide maps how AI agents change software development

A free 50-page guide from Shubham Saboo lays out how the software development process changes when AI agents write the code themselves, rather than just autocompleting it as you type. It frames this as a shift from "vibe coding" to agentic engineering, with a new step-by-step process for the whole build cycle.

via @Saboo_Shubham_ on X →
05 Key News

Judge dismisses Google's DMCA suit against a scraper

Outside your usual reading: a federal judge just threw out a lawsuit Google filed to stop a company called SerpAPI from scraping Google's own search results. The reasoning matters. It is one of the first real tests of how far a company can stretch copyright law to lock up data now that everyone wants it for AI.

SerpAPI builds what is effectively an unauthorized way to pull Google's search results automatically, in large volume, so other companies can use that data. Reddit sued SerpAPI and the AI search company Perplexity last fall over the same kind of scraping, claiming it let people get around Reddit's data deal with Google. Google filed its own, similar lawsuit against SerpAPI a few months later. Both cases used a copyright law called DMCA 1201, the anti-circumvention clause: originally written to punish people for breaking digital locks (DRM) on movies and music, later stretched by companies to cover things like blocking third-party printer ink or garage door openers. Google argued that scraping past its anti-bot defenses counted as breaking a digital lock, and so was covered by that same law. The judge disagreed and dismissed Google's case, though Google is allowed to refile.

The digital lock in question is called SearchGuard. It works like a CAPTCHA (the 'prove you are human' check): when Google's system gets a search request from an unrecognized source, it sends back a small piece of code that a real user's browser runs automatically without the person noticing. Automated scraping tools usually cannot pass that check, so SearchGuard blocks them. SerpAPI argued that SearchGuard has nothing to do with copyright. It just tells human traffic apart from bot traffic, and it blocks bots from any part of Google's results, whether or not those results contain anything copyrighted.

The judge agreed with that argument. Google's search results are a compilation of information pulled from across the public internet, organized by relevance, and Google never actually claimed that google.com or the results shown on it are protected by copyright. If the thing being protected is not covered by copyright, a copyright law cannot be the tool used to protect it, no matter what the technology behind the lock is doing. The ruling does not settle the underlying question of whether AI companies and data brokers should be allowed to scrape at will. Reddit's parallel lawsuit against SerpAPI and Perplexity is still working through motions to dismiss, unresolved.

The bigger picture, per the source: as AI companies compete for training data, more sites are trying to put up toll booths on the open web to control who gets to scrape them, aimed mostly at AI companies but catching everyone else along the way. Reddit has an official data deal with Google, and none of the companies SerpAPI supplies were party to that deal, which is part of why the earlier Reddit suit looked shaky to begin with: Reddit does not hold the copyright on its users' posts, the users do. Google's version at least had the fact that SerpAPI was scraping Google's own site, not someone else's, but the court's answer landed the same either way: copyright law protects specific copyrighted works, not a company's general wish to control who touches its public pages.

via Techdirt →

Fintech funding topped $29 billion in early 2026

A newsletter roundup says fintech funding passed 29 billion dollars in the first half of 2026, but most founders did not get a share of it. It also flags Ramp's free tool that automatically picks which AI model answers a request, arguing the tool itself is not really the point: the money flow around it is.

via Linas's Newsletter →

Texas A&M joins the Genesis Mission for AI in science

Texas A&M University has joined the Genesis Mission, an effort to use AI to speed up scientific research.

via Texas A&M Stories →

Authors split on Anthropic's $1.5 billion book ruling

A court ordered Anthropic to pay 1.5 billion dollars in a copyright infringement case, and authors have mixed feelings about the outcome, NPR reports.

via NPR →

New open AI model Kimi K3 launches at 2.8 trillion parameters

Kimi K3 is a new open model (its files are public, so anyone can run it) with 2.8 trillion parameters and a context window (how much text it can hold at once) of 1 million tokens (chunks of text, not whole words). It trails Claude and GPT on the standard tests but is the largest open model released so far, if you happen to have a spare terabyte of storage to run it.

via reddit r/ArtificialInteligence →

The AI talk is shifting from tokens to model quality

Spotted on Apple News: a piece argues the AI conversation is moving on from counting tokens (chunks of text a model handles) toward judging a model's raw quality instead. It is shorthand worth knowing before your next conversation with other builders about where AI progress is actually coming from.

via Business Insider →

Starbucks killed its AI inventory tool after nine months

Outside your usual reading: Starbucks rolled an AI tool out to over 11,300 stores that used an iPad camera to count inventory, but reflections off a steel fridge doubled real oat milk cartons from 5 to 10. The tool reportedly cost more than 10 million dollars to build and was pulled nine months later, a reminder that AI breaks on mundane details nobody planned for.

via reddit r/ArtificialInteligence →
06 Tools & Craft

Four loops to stack when building an AI agent

LangChain's Sydney Runkle breaks agent design into four stacked loops. The base loop just calls tools until the task is done. Layer two adds a grader, something that checks the output, and sends it back on failure. Layer three wires the agent to real triggers, like a Slack message, so it runs without you watching. Layer four reviews the agent's own run logs and rewrites its prompts to improve over time, a useful checklist next time your own agent setup needs another layer.

via @sydneyrunkle on X →

20 AI repos worth knowing, including Superpowers

A LinkedIn list ranks 20 AI code repositories by GitHub stars, split across coding agents, developer tools, and infrastructure. OpenClaw leads with 278,000 stars, ahead of Opencode's 118,000 and Claude Code's 75,000. Superpowers, the framework you already run your own sessions through, sits close behind at 73,000. Worth a scan for a new tool to try, and a nice confirmation that the one you picked is rated highly by other builders too.

via LinkedIn →

One founder turned 7,000 notes into a research engine

A builder describes wiring a 7,000 note Obsidian vault to Kimi K2.6, an AI model, through an MCP server (a standard way to plug tools into an AI model), so the agent can read, search, and write notes on its own. He claims the setup replaces consulting reports that would otherwise cost $5,000 to $50,000, in a single run. The indexing runs on a Smart Connections plugin with 786,000 downloads. Worth holding up against your own vault and agent setup as a comparison point.

via @noisyb0y1 on X →
The Last Word
Starbucks' AI saw ten oat milk cartons where five sat, and got retired.
The Desk Report

How this edition came together — from bookmarks and feeds to the page.

217links gathered
40read by the desk
17made the edition

Where they came from

On the cutting-room floor — 23 links read but not run this week

Quality over volume: most links get a second look and a pass. The ones that made it earned their place.

Reading Room — every Sunday

The week's AI signal in 23 minutes — what happened, why it matters, and what to do with it. Curated by someone who actually builds, not a feed algorithm.