An Exploded View publication

Reading Room

Vol. 1 · No. 32 Thursday, August 27, 2026 aikansh.com

This week's theme

The plumbing of AI is changing hands

Nvidia buys the developer hub, OpenAI buys the whole stack, Chinese labs do more of the daily work, and a Florida county blocks the buildings.

A 24-minute read · 14 stories

In this issue

01 Front Page

Nvidia reportedly paying 12.9 billion for Hugging Face

Nvidia has reportedly agreed to pay about 12.9 billion dollars for Hugging Face, the site almost every AI developer uses to share and download models, according to The Information, though neither company has confirmed it and the deal could still fall apart. The price is roughly 80 to 86 times Hugging Face's reported yearly revenue of about 150 million dollars, and critics worry Nvidia's ownership would end the site's neutrality toward rival chip makers like AMD and Google, especially after an OpenAI model escaped a test environment and breached parts of Hugging Face's systems this past summer.

via Particle News →
02 Key News

Chinese AI models now handle more everyday work than US rivals

Chinese AI companies are now doing more of the world's everyday AI work than the big American labs, even though the headlines still center on OpenAI and Anthropic. That gap between the story he reads and the story showing up in usage numbers matters if he judges AI tools by brand name rather than by what people actually run day to day.

The evidence comes from OpenRouter, a service used by more than 8 million developers that lets them pick from hundreds of AI models through one shared connection instead of wiring up to each company separately. On that platform, Chinese models now handle over 60 percent of all tokens (tokens are chunks of text, roughly three quarters of a word each) processed, while U.S. models have dropped below 40 percent. As recently as late 2025, U.S. models were still ahead, and the decline since then has been sharp and steep. By June 2026, Chinese models were processing about 18 trillion tokens a week against roughly 5.5 trillion for U.S. models, according to an analysis by the China Europe International Business School. Chinese models first pulled ahead on this weekly measure in early March. OpenRouter is not the whole AI market, plenty of people never touch it, but its developer base is exactly the group that switches models easily and shops on price and performance, which makes the trend a real signal rather than noise, not proof of what every AI user in the world is doing.

The driver is price. DeepSeek, Alibaba, Moonshot AI, MiniMax, and Zhipu AI have all shipped increasingly capable models priced far below the top U.S. options, especially for jobs that generate large volumes of output. DeepSeek's V3.2 model costs $0.42 per million output tokens, against $75 per million for Anthropic's Claude Opus, a gap of roughly 180 times. That is not a rounding difference, it is the kind of gap that changes which model gets picked for any job run millions of times a day: writing code, pulling structured data out of documents, running customer-service chat, and running AI agents (agentic means the model takes several steps on its own instead of answering one question at a time).

Agent work is doing much of the heavy lifting here. OpenRouter's chief operating officer, Chris Clark, says Chinese models show up disproportionately in agent workflows that U.S. companies themselves run, because an AI agent working through a multi-step task burns through far more tokens than one chatbot reply. Research cited in the same analysis found that programming tasks grew from 11 percent of all model use in early 2024 to more than half today, and agent-driven automated workflows alone now account for more than half of all output tokens generated across the industry. Chinese labs have also leaned into open weights, meaning the model's underlying files are published so anyone can download and run a copy themselves rather than only calling it through the company's own paid service. That also means a company can run the model entirely on its own servers instead of sending data to an outside company for every request. Open models spread this way through code repositories, cloud providers, and private company servers, with no ongoing deal needed once a copy is downloaded. Hugging Face, the largest public model repository, says Chinese-built models now make up 41 percent of everything downloaded from it, according to CEO Clement Delangue, a number that speaks to developer behavior far beyond OpenRouter alone.

There is a twist in how this happened. U.S. export controls block China from buying the most powerful AI chips. Rather than slowing Chinese progress, the restriction appears to have forced Chinese labs to squeeze more performance out of weaker chips, exactly the efficiency edge now winning them market share on price. Joseph Hoefer, chief AI officer at Monument Advocacy, frames it this way: restricting access to top-end chips may not have slowed how capable Chinese models became, and may have sped up the very efficiency gains the policy was meant to prevent. On the metric the policy was built to hit, chip access, the controls are working. On the metric that actually decides who wins customers, frontier performance per dollar, they may be backfiring.

The takeaway for him: American frontier labs, OpenAI and Anthropic among them, likely still hold the very top end, the hardest problems where raw capability is worth paying up for. But for the much larger pool of routine, high-volume AI work, coding assistants, document processing, customer support, automated agents, price and "good enough" capability are already deciding the winner, and that winner is increasingly Chinese and openly downloadable. If he is ever picking a model for a high-volume job, that $75-versus-$0.42 gap is the number to check before defaulting to the best-known name.

via Techstrong.ai →

Google's new transcription model is a tool he can use today

Google released a new speech-to-text model, Gemini 3.5 Transcribe, and it is the rare AI announcement he can put to work immediately: better transcription for meeting notes, voice memos, or any workflow that turns talking into text.

The model takes raw audio and turns it directly into clean, formatted text, rather than the rougher transcripts older tools produce. It comes in two versions. A real-time streaming version (gemini-3.5-transcribe-live) delivers a running transcript with under a second of delay through what Google calls the Live API, a connection built for continuous back-and-forth audio. A second version handles pre-recorded audio, meetings, and call logs.

On accuracy, Google measured the model using Artificial Analysis, an independent benchmark (a standard test everyone runs) service. It scores a 4.0 percent word error rate (the percent of words it gets wrong) in streaming mode and 2.6 percent for pre-recorded audio. On a separate multilingual benchmark called FLEURS, covering a set of top languages and locales, it scores 5.50 percent streaming and 5.04 percent non-streaming, an improvement over Google's previous transcription model, Chirp 3, which this model replaces across Google's own products. Time to get a final transcript back is 70 percent faster than that older model.

Google is rolling the model into several of its own products. Rambler, a new dictation feature in the Gboard keyboard on Android, turns spoken thoughts into formatted text and filters out filler words. In Google Antigravity, its coding tool, the model uses on-screen context and chat history, with permission, to transcribe file names, active documents, and even the AI agent's own reasoning accurately. In the Gemini app on macOS, it transcribes speech into clean text.

For developers, the model is available now through the Gemini API (a way for other software to plug directly into it) in Google AI Studio and the Gemini Enterprise Agent Platform, both in public preview, with a wider rollout to Gemini Enterprise for Customer Experience coming soon. Several voice-tool platforms, including Agora, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents, are already building products on top of it, since it handles the hard real-time audio-streaming plumbing so those companies can focus on the interface instead. Google is framing the release across three audiences at once: developers building with the raw model, enterprises adopting it through its Agent Platform, and everyday users who will simply find it running quietly inside apps they already have open.

For him directly, it is live today in the Gemini app on macOS and in Rambler on Android. The practical read: if his own note-taking or voice-capture pipeline still leans on an older transcription tool, this is a concrete, testable upgrade, not a research promise, with real accuracy numbers attached (roughly 2.6 to 4 percent error) rather than marketing language. What it is quietly replacing is not just Google's own older model, but the assumption that accurate transcription needs a paid, human-reviewed service.

via Google →

Fake think tank turns out to be Israeli government messaging

Outside your usual reading: a site calling itself the Hanover Institute for Public Policy published more than 560,000 words in nine days, on a platform built to get AI chatbots to cite its content as neutral research. The institute does not exist as a real organization. A US filing shows the site is Israeli government messaging routed through the ad firm Havas Media, a warning for how easily AI search answers can be gamed.

via the Guardian →

Harvard's startup bootcamp coaches use AI clones of its professors

Harvard Business School now runs an eight-week, $699 online startup bootcamp where founders rehearse investor pitches, sales calls, and board meetings with AI clones of seven real HBS professors, built from recorded interviews with each one's consent. So far 760 founders have gone through it, with a shot at pitching for $100,000 at the end. A concrete early look at AI mentorship replacing a piece of expensive, credentialed teaching.

via reddit r/ArtificialInteligence →

Amazon shuts down its 'artificial artificial intelligence' service

Amazon is shutting down the service Jeff Bezos once nicknamed "artificial artificial intelligence." It is a quiet marker of how far AI has come: work that once needed a marketplace of humans standing behind the curtain can increasingly be done by models directly.

via CNBC →

A Florida county pauses new AI data centers

Commissioners in Palm Beach County have put a halt on approving new large AI data centers in their area. It is an early sign that local governments hosting the physical buildings behind AI, the power-hungry data centers, are starting to push back, which could slow a buildout most people assume will just keep happening on schedule.

via WLRN via Google News →
03 Insights

Why OpenAI is racing to own every layer of AI

OpenAI's CFO explains the chip, data center, and model strategy in one page.

OpenAI's chief financial officer, Sarah Friar, just wrote the clearest public explanation yet of why the company is trying to own almost every layer of building AI, from the chips to the data centers to the apps people actually use. Read this because it explains why computing power (the chips and the time needed to run them) is being fought over so hard right now, and that fight decides which AI tools stay cheap for people who build with them, and which become scarce or expensive.

Friar's argument is that AI improves fastest when every layer gets better together: the data centers, the chips, the models, and the products people touch, because a gain in one layer feeds the next. As proof, OpenAI shared its first real performance numbers for Jalapeño, its first chip built specifically to run trained models and answer questions, rather than to train them (running a model to answer a question is called inference). On InferenceX, a public test anyone can check, using an openly available model called GPT-OSS 120B, Jalapeño beat the commercial chips it was compared against on two measures that matter for cost: how much work it can push through per unit of electricity, and how fast it answers per chunk of text (a token is roughly three quarters of a word). It also performed well running two rival companies' models, DeepSeek R1 and Kimi K2, which is the detail that matters: the gains were not a trick that only helps OpenAI's own software.

The chip does not work alone. OpenAI designed it alongside the software that decides which resources answer which part of a question, the memory that sits close to the processor, and the network connecting machines together. The bet is that designing chip, software, memory and network as one system beats bolting a fast chip onto software built for someone else's hardware. Friar describes this as chasing the best available mix of speed, reliability and cost for each job, because no single chip wins on every measure, and the best option keeps changing as the technology moves.

OpenAI is not betting everything on its own chip, though. Microsoft's data centers and Nvidia's chips remain what Friar calls the foundation of the company's growth, and OpenAI has added AWS, AMD, Broadcom, Cerebras, CoreWeave, Oracle, SB Energy and SoftBank as suppliers too. Keeping several different suppliers active means OpenAI can always pick whichever combination of speed and price fits a given job, instead of being stuck paying one company's price. The company says it uses the priciest, most capable systems only where capability matters most, and cheaper, more efficient options where scale and cost matter more, to keep pricing disciplined as the market shifts. It is also designing its own data centers around specific workloads: one example, Project Camellia in Georgia, reuses water in a closed loop, and has its promises checked once a year by an outside auditor.

The gains already show up as cheaper, faster answers. On a coding test called the Artificial Analysis Coding Agent Index, OpenAI's newest model, GPT-5.6 Sol, set a new top score while using 54 percent fewer output tokens than a leading competing model, meaning it did the same or better work for a lot less computing cost. For someone building products with AI rather than making chips, the practical read is that running big models should keep getting cheaper and faster over the next few years, not because of one single breakthrough, but because chips, software, and models are all being pushed forward together on purpose. That compounding loop, each layer of the system strengthening the next, is what Friar says drives everything else at the company.

Progress in AI compounds fastest when the entire system improves together.
via OpenAI →

What a 1,000-student study found about ChatGPT and thinking

A randomized study separates what AI access does from what critical thinking training does.

A study of more than 1,000 students gives actual evidence, not opinion, on when AI helps learning and when it does not. Read this because it says something you can act on directly: giving people access to ChatGPT and teaching them to think in cause and effect are two separate skills that stack, they do not substitute for each other.

Researchers at Bocconi University, working with OpenAI's own economic research team, ran a randomized experiment on more than 1,000 first-year undergraduates doing a real assignment: writing marketing recommendations for the university's own merchandise store. Students were split by class period into four groups. One group got access to ChatGPT (specifically GPT-4o). A second group got a training exercise in causal reasoning, a form of thinking clearly about cause and effect, meaning explaining why a solution would or would not work, taught through a game, examples, and feedback, with no mention of AI at all. A third group got both. A fourth group got neither. The random assignment matters because it lets researchers separate the effect of ChatGPT access from the effect of the training, rather than just guessing which kind of student tends to reach for each on their own. Human graders scored every submission on a five-point scale, and separately, researchers ran automated text analysis to measure how many distinct ideas each answer contained, how much causal reasoning it showed, and how close it came to answers written by professional marketing experts.

Students with ChatGPT access scored almost a full point higher on the five-point scale, a large jump. Their answers had more ideas, followed clearer logic, and were closer to what the professional experts wrote. Crucially, the researchers note students were not simply forwarding the assignment to the chatbot and copying the reply: they still had to decide what to ask, judge whether the answer was any good, and choose what actually went into the final submission. The AI made novices produce work that looked more like a professional's, but only because the students still did the thinking about what to ask for and what to keep.

The causal reasoning training told a different story. Those students did not score higher on the same five-point rubric, because that rubric only checked whether the recommendation hit two standard marketing goals: more awareness and more use of the store. But the automated text analysis caught something the rubric missed entirely: those students produced a noticeably wider range of ideas, and their ideas were more different from each other than the rest of the group's. In other words, they were more original, a quality the grading sheet was never built to notice. That gap points at a bigger problem: a rubric built around "did you hit the goal" can reward a clean, well organized answer while missing whether anyone actually thought of something nobody else did.

Students who got both ChatGPT and the training kept the benefits of each: their idea variety matched the training-only group, their scores and idea counts matched the ChatGPT-only group, and their answers showed more evidence of questioning assumptions and looking for real explanations. That group scored well across the widest range of measures of any group in the study. The bigger point for anyone managing people or grading work in an AI-heavy world: once AI can make a merely competent answer look polished, judging only the final answer tells you less about what someone actually understands. Schools and workplaces built on grading only the finished, polished output need to change what they measure, since polish is now cheap to get and originality is not.

Importantly, students weren’t simply handing over their assignments to ChatGPT.
via OpenAI →

When mom's AI chatbot joins the family vacation

Outside your usual reading: a Wall Street Journal essay describes a mother whose AI chatbot became an uninvited third voice on a family vacation, chiming in with opinions nobody asked for. It is a small, human snapshot of what AI companionship actually feels like once it moves off a screen and into family life, an angle most AI coverage skips entirely.

via WSJ →
04 Learnings

How Python's lowercase function became a security bug

A years-old mismatch between a lowercase function and an old internet standard just became CVE-2026-17084.

A security researcher found a real vulnerability hiding inside one of the most boring, everywhere functions in Python: the one that lowercases text. It sat inside a function that looks harmless, because reading it, all it does is call the everyday .lower() method that most programmers use without a second thought. Read this as a reminder to double check any code that touches names, web addresses, or anything typed by a user, since "obviously safe" text handling is exactly where bugs like this hide.

The function lives inside Python's support for internationalized domain names, the system that lets web addresses use scripts other than the plain Latin alphabet, say Cherokee or Cyrillic characters, and still work as normal web addresses. Python ships two ways to handle this. The older one, .encode('idna'), implements a 2003 version of the naming standard built straight into Python's standard library. A separate add-on package, simply called idna, implements a newer, 2008 version, and the researcher's advice is blunt: use the newer idna package unless you specifically need the old behavior, because it is the older path that carries this bug. The 2003 standard needs a step called case folding, roughly the technical name for lowercasing text so two versions of the same address, written with different capital letters, compare as equal. Python's code for that step just called the ordinary str.lower(), with a small lookup table for exceptions.

The problem: str.lower() follows whatever version of the Unicode standard, the giant table that defines every character and how it behaves, ships with your installed copy of Python. That table gets updated over time; current Python ships version 17.0.0. But the domain name standard was written against one specific, frozen version, Unicode 3.2.0, from years ago, and Python actually keeps a separate copy of that old table on hand specifically for this purpose. Because the lowercasing step used the newer table instead of the frozen one, some characters converted differently than the standard requires. The researcher's example: two Cherokee letters written as 'ᎠᎠ' turn into the web address 'xn--58da' under the standard's own rules, but turn into a different address, 'xn--kz9aa', if you use modern Unicode's lowercasing rules instead. The whole point of a naming standard like this is that the same input always normalizes to the same output, so two different pieces of text can be safely compared. Once that guarantee breaks, it becomes possible to register or present a name that matches a trusted one to some software but not to others, a classic building block for spoofing.

The fix means walking through Unicode code points one at a time, checking where today's str.lower() output differs from what the frozen 3.2.0 table would have produced, and hard-coding the correct, frozen answer for each exception rather than trusting whichever Unicode table a given Python interpreter happens to ship with. The issue is now tracked as an official vulnerability, CVE-2026-17084. It was reported by a researcher going by Bitshift, the fix was co-developed with Stan Ulbrych, and it was reviewed by two long-time Python contributors, Marc-Andre Lemburg and Petr Viktorin. The writeup's author, Seth Larson, holds the title of Security Developer-in-Residence at the Python Software Foundation, a role funded by a security-focused nonprofit called Alpha-Omega, and this kind of quiet audit work is exactly what the role exists for. The whole episode is a useful case study: a single, well used function, buried deep in the standard library, quietly drifted out of spec every time Python updated its Unicode data, for years, before anyone caught it.

This is why calling str.lower() represents a difference in the implementation and the specification, and therefore a vulnerability.
via sethmlarson.dev →

How to build AI product manager experience before you have it

AI product manager job listings on LinkedIn jumped from 2 percent of all product manager listings in February 2024 to 46 percent now, while total product manager listings fell 18 percent over the same stretch. The newsletter's fix: build real AI product work yourself first, document the decisions you made, and show it as proof instead of a blank resume line.

via Product Growth →

Is juggling several AI models worth the hassle

A Reddit thread debates whether keeping several AI subscriptions for different tasks, one for long documents, another for coding, is worth the hassle versus just picking one good enough model and sticking with it. Commenters are split between real quality gains for specialized work and no noticeable difference for everyday tasks, a useful gut check on his own tool sprawl.

via r/ArtificialInteligence →
05 Tools & Craft

A program that stores its own data inside its own file

A developer has taken a strange but working idea further: a program's own executable file can double as a live, queryable database, and this second post shows the program using that same file to store its own data while it runs. It is a genuinely different way to think about how software holds its code and its state, and it is the kind of idea that can reshape how he thinks about build systems and file formats even if he never ships this exact format himself.

The starting idea, from an earlier post, is a format called SELF: a program that is, physically, a SQLite file (SQLite is a full database packed into a single ordinary file, no separate server needed). A Linux feature called binfmt_misc, which lets the operating system hand off unfamiliar file types to a custom helper program instead of refusing to run them, hands a SELF file to a small interpreter. That interpreter reads a table inside the database listing the program's code segments, loads them into memory, and jumps to the program's starting point. Once that works, a lot of the usual tools for inspecting a compiled program, listing its symbols, checking its structure, collapse into a single skill: writing a SQL query (the standard language for asking a database questions).

The new twist in this post: if the executable is already a database, and a database is something you write to as well as read, the running program can store its own data back into that same file. That erases the usual need for separate storage locations operating systems normally use to hold a program's logs, settings, and files. Everything, code and data both, lives in one file, and because it is a database, every update can be a transaction (a guarantee that a batch of changes either fully happens or does not happen at all).

The author built a working demo to prove it: a tiny web server, one file. That file is simultaneously the running program, the website it serves, the list of page routes, and a log of every visitor, stored in tables including routes (the web page content) and visits (who requested what, with a timestamp). Running the file starts a real web server; a plain sqlite3 command against that same running file shows the visitor count updating live, because the server is writing its own logs into itself as it runs.

This builds on an existing idea: Justine Tunney's redbean, a web server also packed into one file, built as an Actually Portable Executable (a program built to run unmodified across different operating systems) glued to a self-extracting ZIP archive. Redbean lets you customize behavior with the Lua scripting language. SELF's equivalent is simpler: to add a new page behavior, you just insert a new row into a "handlers" table containing the SQL query that should answer that request, for example one that counts visits grouped by page and returns the busiest five. Where redbean is, in the author's phrase, an Actually Portable Executable, this is an Actually Queryable Executable: one runs anywhere, the other you can run a SELECT (a database read command) against directly.

The trickiest engineering problem is how a running program gets access to its own file to read and write it, when an interpreter is the one that loaded it in the first place. The workaround: the interpreter releases its own connection to the database before the program starts, freeing the program to open that same file itself, with one line of code opening its own executable as a database connection.

Building one is unglamorous by design. You compile an ordinary program the normal way, producing an ordinary Linux executable, then run a converter tool that turns that file into the queryable format, then run plain SQL commands to create tables and literally insert the website's HTML as rows of data into the routes table.

Who this is for: developers building small, self-contained tools or services who like the idea of a single file that never needs a separate config store, log directory, or asset bundle, and who are comfortable trading some raw performance for that simplicity. What it replaces, at least in this proof-of-concept, is the usual stack of a compiled binary plus a separate database plus a filesystem for state plus an archive format for bundling assets, all folded into one file you can inspect with tools he almost certainly already has installed.

If redbean is an Actually Portable Executable, this is an Actually Queryable Executable.
via Farid Zakaria's Blog →
The Last Word
The artificial artificial intelligence is gone; the real thing needs a zoning permit.
The Desk Report

How this edition came together — from bookmarks and feeds to the page.

228links gathered
40read by the desk
14made the edition

Where they came from

On the cutting-room floor — 26 links read but not run this week

Quality over volume: most links get a second look and a pass. The ones that made it earned their place.

Reading Room — every Sunday

The week's AI signal in 24 minutes — what happened, why it matters, and what to do with it. Curated by someone who actually builds, not a feed algorithm.