An Exploded View publication

Reading Room

Vol. 1 · No. 34 Saturday, August 29, 2026 aikansh.com

This week's theme

Your model supplier is a choice, not a given

OpenAI drops a customer, a Chinese rival courts the big clouds, and smaller models keep winning: today is about who supplies your AI and whether their numbers hold up.

A 16-minute read · 12 stories

In this issue

01 Front Page

OpenAI cuts SpaceX off from its models inside Cursor

OpenAI is cutting SpaceX off from its models inside Cursor by November 2026, over trust, not price.

OpenAI is walking away from one of its own customers over trust, not price. Today it told SpaceX, which now owns the AI coding tool Cursor, that it will wind down the contract supplying OpenAI's models inside Cursor, with a proposed cutoff date of November 12, 2026. That is the maximum notice period the contract allows, and OpenAI says it is using the full window on purpose, to give developers as much time as possible to adjust. This is a rare, public case of a major AI lab refusing to keep serving a powerful, well funded customer, and it shows that even at the very top of the AI industry, trust still outranks who can pay.

The reason OpenAI gives is blunt: it says it cannot be confident SpaceX will use its technology within its terms of service (the rules for using OpenAI's models), based on its own experience with Elon Musk's companies breaking contracts before. After Musk bought Twitter, which is now folded into SpaceX, that company broke the terms of its OpenAI contract, one of several companies OpenAI says has done that. More specifically, Musk admitted under oath earlier this year, meaning in sworn legal testimony, that xAI, his AI company and also now part of SpaceX, had violated OpenAI's terms of service, terms that are similar to the ones xAI itself uses for its own customers.

The mechanism here matters. OpenAI's custom agreement with Cursor includes a clause giving OpenAI a limited window to cancel the deal after a change of control, meaning a change in who owns the company. SpaceX's acquisition of Cursor triggered that clause. OpenAI also points to its own upcoming model, called Astra, and says it now carries a new level of responsibility to make sure that model is used strictly within its rules before it goes out. Rather than cutting Cursor off immediately, OpenAI chose to hold the cancellation to the latest date the contract allows, while confirming it will not hand Cursor any future models beyond that point.

OpenAI is careful to separate the company it is punishing from the product it is walking away from. It says it has worked with Cursor for nearly four years and has enormous respect for the team, the product, and what it has built for developers, and that the people actually hurt by this decision are the developers who rely on OpenAI's models inside Cursor day to day. It says it is ready to go further than usual to support them through the transition.

Zoom out and this is as much a supply chain story as a Musk story. A tool like Cursor exists on top of models that OpenAI controls, and OpenAI can pull that access over a dispute that has nothing to do with the tool's own team, and everything to do with who now owns the company behind it. It is a reminder that building on someone else's AI models means your access can be revoked over decisions made two ownership layers above you, not just over your own product's behavior.

We are making this choice because we cannot be confident that SpaceX will use our technology within our terms of service, based on our experience with Elon Musk's companies violating contracts.
via OpenAI →
02 Key News

Moonshot in talks to sell its AI model through big clouds

This is worth watching because it could change which AI model your own company defaults to, without anyone deciding it on purpose. Moonshot AI, a Beijing-based startup, is in talks with Microsoft, Amazon, and Google to have its Kimi K3 model hosted and sold directly through their cloud stores, the same shelves where OpenAI and Anthropic's models already sit.

The model itself is enormous: Kimi K3 has 2.8 trillion parameters (the internal settings a model learns during training). That is far too large for almost any company to run on its own computers, so most businesses would rather rent it through a cloud than try to run it themselves. Moonshot is reportedly asking each cloud for a revenue share of up to 30 percent, meaning it would keep 30 cents of every dollar the cloud charges a customer to run Kimi K3. That is a serious cut even by AI industry standards, and it shows Moonshot trying to make money from a model whose underlying files are free to download (an open-weight model: anyone can download the files and run it themselves).

If a deal like this closes, it removes the biggest practical barrier to using a Chinese model at a real company: billing, compliance, support, and infrastructure would all run through a cloud your company might already trust and already pay. Chinese models have gained ground this year mostly on price, especially in agentic work (tasks where the model takes several steps on its own), because a cheaper model can cut costs a lot even after the cloud adds its own markup.

NVIDIA is leaning into this shift rather than fighting it. The chipmaker has been adding technical support for Chinese open models even while warning that U.S. government restrictions could limit that work. This week NVIDIA announced optimizations for DeepSeek's V4 Flash and Alibaba's Qwen 3.8, plus day-zero support (support ready the day a model launches, not weeks later) for a smaller Qwen 3.8-27B model on its RTX chips, and updates that make clusters of its DGX Spark computers easier to use for two more Chinese models, Z.ai's GLM 5.2 and DeepSeek V4 Flash. Forrester analyst Charlie Dai told CNBC that NVIDIA's moves reflect the growing influence of Chinese models and NVIDIA's goal of staying the preferred hardware layer no matter which country's model wins. In plain terms, NVIDIA does not care who makes the best model as long as it runs best on NVIDIA chips.

The politics are getting sharper. The Trump administration is reportedly considering rules that would restrict NVIDIA's ability to support Chinese-made models and apps, and NVIDIA has warned that restrictions on models like DeepSeek, Qwen, or Kimi could hurt its own business. Moonshot itself is under direct scrutiny: Treasury Secretary Scott Bessent has said he might add the company to a trade blacklist, and U.S. officials have accused Moonshot of using distillation (training a smaller model to copy a bigger one, in this case allegedly Anthropic's Fable model) and of illegally obtaining NVIDIA chips. Complicating any crackdown, more than 20 companies including Microsoft, Meta, Palantir, and NVIDIA itself asked policymakers in July not to impose premature restrictions on open-weight models, since anyone, not just Chinese developers, can download and modify them.

None of these cloud deals have closed yet. But the fact that Moonshot is negotiating shelf space on the three biggest clouds at all is the real story: Chinese model makers are no longer just competing on downloads and cheap pay-per-use pricing, they are now going after the recurring, enterprise-grade distribution that OpenAI and Anthropic have relied on as their edge.

via Techstrong.ai →

DeepMind tests whether AI benchmark scores can be trusted

Every time you read that a new model beat a rival on some test, that claim is only as good as whether the model got to see the test questions beforehand. Google DeepMind just ran what it calls the world's first double-blind evaluation of a major, closed AI model, meaning neither DeepMind nor the outside group testing it could see the other side's material in advance. That is aimed straight at a problem worth being suspicious of whenever you read a benchmark score: benchmark contamination, where a model has already seen the test questions during its training and aces a test it never really had to solve.

DeepMind tested a Gemini Flash Lite model (a smaller, faster member of its Gemini family) against confidential questions, working with four outside partners: the Singapore AI Safety Institute, a privacy-technology nonprofit called OpenMined, an evaluation group called AVERI, and MLCommons, the organization behind many of the industry's standard AI tests.

The trick is a tool called Confidential Space, part of Google Cloud's confidential computing product line. It lets two parties run a test inside something like a sealed box: the outside evaluator's questions go in, DeepMind's Gemini model goes in, and the box runs the test with cryptographic proof (a mathematical guarantee, not just a promise) that neither side can see the other's material. The evaluator never sees Gemini's underlying model files, and DeepMind never sees the evaluator's questions, before or after the test.

Historically, testing a closed AI model rigorously meant one side had to trust the other with something valuable. Either the outside evaluator handed its questions to the AI company, risking that a future model gets trained on those exact questions, or the AI company handed over its model's internals to the evaluator, risking its intellectual property leaking out. Contracts and no-logging promises have covered this gap so far, but this is the first time the guarantee comes from cryptography instead of trust.

DeepMind says this matters most for the highest-stakes tests, such as cybersecurity evaluations or ones run by government safety bodies, where neither the questions nor the model can be allowed to leak. If double-blind testing spreads across the industry, it gives outside groups, and eventually him, real grounds to trust a benchmark score instead of discounting it for possible contamination. For now this is a pilot on one mid-size model, not yet a routine industry practice, so treat it as the first real attempt at a fix rather than a finished standard.

via DeepMind →

Samsung puts tiny calculators inside memory chips

Outside your usual reading: this is a hardware story about memory chips, not a product launch, but it explains why AI could eventually get much cheaper to run.

Normally, a computer's memory chip (DRAM) just stores data, and the processor has to fetch that data over a slow connection before it can do any math with it. That trip is one of the biggest hidden costs of running AI models: pulling the model's numbers out of memory usually takes longer than doing the actual arithmetic on them. At the Hot Chips 2026 conference, Samsung showed a different approach called Processing-in-Memory (PIM): put small calculators directly inside the memory chip, next to the data, so it never has to leave.

Samsung built this into a standard chip type called LPDDR5X, the kind used in phones and laptops. A normal LPDDR5X chip is divided into 16 separate storage sections called banks, and normally the computer can only pull data from about two of those banks at once, capping speed at 76.8 gigabytes per second. Samsung's version, LPDDR5X-PIM, puts a small calculator, built to do MAC operations (multiply then add, the basic operation behind almost all AI math), inside every one of the 16 banks. Because each calculator works directly with its own bank, all 16 can run at once, for a combined internal speed of 614 gigabytes per second, about eight times faster.

Each in-memory calculator is small: it holds a scratchpad of numbers to work with and a list of up to 64 instructions to follow, and it multiplies those against whatever model data already sits in that bank, using low-precision number formats (8-bit numbers, coarser than a full-size number but far cheaper to compute with and already standard for AI). Wire together eight of these chips and you reach roughly 9.6 trillion operations per second, in the same range as the small AI chip built into Intel's Meteor Lake laptop processors, though it costs you 128 gigabytes of ordinary memory you may not have needed.

The clever part is that Samsung did this without inventing a new memory standard the industry would need to redesign around. The chip still speaks the normal memory language: Samsung just reserved a handful of special memory addresses that act like hidden switches. Writing to one flips the chip into compute mode, at which point ordinary write commands load numbers into the calculators instead of storing data, and ordinary read commands trigger the math and return results instead of stored data. Flip it back, and the chip behaves like normal memory again. This is a conference presentation, not a shipping product, and Samsung's own materials leave open how much software rework it would take to actually use, but it is a real preview of chipmakers chasing the true cost of AI: moving data, not doing the math.

via Chips and Cheese →

OpenAI's free teacher tool reaches 55 more school districts

OpenAI is expanding ChatGPT for Teachers, its free tool for schools, to 55 more U.S. school districts, reaching over 100,000 additional teachers and staff with secure AI tools and training. It is the enterprise software playbook applied to education: give the tool away early so a whole generation of students and teachers grows up defaulting to it.

via OpenAI →

Investor group drops $50 billion bid for PayPal

An investor group has dropped its roughly $50 billion attempt to buy PayPal. A deal that size falling apart is a sign that even well-funded buyers are getting more cautious about paying up for a slowing fintech giant.

via Bloomberg →

OpenAI expands its footprint in Brazil

OpenAI is putting more people and resources into Brazil to work directly with local developers, businesses, and communities on AI adoption. It is another sign the big AI labs are racing to lock in international markets early, not just fight it out in the U.S.

via OpenAI →

Google gives its AI video model finer creative controls

Google's Gemini Omni 1.1 Flash video model now lets developers extend a clip using up to 10 seconds of prior footage for context, versus just the last second before, set exact start and end frames, and draft cheap 360p previews up to 60 percent faster before upscaling to 4K. AI video tools are moving fast from novelty demos toward something a production team could actually use day to day.

via Google →
03 Insights

Smaller AI models that think longer are beating bigger ones

Smaller AI models that think longer are starting to beat older, bigger, pricier models on hard tests.

The lesson underneath this piece is simple: raw model size is no longer what decides who wins in AI, so before you default to the most expensive model for a job, it is worth checking whether a smaller one that thinks longer can already match it for less money.

OpenAI's smaller reasoning model, called o3-mini, is built for math, coding, and other step by step problems. Instead of relying only on what it learned during training, it spends extra time thinking before it answers: it works through an internal chain of reasoning first, then gives its final response, with a variable reasoning effort setting that trades speed for accuracy depending on how hard the task is. At medium reasoning effort, o3-mini matches the performance of OpenAI's older, bigger o1 model on two hard tests, AIME 2024 (a competitive math exam) and GPQA Diamond (a difficult science test), while answering about 2,500 milliseconds, or two and a half seconds, faster per turn. It also uses a training method OpenAI calls deliberative alignment, meant to make it harder to trick into ignoring its limits.

DeepSeek, an AI lab, made a stronger point by releasing its reasoning models with the model files public, so anyone can download and run them themselves. It calls them DeepSeek-R1 and DeepSeek-R1-Zero, and the point of R1-Zero was to prove that a model can learn to reason well purely through reinforcement learning, meaning learning from a reward signal for getting the right answer, without first being shown human written example answers. Trained this way, using a method called GRPO, the model taught itself to double check its own work and reason through multiple steps, on its own, with no worked examples. To fix some rough edges, since the model would mix languages and write in a hard to follow way, DeepSeek added a short stage of curated starting examples before the reinforcement learning for the full R1 version.

DeepSeek then trained six smaller models to copy R1's reasoning, a technique called distillation (training a small model to copy a big one), ranging from 1.5 billion to 70 billion parameters, parameters being the model's internal adjustable settings and a rough stand in for size. Built on other open model families, Qwen2.5 and Llama3, these smaller copies are free for anyone to download and run on their own machines. The 32 billion parameter version, DeepSeek-R1-Distill-Qwen-32B, beat OpenAI's own smaller model, o1-mini, on several math and coding tests, despite being a fraction of the cost to run. That matters for anyone building with AI: a company can now run a frontier grade reasoning model, meaning close to the best available, on its own hardware, cutting reliance on paid outside models and easing worries about sending private data to an outside company.

The shift is also showing up in the tools around these models, not just the models themselves. Code based agent frameworks such as smolagents, a lightweight toolkit for building AI agents that write and run code, and fast serving systems such as vLLM-Omni, software that runs many AI requests at once efficiently, are making it easier to actually deploy these reasoning models in real products.

via Business Analytics Review →
04 Learnings

ChatGPT can now read your Mac texts and draft replies

OpenAI's new Apple Messages plugin lets ChatGPT read your iMessage conversations on a Mac and draft replies for your approval, like a quick I'm on it text back to your spouse. It works in ChatGPT's Work and Codex desktop modes, even on the free plan, after you grant Mac permissions once. Worth trying if small forgotten promises pile up.

via Matt Paige's Substack →

Reddit: AI's best use might be small home fixes

Outside your usual reading: a Reddit thread on using AI for home repairs, like fixing a stripped screw hole in a door hinge with gel super glue instead of the usual toothpick trick. The poster warns AI can also overreach, once suggesting a full teapot repair when replacing the unit outright would have been cheaper and easier.

via r/ArtificialInteligence →
05 Tools & Craft

TurboKV: a fast, open source data store written in Rust

TurboKV is a new open source key value database, a simple way to store and look up data by a label, built in the Rust programming language, with atomic batch writes, ordered range scans, and configurable durability and compression. It hit 114 points on Hacker News. Worth bookmarking if you or anyone in your orbit needs a fast, lightweight storage layer for a side project.

via GitHub →
The Last Word
The model you build on is a vendor, and vendors renew, or don't.
The Desk Report

How this edition came together — from bookmarks and feeds to the page.

225links gathered
40read by the desk
13made the edition

Where they came from

On the cutting-room floor — 27 links read but not run this week

Quality over volume: most links get a second look and a pass. The ones that made it earned their place.

Reading Room — every Sunday

The week's AI signal in 16 minutes — what happened, why it matters, and what to do with it. Curated by someone who actually builds, not a feed algorithm.