Companies beat top AI models by training their own smaller ones
Three real companies proved a smaller, custom-trained model can beat giant general AI models at one specific job.
Here is the direct takeaway for how you build your own AI agents: for a narrow job you do over and over, a small model trained specifically for that job can beat an expensive general purpose model, and cost much less to run each time you use it. Three real companies just proved this, each in a completely different business, and over the past two years the approach has hardened into a repeatable playbook. Step one, take a model whose files are public, called an open weights model, meaning anyone can download it and retrain it themselves. Step two, teach it your own real examples using reinforcement learning, which means training a model by rewarding it when it gets your specific task right and correcting it when it does not. Step three, score it against your own version of the actual job, not a generic public test that has nothing to do with what you actually need done.
Bridgewater Associates, one of the largest hedge funds in the world, has analysts who sift a constant stream of news articles, regulatory filings, and emails, judging which ones matter to the firm's investment views and where the useful content stops and boilerplate starts. The catch is that relevant means relevant by Bridgewater's own internal judgment, built up over decades, and no amount of clever prompting got the big frontier models, meaning the best general purpose AI models from the top labs, to reliably match that judgment. So Bridgewater trained an open weights model on labels created by its own expert investors, essentially teaching the model to think like their analysts already think. The trained model makes roughly 30 percent fewer mistakes than the best frontier model, and it costs a fraction as much to run each time, what the piece calls its inference cost, meaning the cost of getting one answer out of the model once it is trained and running.
Harvey, which builds AI agents for law firms, hit the same wall on its hardest work: due diligence on business transactions and drafting legal memos. These jobs are long horizon, meaning the agent has to take many steps in a row through large sets of documents, and small errors stack up and compound as the steps pile up, the way one wrong assumption early in a memo can poison everything that follows. Even the best frontier models, run at their maximum reasoning effort, the setting that makes a model think longer before it answers, kept falling short of the quality bar that law firms actually need to trust the output. Harvey's fix was to run reinforcement learning on an open weight model trained specifically on legal work, rather than continuing to push a general purpose model harder. The result outperforms both GPT-5.5 and Claude Opus 4.8, Anthropic's own top model, on Harvey's internal scoring rubrics, the standards it grades answers against, on the exact tasks its lawyers actually do.
Intercom's AI support agent, called Fin, resolves close to two million customer issues a week, and at that volume the real problem stops being accuracy alone and becomes unit economics, meaning what is left over once you subtract the cost of each answer from what it is worth to the business. Frontier model pricing per call adds up fast at that scale, and every extra point of issues resolved without a human matters directly to the bottom line. So Intercom's AI team post-trained its own vertical model, named Fin Apex, on billions of real customer service conversations gathered from its own platform, rather than continuing to lean on a general purpose model built for every task at once. Intercom reports that Fin Apex resolves more issues than the best frontier models, while costing less to run per conversation, a combination a general model could not match.
The article says the same shape shows up in eight more deployments beyond these three, collected in an appendix, each running its own version of the same recipe: pick an open weights model, gather your own real task data, and train against a scored version of the actual job instead of a generic benchmark that was never built for your use case. None of these three companies replaced their frontier model everywhere, and none of them are claiming a small model beats a big one in general. They replaced it for one narrow, high volume, well defined job, where they already had enough of their own labeled examples to teach a smaller model the one specific judgment call that mattered most to their business. That is the part worth carrying into your own agent work: before reaching for the biggest, most expensive model on a job you do over and over, ask whether you already have enough real examples of the right answer sitting around to train something smaller and cheaper to do it instead.
The trained model makes roughly 30% fewer mistakes than the best frontier model, at a fraction of the inference cost.via Fermisense →