Why one team's AI feature cost exploded in production
Real users pasted huge amounts of text into a feature built for short prompts, and the cloud bill followed.
A team building an AI feature explains, in their own postmortem, how the bill quietly got out of control once real users showed up, and it is worth reading because it is the exact trap waiting for anyone shipping an AI product, including your own work.
The feature looked fine in testing. Prompts were small: summarize this, draft that, maybe using a 2,000 token, a token is a chunk of text roughly three quarters of a word, context window, how much text the model can hold in mind at once, on a busy day. Then production users arrived with 14 paragraph questions, pasted half their customer database into the box, and kept asking follow ups without ever resetting the conversation. On top of that, the retrieval layer, the part of the system that looks things up in the company's own documents before answering, kept pulling in the same three chunks of text repeatedly, because one copy of a stale policy document did not feel like enough to the system.
The team's own account of how it happened is the useful part. Accuracy looked better when they stuffed more text into the prompt, so they did that. Then answers got slow, so they trimmed the text back. Then quality dropped on harder questions, so they added context back in. Only later did someone notice the system was sending nearly identical retrieved text, plus a large system prompt, plus the full conversation history, on every single turn, whether or not any of it was new information.
The number that stuck: one user asking a single long, unusual question could cost more to answer than the company's entire demo flow used to run end to end. That is what showed up in the budget meeting, and by their account it was not a pleasant one.
Their fix was not use fewer tokens as a slogan. It was specific engineering: breaking long documents into smaller chunks, removing duplicate retrieved text before it reaches the model, capping how much old conversation gets carried forward, and testing different ways to trim context without losing accuracy. They are still deciding, case by case, whether slower answers or lower accuracy is the tradeoff worth making that week. If you or anyone building on your behalf ships something with an open text box, this is the checklist to have ready before real users arrive, not after the first invoice.
Nothing like explaining that a user asking a long weird question can cost more than the entire happy path demo flow.via Reddit r/ArtificialInteligence →