Skip the fancy AI search stack, most teams don't need it
A simple keyword search often beats an expensive embeddings pipeline.
You are designing the search and memory system for memory-os right now, so this is a direct check against building something fancier than the job needs.
The piece argues that most teams jump straight to the complicated version of retrieval, called RAG (retrieval-augmented generation, meaning the AI looks things up in your own documents before it answers). They reach for embeddings (turning text into numbers so a computer can compare meaning), vector databases, and reranking pipelines (a second pass that reorders results by relevance), when their users are really just typing something like "how do I reset my password."
The author says the right method partly depends on how people search: keyword-style questions should start with plain search, conversational questions benefit from embeddings, and a mix of both needs a blended approach.
The starting point, called the MVP (minimum useful version), is plain keyword search: the BM25 method, or tools like Elasticsearch and Postgres full-text search, the technology that predates embeddings. It is built for cases like a user typing "pandas merge dataframe" or searching for an exact string like "invoice #12345." It costs nothing per search, answers in under 10 milliseconds, and is easy to debug because you can see exactly why a document matched. It needs no decision about how to split documents into chunks, no complicated way to test whether it is working, and no risk of a model being retired out from under you, since BM25 itself never changes. Its weakness is that it misses synonyms and struggles with a conversational question like "how do I fix this." The author says it still handles a large share of real use cases on its own and should not be skipped.
Here is why that first step is underrated: jumping straight to embeddings forces you to answer questions you don't need yet, like what size to cut documents into, how much those chunks should overlap, whether to split by meaning or by fixed size, and how to even tell if your split was any good. Plain keyword search skips all of that. Your documents stay whole, and search just works on them directly.
The next step up uses an AI model to clean up a messy question into good keywords, at a cost of roughly a tenth of a cent per search (the source's example uses GPT-4o-mini, a small, cheap OpenAI model, priced near $0.001 per search). The insight behind it: most "semantic search" problems are really just a badly phrased question, not a limitation of keyword search itself. The AI model can strip filler words ("how do I" disappears). You can also give it your own glossary of terms up front. The example given: a company has an internal tool called "Atlas." A general-purpose embedding model reads "Atlas" and thinks Greek mythology or maps, scoring the real match at just 0.15 out of 1, essentially useless. Telling the AI model in advance, through a short instruction, which words are exact domain terms to preserve (Atlas: the internal data processing framework, Mercury: the messaging system, Zeus: the auth service) gets a perfect keyword match instantly, with no retraining needed.
This approach is also more forgiving to fix. With embeddings, a bad result means adjusting how you split documents, redoing the embeddings for the whole collection, and rerunning tests to see if it actually improved. With query rewriting, a bad result just means editing the instruction and testing again immediately. The source also describes a multi-turn version: the AI model rewrites the question, searches, checks whether the results look good enough, and if not, tries again with what it learned, up to a handful of times, all without touching the underlying documents.
The most advanced version blends both: keyword search pulls the top 50 to 100 candidates, then embeddings rerank them down to the best 10, because keyword search is fast and precise on exact terms while embeddings are better at catching meaning.
The lesson for memory-os: start with the plain keyword version, and only add the expensive layers once real usage data proves you need them, not before.