Hybrid search ranks a result with both lexical match (keywords, BM25) and vector similarity (embeddings). Pure vector search is good at “something like this”. Learning products also need “exactly this word, this CEFR tag, this unit”. When we first shipped embeddings-only retrieval for Polylingo explanations, learners got fluent paragraphs about the wrong tense because “past” sat near “passed” in vector space and the unit filter was a suggestion, not a constraint.
The fix was not a bigger model. It was metadata. Every chunk now carries locale, level, unit id, and a stable document id. The query is: filter first, then blend BM25 and cosine, then ask the model to answer only from those rows. If the blend is empty, we say we do not have a lesson for that — we do not invent one.
Chunking for lessons, not books
We do not slice by token count alone. A grammar card is already a chunk. A Kitapix article is split on headings, then we keep the heading in the chunk so BM25 still sees “conditional sentences”. Overlap is small. Giant overlapping windows made the index look smart in demos and dumb in the classroom, because two chunks from different units mixed in the context window.
Re-embed when the course changes
Authors edit courses weekly. We version the index per course and rebuild the changed units, not the whole catalog. The assistant’s system prompt includes the index generation. If a teacher publishes unit 4 at 16:00, a learner at 16:05 should not get unit 3’s explanation with yesterday’s examples. That is a product bug with an SEO-shaped name: stale retrieval.
What we measure
Citation precision (did the answer point at the right card?), teacher override rate, and “I meant the other meaning” tickets. We do not measure “how often the model spoke”. Speaking is easy. Being on-unit is the job.
Questions we keep getting
Do we need a dedicated vector database? At our size, SQL Server plus a compact vector index has been enough for several products. We will split when recall or ingest forces it — not because a vendor keynote said so.
Why keep BM25 at all? Proper nouns, exercise codes, and Turkish suffixes. Embeddings blur them. Teachers search for codes.
Is this the same as GEO? Related. Hybrid search is how the product finds its own content. GEO is how answer engines find our public pages. Both reward structure and honest citations.