TL;DR: For Turkish RAG, the embedding model should be chosen with a small evaluation set built from your own documents and real user queries, not from a generic benchmark ranking. Decision criteria are language coverage, dimensionality, cost and latency; the hidden cost is that switching models later means regenerating all vector data.
In Retrieval-Augmented Generation (RAG) systems, answer quality depends less on the generation model than on the quality of the retrieved passages. The component that retrieves them is the embedding model, and in an agglutinative language like Turkish this choice is less straightforward than English examples suggest.
What does an embedding model do in RAG?
An embedding model converts text into a fixed-size numeric vector and places semantically similar texts close together in vector space. In RAG, both document chunks and the user query are vectorized with the same model; the nearest chunks are passed to the generation model as context. If the model retrieves the wrong chunk, the answer stays wrong no matter how strong the generation model is.
Why does Turkish need its own evaluation?
Turkish is agglutinative: a single word can take dozens of suffixes. "Sözleşmelerimizdeki", "sözleşmeye" and "sözleşmeden" carry the same concept in different surface forms. A model trained mostly on English data may place these forms farther apart than it should. General multilingual benchmark scores do not measure this behavior because they do not contain your domain terminology, abbreviations or writing habits.
How do you compare models?
A comparison means measuring candidate models on the same data with the same metric. A practical flow:
- Build an evaluation set. Pick 50-100 questions from your real documents and manually mark the correct passage for each.
- Index with identical chunking. Chunk size and overlap must be the same for all candidates; otherwise you measure chunking, not the model.
- Compute metrics. The hit rate of the correct passage in the top 5 (recall@5) and its mean rank (MRR) are enough.
- Read the failures. A score table does not answer "why"; looking at ten failed queries shows whether the problem is suffix handling or domain terms.
The table below summarizes which criteria to check alongside the score:
| Criterion | Why it matters | How to measure |
|---|---|---|
| Turkish retrieval quality | The actual target metric | recall@5 and MRR on your own set |
| Vector dimensionality | Drives storage and search cost | Document count × dimensions × 4 bytes |
| Maximum input length | Limits chunk size | Model documentation |
| Latency and price | Per-query cost and user experience | Measure under real load |
| Hosting model | Data residency and privacy impact | Provider contract and region |
Why is switching models later expensive?
Switching the embedding model requires regenerating the entire vector index. Vectors from different models do not share a space; comparing an old vector with a new query vector gives meaningless results. If the dimensionality changes, the index schema changes too. So "we can swap it later" is a misleading assumption: choose by measurement, record which model produced each vector, and plan the switch in two stages — build the new index alongside, validate it, then move traffic.
Which practices protect you in production?
- Keep the embedding model name in a single configuration point and tag every vector record with the model that produced it.
- Store the evaluation set in the code repository and rerun the same test on every model or chunking change.
- Do not leave retrieval to vectors alone; combine it with keyword search for queries that need exact matches, such as product codes, abbreviations and proper names.
At Exponential Yazılım we build production AI components end to end, measurable and sustainable in the long term. The principle for embedding selection is the same: measure with your own data first, then decide.
