

RAG Without Frontier Models: Why Good Retrieval Matters More Than the Biggest LLM
2026-06-24 · Manuel Spörer
RAG systems are often reduced to the language model. Architecture diagrams then show a vector database on the left, a large LLM on the right, and an arrow labeled "context" between them. If the answers are not good enough, the next step seems obvious: use a larger model, add more parameters, or call a frontier model in the cloud.
In practice, the problem often lies elsewhere. The system retrieved the wrong passage, destroyed a table during ingestion, separated important terms through unsuitable chunking, or placed eight results from the same document in the prompt. A stronger language model may hide these errors more eloquently, but it cannot recover missing evidence.
That is exactly why our RAG setup is not built around a single model that is as large as possible. It is a pipeline of specialized models and deterministic steps: documents are processed while preserving their structure, split according to the tokenizer, indexed with additional context, searched using a hybrid approach, reranked, and only then passed to the answer model. This article explains the difference between a simple and an advanced RAG system, describes our current setup, and summarizes the most important dos and don'ts from implementing it.
What RAG Is Actually Supposed to Solve
Retrieval-Augmented Generation, or RAG, combines a generative model with an external knowledge base. Instead of answering a question exclusively from its model parameters, the system first searches for relevant evidence and passes it to the LLM together with the question [1].
This fundamentally changes the language model's job. It does not need to store all domain knowledge itself. It needs to read a small, relevant context, answer the question based on that context, and correctly attribute the sources it used. Domain knowledge, freshness, and traceability therefore reside largely outside the model.
This is also why sheer model size is less dominant in RAG than it is for open-domain knowledge questions. Without retrieval, a frontier model must combine as many capabilities and facts as possible in its parameters. Ideally, a RAG model receives exactly the eight passages it needs for the specific answer. The hardest question is then often not "Can the model know this information?" but "Did our pipeline find the right information and present it cleanly?"
What Is a Simple RAG System?
A simple RAG system usually follows a straightforward process:
Documents -> extract text -> fixed chunks -> embeddings -> vector database
Question -> query embedding -> top-k vector search -> chunks in prompt -> answer
This is a sensible starting point. Such a setup is quick to implement, easy to explain, and often sufficient for small, homogeneous document collections. It also reveals quickly whether the basic principle works for a particular use case.
Simple RAG does, however, have typical limitations. Pure vector search finds semantically similar text, but it can perform worse than traditional keyword search for exact product names, legal sections, abbreviations, or numbers. Rigid chunking separates tables from their headings or statements from the section that gives them meaning. A flat top-k can also consist almost entirely of passages from a single long PDF. The answer model formally receives context, but not good context.
RAG research therefore often distinguishes between naive, advanced, and modular RAG [2]. The label itself is not what matters. What matters is where the pipeline loses quality and which countermeasures actually help at those points.
What Makes Our System Advanced RAG?
Our query path currently consists of four core retrieval steps and a separate generation step:
Question
-> BGE-M3 query embedding
-> Weaviate hybrid search: BM25 + three vector spaces
-> document grouping: up to 12 documents x 4 chunks
-> BGE-reranker-v2-M3: around 48 candidates down to top 8
-> final document cap: no more than 3 chunks per document
-> Qwen3-30B-A3B generates the answer with source markers
Each step solves a different problem. The embedding model enables semantic search. BM25 catches exact terms and rare words. The cross-encoder scores the question and passage together. The document caps prevent a single extensive source from occupying the entire context. In the end, the generative model sees only a small, already curated evidence space.
1. Structure Comes Before Embedding
HTML, Markdown, and PDF files do not simply pass through the same text splitter in our system. Web pages are processed based on their readable content. PDF pages are converted by granite-docling-258M into DocTags and then into a structured Docling document. For this purpose, Docling supports a VLM pipeline that turns page images into structured document representations [8, 9].
A HybridChunker then splits the documents. Its tokenizer is pinned to BGE-M3—the exact model that later embeds the chunks. Our current limit is 512 tokens. Heading paths, section IDs, and page numbers for PDF documents are preserved as metadata.
That may sound like a detail, but it is one of the most important architectural decisions. If you chunk using an arbitrary character limit and later use a different tokenizer, you are measuring two different things. If you flatten the document structure beforehand, you cannot reconstruct it after embedding.
2. Contextual Retrieval for Isolated Chunks
A single chunk is often ambiguous. A sentence such as "The entitlement lasts for a maximum of 78 weeks" is hard to retrieve without the document title and section. Is it about sickness benefits, benefits for caring for a sick child, or something else?
For that reason, during ingestion our system generates a short context prefix for every non-degenerate chunk. The model sees the document header, the surrounding section, and the chunk itself. The resulting prefix contains one to three sentences and is placed before the chunk prior to creating both the embedding and the BM25 index. The raw content remains available separately.
This approach is based on Contextual Retrieval. Anthropic describes precisely this combination of contextualized embeddings and contextualized BM25 [3]. Timing is essential: the additional context must be generated before indexing. Appending it only after retrieval no longer improves the search.
In our system, the context pass does not receive the complete document for every chunk. Instead, we use the document header and section—or, for very long sections, a window around the current chunk. This limits cost and context size. Calls run sequentially so that llama.cpp can reuse its prefix cache.
3. Three Vectors Instead of One Generic Embedding
Each chunk receives three BGE-M3 vectors in our system:
| Vector | Content | Purpose |
|---|---|---|
content_vector | Context prefix plus chunk text | Semantic relevance of the passage |
title_vector | Heading path or document title | Thematic and structural classification |
description_vector | Document description | Document-level relevance and broad topic |
BGE-M3 is multilingual, supports long inputs, and was developed for several retrieval methods. Its model card explicitly recommends combining hybrid search and reranking for RAG [4]. In our setup, we use the model for dense vectors and combine them with Weaviate's BM25 index.
Weaviate itself does not create embeddings in this setup. The collection uses vectors supplied by our services. This keeps the model choice within those services, makes reindexing more reproducible, and prevents unnecessary coupling between database configuration and the model lifecycle. When using custom embedding models, Weaviate likewise recommends disabling automatic vectorization [6].
4. Hybrid Search Instead of Vector Search Only
The search runs simultaneously across BM25 and all three vector spaces. Weaviate fuses both result lists; the alpha parameter controls the share contributed by vector search [6]. Our starting value is 0.3, deliberately giving more weight to the keyword side than to the semantic side.
This is not a universal best practice. It is a starting value for our German-language corpus, which contains many technical terms, benefit names, years, and legal formulations. Exact tokens are especially valuable in this setting. A question about a particular legal section should not lose just because a different text sounds more similar semantically.
It is also important that filters take effect before scoring. Sources, individual URLs, and tags narrow the search space before results are scored. Filtering afterward would displace relevant candidates that fall outside the initially loaded top-k.
5. Reranking and Source Diversity
Hybrid search is fast and broad, but not precise enough for the final prompt. bge-reranker-v2-m3 therefore reads the question and each candidate together. Unlike a bi-encoder, the cross-encoder does not create independent vectors for each input; it scores the text pair directly. This is more accurate but also more expensive, making it suitable for a few dozen candidates rather than the entire index [5].
Before reranking, Weaviate groups the results by document. By default, we load up to 12 document groups with no more than four chunks each, resulting in roughly 48 candidates. Eight chunks remain after reranking. A second cap limits the final context to no more than three chunks per document.
Both stages are necessary. Grouping only after a flat top 30 would not help if the first 25 results already came from the same PDF. Results from other documents would never have been loaded. Conversely, despite a diverse candidate set, the reranker can once again move several passages from the same source to the top. That is why the second cap follows reranking.
This rule is relaxed for document-specific questions. If the scope consists exclusively of no more than three specific URLs, we search deeper within those documents. Advanced RAG does not mean applying every heuristic all the time. It also means selectively suspending heuristics to fit the type of question.
6. A Limited Job for the Answer Model
Only now does the generative model enter the process. The context is grouped by document and section, restored to reading order within each document, and enriched with title, URL, heading path, and page information. The system prompt permits answers based exclusively on this context and requires source markers in the format [N]. If the evidence is insufficient, the model is expected to say exactly that.
The current answer model is Qwen3-30B-A3B-Instruct-2507 in a quantized GGUF variant. It is a mixture-of-experts model with 30.5 billion total parameters, of which only around 3.3 billion are activated per token [7]. "Read eight relevant pieces of evidence, summarize them in German, and cite them correctly" is a much narrower job than open-ended research or general problem-solving.
Which Models We Use
Our setup is not a single-model system. Models are assigned by task:
| Model | Size or variant | Role in the system |
|---|---|---|
bge-m3 | GGUF F16, around 0.6B parameters | Query, content, title, and description embeddings |
bge-reranker-v2-m3 | GGUF Q8, around 0.6B parameters | Cross-encoder reranking of search candidates |
Qwen3-30B-A3B-Instruct-2507 | GGUF Q5_K_M, 30.5B total / 3.3B active | Answer generation in the current text RAG path |
Qwen3-VL-30B-A3B-Instruct | Local via llama.cpp | Context prefixes, tags, and document headers; potentially multimodal answers |
granite-docling-258M | 258M, VLM | Structured PDF conversion into DocTags |
All GPU-bound model calls run through an OpenAI-compatible interface. llama.cpp serves the GGUF models; llama-swap acts as a router in front of it and can load and swap models on demand [10, 11]. Serving runs on our infrastructure separately from the Docker application stack. Small embedding and reranking models can remain resident while larger generative models are loaded as needed.
Why We Do Not Need a Frontier Model for This
The short answer is that the generative model does not do all the work in this system.
A frontier model would undoubtedly be stronger at open-ended reasoning, ambiguous tasks, broad world knowledge, and highly complex synthesis. Our standard RAG path, however, deliberately narrows that task. Retrieval and the reranker handle knowledge selection. Document structure is produced during ingestion. Source diversity is enforced deterministically. Citation numbers come from a fixed mapping. The answer model must reliably process the supplied context, not reconstruct the internet from its parameters.
This produces four practical advantages:
- Data remains under our control. Documents and prompts do not have to be sent to an external model provider for every request.
- Costs become more predictable. After the hardware investment, there are no token-based API costs for each query or reindexing run. Electricity, operations, and maintenance do not disappear, of course.
- Models remain interchangeable. Embedding, reranking, and generation are separated by clearly defined interfaces. A new answer model does not require changing the vector database.
- Errors become easier to locate. A poor result is a retrieval problem. Incorrect ordering is a reranking problem. An unsupported statement is a generation or prompt problem. These layers blur together in a monolithic "large model does everything" setup.
The claim is therefore not that small models are fundamentally better than frontier models. It is this: For a clearly bounded, evidence-based RAG task, a frontier model is not automatically the most effective or economical lever.
Lessons Learned: What Really Matters in Practice
Retrieval Errors Matter More Than Generation Errors
If the correct passage is not in the context, the answer model can only refuse or guess. Changing models may improve the style, but not the evidence base. We therefore debug questions backward: Which chunks did the model see? Which candidates did the reranker see? What did hybrid search return? What was originally indexed?
Chunking Is Part of the Model System
Chunk size, tokenizer, table handling, headings, and page assignment all help determine which information can later be found. Chunking is not cosmetic preprocessing. It is a core part of retrieval design.
More Results Do Not Automatically Mean More Context
A larger top-k can make the answer worse. Similar or redundant passages compete for attention, long documents dominate the prompt, and citation mapping becomes less clear. Our two-stage document cap was therefore more important than another increase in context-window size.
Degradation May Work, but It Must Not Be Invisible
If the reranker fails, the system continues to answer using the hybrid-search order instead of turning every query into an HTTP 500 error. The state is marked as degraded and logged. One remaining gap is that the frontend does not yet display this flag to end users. Failing soft is right; silent quality loss is not.
Model Serving Is an Architectural Topic of Its Own
Contextual Retrieval creates many similar prompts. Without a prefix cache, our model had to process the same document context again for every chunk. Cache reuse and separate parallel slots for ingestion and chat significantly reduce prefill work. Model choice alone says little about actual throughput; caching, parallelism, quantization, and swap behavior are at least as relevant.
An "Advanced" System Does Not Need Every Advanced Feature
We deliberately do not use query rewriting for follow-up questions at present. The user's latest message goes directly into retrieval. This has a known limitation for fragmentary questions such as "And how much does it cost?", but it also avoids another generative step that could alter the search intent. If measurements show that this limitation matters, there is a clear integration point. Until then, the pipeline remains simpler.
Without Evaluation, Even a Good Architecture Remains a Hypothesis
The current values—512 tokens per chunk, alpha = 0.3, 12 times 4 candidates, top 8, and no more than three final chunks per document—are reasoned starting points. They have not yet been calibrated against a complete golden set. The planned comparison between Qwen3-30B and Qwen3-VL for text answers is also still pending.
This makes the article a robust architecture and experience report, but not scientific proof that our local model beats every frontier model. The correct next step is not a larger model based on a hunch, but a reproducible evaluation of retrieval results, answer faithfulness, citations, latency, and cost.
Dos and Don'ts for RAG Systems
| Do | Don't |
|---|---|
| Measure retrieval and generation separately | Judge answer quality only by writing style |
| Align chunking with the embedding model's tokenizer | Split documents blindly by characters or paragraphs |
| Combine BM25 and vector search | Rely exclusively on semantic similarity |
| Use a specialized reranker for a limited candidate pool | Run a cross-encoder across the entire index |
| Ensure source diversity before and after reranking | Let one long document flood the entire context |
| Generate context prefixes before embedding and BM25 indexing | Attach context only after retrieval |
| Store raw text, search text, and metadata separately | Irreversibly merge generated context into the original text |
| Apply filters before scoring | Load top-k first and filter unsuitable results afterward |
| Make degradation visible and measure failures | Silently swallow losses in quality |
| Keep models interchangeable behind stable interfaces | Tightly couple the database, framework, and model lifecycle |
| Evaluate retrieval and answers separately with a golden set | Use model size as a substitute for measurement |
When a Frontier Model Still Makes Sense
There are tasks for which we would still seriously consider a frontier model: open-ended research across many unknown sources, highly complex multi-step reasoning, demanding agent planning, weakly structured requests with substantial ambiguity, or a quality baseline for evaluating local models.
A poor or incomplete corpus can also shift the model requirements. If the answer is not stated explicitly in the documents and must be inferred from many scattered clues, the LLM's synthesis capabilities become more important. That is then a deliberate product feature, not a reason to route every standard RAG request to the largest available model by default.
There is also a clear boundary in specialized and regulated domains: neither a local model nor a frontier model replaces expert validation, good sources, and systematic evaluation. A larger model is not a compliance mechanism.
Conclusion: The Best RAG Model Is a Good Pipeline
With RAG, no single model determines quality. What matters is the interaction between document processing, chunking, embeddings, search methods, reranking, context construction, and generation.
Our setup therefore uses several specialized open-weight models instead of one frontier model for everything. BGE-M3 makes text retrievable. BGE-Reranker-v2-M3 sorts candidates more precisely. Granite Docling preserves the structure of PDF pages. Qwen3 turns a small number of selected pieces of evidence into a German-language answer with citations. Weaviate, llama.cpp, and llama-swap keep the components technically separate and locally deployable.
The crucial point is not that 30 billion parameters are always enough. The point is that a well-designed RAG system gives the model a task for which they can be enough. Before buying the next frontier model, check whether the current system is finding the right eight passages in the first place.
Sources
[1] Lewis et al. Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. 2020.
[2] Gao et al. Retrieval-Augmented Generation for Large Language Models: A Survey. 2023.
[3] Anthropic. Introducing Contextual Retrieval. 2024.
[4] Beijing Academy of Artificial Intelligence. BGE-M3 Model Card.
[5] Beijing Academy of Artificial Intelligence. BGE Reranker v2 M3 Model Card.
[6] Weaviate. Hybrid Search and Bring Your Own Vectors.
[7] Qwen Team. Qwen3-30B-A3B-Instruct-2507 Model Card.
[8] Docling. Vision Models.
[9] IBM Granite. granite-docling-258M Model Card.
[10] ggml-org. llama.cpp.
[11] mostlygeek. llama-swap.