A RAG pipeline that answers from your documents
not from whatever the model happens to remember
An AI assistant that answers confidently from memory and one that answers correctly from your own documents are two very different products. We build the retrieval pipeline behind the second kind. Chunking, embeddings and vector search ground every answer in content you actually wrote, with a source it can point to.
Grounded answers instead of confident guesses
Retrieval augmented generation, known as RAG, is the technique behind most AI assistants that answer accurately about your business specifically. Your documents get split into chunks. Each chunk gets converted into a vector embedding that captures its meaning, and those vectors live in a vector database. When a question comes in, the pipeline finds the chunks closest in meaning to that question. It hands them to the model along with the question. The model then answers from that retrieved text, not from whatever it happened to pick up during training.
When a knowledge base earns its keep
You need this once an AI assistant has to answer questions about content specific to your business. That means your actual pricing, your policies, your product documentation, your internal knowledge base, content a general-purpose model was never trained on. It cannot know what it never saw. This is also the right build when accuracy and traceability matter. A support agent that can cite the exact policy document it answered from is trustworthy in a way a model’s confident-sounding guess never is.
You do not need a RAG pipeline for a general-purpose assistant with no proprietary knowledge to ground answers in. The same goes for a narrow task with a small, fixed set of answers that fits comfortably in a prompt with no retrieval step at all. Building a vector database for twenty FAQ entries is more infrastructure than the problem needs.
How we tune the pipeline, not just ship it
Most RAG pipelines get chunking wrong by defaulting to a fixed character count no matter what the content looks like. We chunk around your actual document structure instead: sections, headings, logical units. A retrieved chunk ends up a coherent piece of meaning, not an arbitrary slice that cuts a sentence in half.
We default to pgvector on PostgreSQL. That keeps vector search inside a database you likely already run, so there is no separate system to operate. We move to a dedicated vector store like Qdrant only when scale or filtering needs genuinely call for it.
Retrieval gets tuned against a real evaluation set, a list of representative questions with known correct answers. We do not ship on a default configuration and hope it works. We adjust chunk count, similarity threshold and reranking until the evaluation set’s accuracy is actually good enough to launch. Every answer carries a citation back to its source chunk. That lets a user verify it, and it lets us see exactly where a wrong answer came from if one shows up. We built pipelines on this same pattern twice before. One powers an AI persona answering as a subject-matter expert. The other powers a multi-channel sales agent that answers from a real product catalogue, not an approximation of one.
What to watch
The most common RAG failure is not the model making things up. It is retrieval coming back empty and the model answering anyway, as if it had found something real. We build an explicit “I don’t have information on that” fallback specifically to stop the model from papering over a retrieval gap with a confident guess.
Document freshness is the ongoing cost of ownership. A knowledge base that updates its policies but never re-indexes the change keeps answering from the outdated version. The refresh pipeline matters as much as the initial build. Lock-in stays low: pgvector and open embedding models keep the pipeline portable across LLM providers. That matters because the model layer on top of retrieval is exactly where provider pricing and capability shift fastest.
Price and timeline
| Scope | Price | Timeline |
|---|---|---|
| Single knowledge base, pgvector | from $1,500 | 2 to 3 weeks |
| Multiple sources, reranking, evaluation set | from $4,000 | 4 to 5 weeks |
Related
Built as part of AI agents and custom development. Pairs with an LLM gateway with cost control and prompt and knowledge versioning to keep the whole AI layer maintainable. See it grounding an AI persona answering as a digital expert and a seven-channel AI sales agent. Tell us what documents your assistant needs to answer from: get in touch.
FAQ
How much does a RAG pipeline cost?
A setup covering a single knowledge base starts at $1,500. That covers chunking, embeddings, retrieval and citation. A fuller pipeline with reranking, multiple document types and an evaluation set runs $3,000 to $7,000.
How long does it take?
2 to 5 weeks. The pipeline itself goes together faster than the tuning. Getting chunk size, retrieval count and similarity thresholds right against your actual documents takes real iteration, not a default configuration pulled off a shelf.
What vector database do you use?
pgvector on PostgreSQL for most setups. It avoids adding a separate database to operate, and it performs well at the scale most of our clients need. A dedicated vector store, Qdrant or Weaviate, makes sense at larger scale or with specific filtering needs.
Does this stop the AI from making things up?
It cuts the problem down a lot by grounding answers in retrieved text rather than the model's trained knowledge, and citation lets a person check the source. No RAG setup removes the risk completely though. We build in explicit fallback behavior for when retrieval finds nothing relevant, rather than letting the model guess.
Who owns the knowledge base and the pipeline?
You do. The documents, the embeddings and the retrieval code live on your own infrastructure. The pipeline works with any LLM provider, nothing is locked to a single vendor's platform.