Search that answers in sentences,
and always shows where the answer came from
Keyword search returns a list of links. A retrieval-augmented search answers the actual question and shows exactly which document it came from. We build the retrieval layer to be precise first, since a confident wrong answer is worse than a slow right one.
Why a plain-language answer beats a list of links
A RAG search product answers questions in plain language. It retrieves relevant passages from your own documents or catalogue, generates an answer grounded in them, and attaches a citation showing exactly where that answer came from. It fits any large document collection, product catalogue, or knowledge base where keyword search returns too many irrelevant results and a person just wants an actual answer. It is not worth the complexity for a handful of pages a person can read directly. The value shows up once volume is large enough that finding the right passage by hand takes real time.
How the retrieval actually works
Documents are chunked and indexed with careful attention to chunk size and overlap. Retrieval quality depends heavily on getting this right for your specific content, not a one-size-fits-all default. Retrieved passages feed into the model, which generates an answer grounded specifically in what was retrieved. The source citation stays attached, so a person can verify the claim directly, rather than trusting it blindly. A relevance check on retrieved results catches the case where nothing in the index actually answers the question. It declines to answer rather than generating a confident response from weak or unrelated matches. Re-indexing keeps the search current as documents change, so an outdated document never continues to surface as if it were still accurate.
How we get retrieval right before adding generation
We start by understanding your actual document structure and content type. Chunking strategy differs significantly between, say, a legal document and a product catalogue, and getting this wrong is the most common cause of bad retrieval. We test retrieval quality against a set of real questions with known correct answers, before connecting a generation model on top. A retrieval problem disguised as a generation problem wastes time fixing the wrong layer. The relevance check gets tuned against deliberately irrelevant queries, to confirm it declines appropriately rather than hallucinating from unrelated context. We launch against one document source, measure real answer quality, and expand once that foundation is solid.
Where retrieval quietly goes wrong
The real risk is a confident answer built on a weak or irrelevant retrieved passage. That is why the relevance check exists, and why it gets tested against deliberately irrelevant queries before launch, rather than assumed to work from the architecture alone. Chunking strategy is the other place mistakes compound quietly. A chunk boundary that splits a critical sentence in half can leave a document technically indexed but practically unretrievable for the exact question it should answer. Re-indexing discipline matters as documents change. A RAG system is only as current as its last index update, and a stale index producing a plausible but outdated answer is worse than an obvious gap.
Timeline and price
| Option | Price | What it covers | Timeline |
|---|---|---|---|
| MVP | from $2,000 | One document source, citation-backed answers, relevance check | 4 to 5 weeks |
| Production | from $5,000 | Multiple sources, access control, scheduled re-indexing, query logging | 6 to 8 weeks |
| Full control (handover-ready) | from $6,000 | Everything in Production, plus a full handover package: architecture docs, test suite, admin access audit, and a walkthrough so your own team or another vendor can run it without us | 8 to 9 weeks |
Running cost on top of the build is usually $20 to $65 a month in vector-store and model costs, depending on document volume and query frequency.
What you own at the end
You own the document index, the retrieval pipeline, the citation logic and the full source code, running on your own infrastructure. The handover package documents the chunking and retrieval strategy, so adding new document sources later follows an established, tested pattern.
Related
Pairs with the AI knowledge base assistant for staff-facing use and the AI support agent product for customer-facing use, both built on this same retrieval foundation. See the AI agents service page for the full range of agent and search builds. Real builds: the visa consulting centre support bots case study and the SENET memory engine case study, both built around retrieval that cites its sources. Have a document collection too big to search by keyword alone? Get in touch.
FAQ
How much does a RAG search product cost?
From $2,000 for retrieval over a single, reasonably organized document set with citation-backed answers. Multiple sources, access control and re-indexing on a schedule runs $5,000 to $6,000.
How long does it take?
Four to five weeks against a well-structured document set. Messy or very large document collections need more time upfront for chunking and indexing strategy, typically extending to seven to nine weeks.
What is the stack?
A vector store, either pgvector on PostgreSQL or a dedicated vector database for large corpora. Python for chunking and retrieval logic. Claude or GPT to generate the final answer from retrieved context, with citations tied to source IDs.
Who owns the index and the retrieval logic?
You. The document index, the retrieval pipeline and the code run on your own infrastructure, with no dependency on a search-as-a-service subscription.
What happens when no document actually answers the question?
A relevance check on retrieved results catches weak matches and the system declines to answer confidently from them, rather than generating a plausible-sounding response from unrelated context.