RAG and LLM features

RAG implementation inside your existing product

Retrieval over your own knowledge, wired into the product you already ship, with answers that cite where they came from.

From
1,200 EUR
Timeline
Two to four weeks depending on the state of the documents

RAG has a reputation for being easy, because a working prototype takes an afternoon. Chunk the documents, embed them, search, stuff the results into a prompt. It demos well and it falls apart on contact with a real knowledge base, where half the documents are outdated, three of them contradict each other, and the important answer lives in a table inside a PDF.

The engineering that matters is not the retrieval call. It is deciding what goes into the index, what gets excluded, how a stale document stops being served, and how the answer shows its source so a person can check it. That is most of the work, and it is what separates a chatbot people trust from one they quietly stop using.

Who this is for

  • You have a body of knowledge worth searching: documentation, contracts, support history, product specs, internal wiki.
  • There is an existing product, portal or internal tool the feature can live inside.
  • Someone on your side can say which documents are authoritative and which are outdated.

If the knowledge lives only in people's heads, or nobody can rule on which version of a document is current, retrieval will amplify the mess. Fixing the source comes first.

What you get

  • An ingestion pipeline for your formats, including the awkward ones: PDFs, tables, scans, exports.
  • Retrieval tuned on your own questions, not on a generic benchmark.
  • Answers with citations back to the source document and section.
  • The feature integrated into your product or internal tool, not a separate site nobody opens.
  • An evaluation set of real questions with correct answers, so quality is a number you can track.
  • A refresh path: how new and changed documents enter the index, and how stale ones leave it.

How it runs

01
Audit the knowledge

What exists, in what formats, what is authoritative, what should never be served. Usually the most valuable step.

02
Build ingestion and retrieval

Parsing, chunking and indexing suited to your documents, then retrieval tuned against your own question set.

03
Wire it into the product

The interface where your users already are, with citations and a clear behaviour when the answer is not in the index.

04
Measure and hand over

Evaluation results, the refresh path, documentation and a walkthrough for your team.

Questions people ask first

Can you do this without moving our documents to a third party?
Yes. The index can live in your own infrastructure, and the model call can go to a provider with no retention or to a model you host. Each option costs differently in quality and money, and we decide it during the audit rather than after.
How do we know the answers are right?
Every answer cites the document and section it came from, so a person can verify in one click. Beyond that, we build an evaluation set from your real questions with known-correct answers, and every change is measured against it.
What happens when a document is updated?
The refresh path is part of the delivery. New and changed documents enter the index on a schedule or on a trigger from your system, and superseded versions stop being retrievable so the assistant cannot quote last year's price list.
Is a vector database required?
Often, but not always. For a few hundred documents, well-built keyword search plus reranking can beat a vector store on both accuracy and cost. I pick based on your corpus, not on what is fashionable.
Can this run on top of our existing chatbot?
Usually yes. If you already have a support bot or an in-product assistant, retrieval can be added as a source behind it rather than replacing the whole thing.

Related reading

All services