What is RAG in artificial intelligence?
RAG (Retrieval-Augmented Generation) is a technique that combines a language model (LLM) with a search system over your own knowledge base. Before answering, the system looks for relevant documents in your content (FAQs, products, pricing, contracts) and hands them to the LLM as context, so it answers from verifiable data rather than whatever the model remembers from its training.
Without RAG, an LLM answers based on what it learned during training. That means information that may be out of date, generic facts that aren't specific to your business, and a high risk of hallucination (inventing prices, policies or products that don't exist). With RAG, the model cites what you've given it and can show its source.
The typical RAG stack in 2026 uses a vector database (pgvector, Pinecone, Weaviate, Qdrant) that stores embeddings of your documents. When a question comes in, it is turned into an embedding, the most similar chunks are retrieved and they are injected into the LLM's prompt. Done well, the answers are specific and verifiable.
An applied example
An accountancy firm uploads all its service contracts, real FAQs from the past year and up-to-date regulations to a vector database. When a client asks on WhatsApp, "what happens if I'm late with modelo 130?", the AI agent searches for the relevant chunks, finds the firm's own exact answer (not generic answers from the internet) and replies by citing the internal document. The client gets information that is correct and specific to their contract.
When it is worth it
- You have internal documentation that rarely changes (FAQs, manuals, products, pricing) that the AI can use as a source.
- Your customers ask questions specific to your business that a generic LLM would answer badly.
- You need traceability: being able to explain where each answer comes from.
- Your volume of enquiries justifies a serious technical implementation with preprocessing, indexing and maintenance.
Common mistakes
- Uploading documents in poor formats (scanned PDFs, Word files with awkward tables) without preprocessing. Vector search then returns rubbish.
- Chunks that are too big (4000+ tokens) or too small (50 tokens). The ideal size is usually 200-800 tokens, depending on the content.
- Not re-ranking the results before passing them to the LLM. Cosine similarity is noisy when it comes to fine distinctions.
- Not filtering by metadata (category, date, author). Without filters, the AI can mix new content with outdated content.
- Expecting RAG to fix a poorly trained LLM or a badly designed prompt. RAG is only one piece of the system.
How we do it at STAKKER
STAKKER SYSTEMS implements RAG on pgvector (Postgres with a vector extension). We use it in production for web chatbots, WhatsApp and AI agents. It includes document preprocessing, smart chunking, re-ranking and observability of which chunks were served in each answer.
Frequently asked questions
Does RAG completely prevent hallucinations?
It reduces the risk a lot but doesn't eliminate it. If the cited document contains incorrect information, the AI will repeat it. RAG improves traceability, not the quality of the source.
Do I need to re-index when documents change?
Yes. If your pricing or a process changes, you need to update the document and re-index its embedding. With n8n you can automate this (a daily cron job or a trigger on edit).
Which vector database makes sense in 2026?
For small and medium volumes, pgvector on Postgres is the most sensible choice (no vendor lock-in, free, easy to run). To scale to millions of chunks, Qdrant or Weaviate.
How much does it cost to run RAG in production?
For a small business: no licence costs (it's all open source) and infrastructure costing a few tens of euros a month. The real cost is the initial preprocessing and keeping the quality of the documents up to scratch.
Related terms
- AI AgentAn AI agent is a software system that combines a large language model (LLM) with access to tools (APIs, databases, messaging) and a goal, so it can carry out tasks with minimal human involvement. The difference from a traditional chatbot is that an agent doesn't just chat: it acts.
- n8nn8n is an open-source automation platform that connects APIs, databases and SaaS tools through visual workflows. It lets you automate business processes without writing glue code, while keeping control of your data.
- AI AutomationAI automation combines traditional workflows (n8n, Zapier, Make) with large language models (LLMs) to carry out tasks that used to need human judgement: sorting messages, drafting replies, summarising documents, deciding escalation routes, creating content. It goes a step beyond the classic "if X happens, do Y".