A RAG chatbot doesn't answer from memory. For each question it searches your help center, policies and product sheets. It pulls the few most relevant passages, and the language model writes an answer from those passages only, ideally with a link to the source. Your documents decide the quality more than the model does. If you want a team to set it up, start with a support bot that answers from your knowledge base.
Below: how retrieval works in plain terms, what the bot should never answer, a readiness checklist for your documents, a test set to run before launch and escalation rules for the moments it gets things wrong.
Short answer: four steps per question
Every customer message runs through the same loop.
- Question. The customer types "Can I return a sale item after 30 days?"
- Search. The system looks through an index of your documents for passages that match the meaning of the question.
- Passages. It takes the top few matches — say, the returns policy and the sale terms page.
- Answer with source. The model gets the question plus those passages and an instruction: answer only from this text, cite it, and say "I don't know" if the text does not cover it.
That is the whole idea, and the term comes from a 2020 research paper by Lewis and colleagues, who described models that combine a language model with a searchable "non-parametric memory" and found they produce "more specific, diverse and factual language" than the model alone (Lewis et al., arXiv:2005.11401, as of September 30, 2026).
For support, this has one big consequence. When a policy changes, you edit the document and the next answer uses the new text, with no retraining.
How retrieval-augmented generation works, without the jargon
Four parts do the work. You'll hear their names in every vendor call, so here is what each one means.
Chunks. Long documents get split into short pieces. A returns policy might become six chunks of a few paragraphs each. OpenAI's hosted file search, for example, splits files into chunks of 800 tokens with a 400-token overlap by default (OpenAI retrieval guide, as of September 30, 2026). A token is roughly part of a word. Chunk size matters: too small and the chunk loses context, too large and the search gets fuzzy.
Embeddings. Each chunk is turned into a list of numbers that captures its meaning. Two chunks about refunds end up close together, even if one says "refund" and the other says "money back."
Search. The customer's question gets the same treatment, and the system finds the chunks closest in meaning. Good setups add plain keyword search too, because product codes and order statuses are exact strings that meaning-based search can miss.
Prompt. The winning chunks are pasted into the instructions for the model, along with rules: stay within this text, cite the source, hand off when unsure.
Retrieval is where most answers go wrong. If the right passage never reaches the model, no model can fix it. Anthropic published test results in September 2024 on this exact point. Adding context to each chunk cut the top-20 retrieval failure rate by 49% when combined with keyword search. Adding a reranking step raised the cut to 67%, from 5.7% of queries down to 1.9% (Anthropic, "Introducing Contextual Retrieval," September 19, 2024, as of September 30, 2026). Those are the vendor's own benchmarks, so treat them as direction and run your own test set.
The same post adds a useful caveat. If the whole knowledge base is under about 200,000 tokens, you can put all of it in the prompt and skip retrieval entirely. A small help center may not need RAG at all.
What a support bot should and shouldn't answer
Scope is a business decision, and you make it before anyone writes a prompt.
Good candidates share one trait: a stable answer and a clear data source. Think opening hours, order status, price lists, booking and rescheduling, and requests for documents such as invoices or warranty terms. These questions repeat, and the right answer sits in one place.
Bad candidates carry consequences. Anything that sounds like legal, medical or financial advice belongs here, along with refund exceptions, complaints and contract terms that depend on the individual account. The rule is simple. Open-ended advice and anything with legal or financial consequence belongs with a person from the first version.
Write the out-of-scope list down and give it to the bot as explicit instructions. Put it in the test set too, as questions the bot must refuse.
One more line to draw: account data. A bot that reads order status needs a verified customer session and an API call. Document search is the wrong tool there. Keep transactional lookups on scripted paths and free-form answers on RAG. Most working support bots combine the two.
Is your knowledge base ready?
Run this checklist before you choose a platform, because every "no" becomes a wrong answer later.
| # | Check | Why it matters |
| 1 | One source of truth per topic | Two returns pages means two possible answers |
| 2 | Every document has a "last reviewed" date | The bot can't tell stale text from current text |
| 3 | Every document has a named owner | Someone must fix it when the bot quotes it wrong |
| 4 | FAQ and policy pages agree | The search may pull the FAQ, while the policy governs |
| 5 | Prices, fees and deadlines live in a table or structured field | Numbers buried in prose get misread or mixed |
| 6 | Internal and confidential documents are kept out of the index | Whatever is indexed can end up in an answer |
| 7 | Region-, plan- or product-specific rules are labeled | "Free shipping" in one state is not free in another |
| 8 | Each page covers one question or one topic | Mixed pages produce mixed chunks |
| 9 | Headings describe the content | Chunks often carry the heading as context |
| 10 | No answers live only in images or PDFs scanned as pictures | Text inside images is usually not searchable |
| 11 | Retired products and old policies are removed or marked | Old text keeps winning searches |
| 12 | A review schedule exists, with a date on the calendar | The bot degrades as the business changes |
Expect to spend real time here. The fixes are editorial work: merging duplicates, dating pages, deleting old ones. It is dull. It also improves the help center for human readers.
Test before launch: a question set
A demo proves little. The real exam is 30 to 50 questions pulled from your ticket history, plus a separate block the bot must refuse.
Use this template. One row per question.
| Customer question | Expected answer | Source document | Must cite? | Hand off to human? | Pass/fail |
| "Can I return a sale item?" | Yes within 14 days, store credit only | Returns policy, section 3 | Yes | No | |
| "Where is order 48213?" | Status from the order system after login | Order API (scripted path) | No | No | |
| "Can you waive the restocking fee?" | "I can't decide that — connecting you to an agent" | Out-of-scope list | No | Yes | |
| "Is this product safe for my pregnancy?" | "I don't know" + handoff | Out-of-scope list | No | Yes | |
| "What's your price for the Pro plan in Canada?" | Price from the pricing table, with link | Pricing table | Yes | No |
How to build it:
- Pull questions from tickets exactly as customers wrote them. Typos included.
- Cover your top ticket categories in proportion to their volume.
- Add 8 to 10 "must say I don't know" questions: topics you don't cover, competitor questions, advice requests.
- Add 3 to 5 questions where two documents disagree. See which one wins.
- Write the expected answer before you run the test. Otherwise you'll grade generously.
Track three numbers per run:
- Answers with a correct source. A right answer citing the wrong page is a failure waiting to happen.
- Correct "I don't know" rate. Out of the must-refuse block, how many did it refuse?
- Escalation rate. Too low can mean the bot is bluffing. Too high means your documents have gaps.
Rerun the full set after every change to documents, prompts or models, and keep the results. When someone asks why the bot said something odd, you'll have a baseline.
NIST's Generative AI Profile (NIST AI 600-1, July 26, 2024) lists confabulation — confident, false output — among the risks of generative AI and recommends pre-deployment testing (NIST, as of September 30, 2026). It builds on the voluntary AI Risk Management Framework 1.0, released January 26, 2023 (NIST AI RMF, as of September 30, 2026). Neither is a law. Both are good checklists for your own test plan.
When the bot is wrong: liability and handoff
Wrong answers will happen. Plan for them.
The best-known case is Canadian. In Moffatt v. Air Canada, 2024 BCCRT 149, decided February 14, 2024, British Columbia's Civil Resolution Tribunal heard a claim from a customer whose airline chatbot said he could apply for a bereavement fare after travel. The airline's policy said otherwise. Air Canada argued, in effect, that the chatbot was a separate legal entity. The tribunal called that "a remarkable submission" and wrote that the airline "is responsible for all the information on its website," whether it comes "from a static page or a chatbot." It ordered Air Canada to pay $812.02 in total, including $650.88 in damages (Moffatt v. Air Canada, 2024 BCCRT 149, paras. 27 and 44, as of September 30, 2026).
That is a small-claims decision in one Canadian province. It says nothing about US law, but the practical lesson travels anyway: your customer will treat the bot's answer as your answer. This article is not legal advice.
So design the exit first. The bot needs clear triggers to stop and pass the conversation to a person.
| Trigger | Where it goes | What goes with it |
| Customer asks for a human | Live agent queue, or ticket if after hours | Full transcript, customer ID, detected topic |
| No passage above the relevance threshold | Agent queue with "no source found" tag | Question, top passages that were rejected |
| Two failed answers in a row ("that's not what I asked") | Agent queue, priority bump | Transcript, both bot answers |
| Refund, cancellation or complaint keywords | Billing or retention team | Order number, account status, transcript |
| Legal, medical, safety or financial advice request | Human specialist or a fixed safe reply + ticket | Transcript, flag for review |
| Angry language or threats | Senior agent | Transcript, sentiment flag |
| Outside working hours | Ticket with promised response time | Everything above, plus contact preference |
The last column is where most bots fail. Customers forgive a bot that can't answer. They don't forgive repeating their order number to the agent it hands them to. Pass the transcript along with whatever the bot already identified about the customer and the issue.
Then read the logs and review unanswered questions weekly for the first month. Each one points to a missing or unclear document. Fix the document, add the question to the test set, rerun.
Build on a platform or custom
You have two broad routes, and both use the same retrieval idea.
Built-in bots in your helpdesk. Most major helpdesk and chat platforms now ship an AI agent that reads your help center. Setup is fast. The bot sits where your agents already work, and handoff inside the same tool is usually smooth. The limits show up later: you get the vendor's retrieval settings, the vendor's model choices and the vendor's reporting. Pricing is often per resolution or per conversation, so check the vendor's pricing page and model the cost at your real ticket volume.
Your own RAG setup. You pick the model, the search, the chunking and the logging. You can pull from sources the helpdesk can't see: product databases, internal wikis with a filtered export, order systems through an API. You own the test set, the transcripts and the prompts, and the cost is engineering time up front and someone to run it after launch.
A rough way to choose:
- Help center under a few hundred articles, one channel, standard questions: start with the built-in bot. Use the checklist and test set anyway.
- Answers depend on data outside the help center, or you need strict control over what gets indexed: a custom setup pays for itself in control.
- High volume with per-resolution pricing: run the numbers for year two, not month one.
Support is rarely the only process worth handing to software. If you are mapping the rest, see other support workflows worth automating before you commit budget to a single bot.
FAQ
Do I need to fine-tune a model for a support bot?
No, for most support use cases retrieval is enough. Fine-tuning changes how a model writes. It is a poor way to teach facts that change every quarter, like prices or policies. With retrieval, you update a document and the next answer reflects it. Fine-tuning can help with tone or format later, once the retrieval side works.
How long does it take to prepare the knowledge base?
It depends on how many documents disagree with each other. A tidy help center with dated, owned pages may need a week of review. A help center with years of duplicates, old PDFs and conflicting FAQ pages can take much longer. Run the 12-point checklist first. The number of "no" answers gives you a realistic estimate.
Can a RAG bot use documents in several languages?
Yes, many embedding models handle multiple languages, and a question in Spanish can match an English passage. Quality varies by model and language pair. Test it. Add questions in each language to your test set, and check that answers cite the right regional policy. Label region-specific documents clearly so the search does not mix them.
What should the bot say when it doesn't know?
It should say so in one sentence and offer the next step. For example: "I don't have that information. I can connect you with an agent, or you can leave your email for a reply within one business day." Avoid guessing, apologizing at length or repeating the question. A clean refusal plus a working handoff keeps trust.
Is customer data sent to the model provider?
Usually yes, at least the question and the retrieved passages. Check each provider's data retention and training terms before launch. Keep confidential documents out of the index, and mask card numbers and similar fields before they reach the model. For account lookups, prefer a scripted API call that returns only the needed field.