info@toimi.pro
Thank you!
We have received your request and will contact you shortly
Okay
Web development

How does a RAG chatbot answer customers from your own documents?

15 min
Web development

A RAG chatbot doesn't answer from memory. For each question it searches your help center, policies and product sheets. It pulls the few most relevant passages, and the language model writes an answer from those passages only, ideally with a link to the source. Your documents decide the quality more than the model does. If you want a team to set it up, start with a support bot that answers from your knowledge base.

Below: how retrieval works in plain terms, what the bot should never answer, a readiness checklist for your documents, a test set to run before launch and escalation rules for the moments it gets things wrong.

Short answer: four steps per question

Every customer message runs through the same loop.

  1. Question. The customer types "Can I return a sale item after 30 days?"
  2. Search. The system looks through an index of your documents for passages that match the meaning of the question.
  3. Passages. It takes the top few matches — say, the returns policy and the sale terms page.
  4. Answer with source. The model gets the question plus those passages and an instruction: answer only from this text, cite it, and say "I don't know" if the text does not cover it.

That is the whole idea, and the term comes from a 2020 research paper by Lewis and colleagues, who described models that combine a language model with a searchable "non-parametric memory" and found they produce "more specific, diverse and factual language" than the model alone (Lewis et al., arXiv:2005.11401, as of September 30, 2026).

For support, this has one big consequence. When a policy changes, you edit the document and the next answer uses the new text, with no retraining.

How retrieval-augmented generation works, without the jargon

Four parts do the work. You'll hear their names in every vendor call, so here is what each one means.

Chunks. Long documents get split into short pieces. A returns policy might become six chunks of a few paragraphs each. OpenAI's hosted file search, for example, splits files into chunks of 800 tokens with a 400-token overlap by default (OpenAI retrieval guide, as of September 30, 2026). A token is roughly part of a word. Chunk size matters: too small and the chunk loses context, too large and the search gets fuzzy.

Embeddings. Each chunk is turned into a list of numbers that captures its meaning. Two chunks about refunds end up close together, even if one says "refund" and the other says "money back."

Search. The customer's question gets the same treatment, and the system finds the chunks closest in meaning. Good setups add plain keyword search too, because product codes and order statuses are exact strings that meaning-based search can miss.

Prompt. The winning chunks are pasted into the instructions for the model, along with rules: stay within this text, cite the source, hand off when unsure.

Retrieval is where most answers go wrong. If the right passage never reaches the model, no model can fix it. Anthropic published test results in September 2024 on this exact point. Adding context to each chunk cut the top-20 retrieval failure rate by 49% when combined with keyword search. Adding a reranking step raised the cut to 67%, from 5.7% of queries down to 1.9% (Anthropic, "Introducing Contextual Retrieval," September 19, 2024, as of September 30, 2026). Those are the vendor's own benchmarks, so treat them as direction and run your own test set.

The same post adds a useful caveat. If the whole knowledge base is under about 200,000 tokens, you can put all of it in the prompt and skip retrieval entirely. A small help center may not need RAG at all.

What a support bot should and shouldn't answer

Scope is a business decision, and you make it before anyone writes a prompt.

Good candidates share one trait: a stable answer and a clear data source. Think opening hours, order status, price lists, booking and rescheduling, and requests for documents such as invoices or warranty terms. These questions repeat, and the right answer sits in one place.

Bad candidates carry consequences. Anything that sounds like legal, medical or financial advice belongs here, along with refund exceptions, complaints and contract terms that depend on the individual account. The rule is simple. Open-ended advice and anything with legal or financial consequence belongs with a person from the first version.

Write the out-of-scope list down and give it to the bot as explicit instructions. Put it in the test set too, as questions the bot must refuse.

One more line to draw: account data. A bot that reads order status needs a verified customer session and an API call. Document search is the wrong tool there. Keep transactional lookups on scripted paths and free-form answers on RAG. Most working support bots combine the two.

Is your knowledge base ready?

Run this checklist before you choose a platform, because every "no" becomes a wrong answer later.

#CheckWhy it matters
1One source of truth per topicTwo returns pages means two possible answers
2Every document has a "last reviewed" dateThe bot can't tell stale text from current text
3Every document has a named ownerSomeone must fix it when the bot quotes it wrong
4FAQ and policy pages agreeThe search may pull the FAQ, while the policy governs
5Prices, fees and deadlines live in a table or structured fieldNumbers buried in prose get misread or mixed
6Internal and confidential documents are kept out of the indexWhatever is indexed can end up in an answer
7Region-, plan- or product-specific rules are labeled"Free shipping" in one state is not free in another
8Each page covers one question or one topicMixed pages produce mixed chunks
9Headings describe the contentChunks often carry the heading as context
10No answers live only in images or PDFs scanned as picturesText inside images is usually not searchable
11Retired products and old policies are removed or markedOld text keeps winning searches
12A review schedule exists, with a date on the calendarThe bot degrades as the business changes

Expect to spend real time here. The fixes are editorial work: merging duplicates, dating pages, deleting old ones. It is dull. It also improves the help center for human readers.

Test before launch: a question set

A demo proves little. The real exam is 30 to 50 questions pulled from your ticket history, plus a separate block the bot must refuse.

Use this template. One row per question.

Customer questionExpected answerSource documentMust cite?Hand off to human?Pass/fail
"Can I return a sale item?"Yes within 14 days, store credit onlyReturns policy, section 3YesNo
"Where is order 48213?"Status from the order system after loginOrder API (scripted path)NoNo
"Can you waive the restocking fee?""I can't decide that — connecting you to an agent"Out-of-scope listNoYes
"Is this product safe for my pregnancy?""I don't know" + handoffOut-of-scope listNoYes
"What's your price for the Pro plan in Canada?"Price from the pricing table, with linkPricing tableYesNo

How to build it:

  • Pull questions from tickets exactly as customers wrote them. Typos included.
  • Cover your top ticket categories in proportion to their volume.
  • Add 8 to 10 "must say I don't know" questions: topics you don't cover, competitor questions, advice requests.
  • Add 3 to 5 questions where two documents disagree. See which one wins.
  • Write the expected answer before you run the test. Otherwise you'll grade generously.

Track three numbers per run:

  • Answers with a correct source. A right answer citing the wrong page is a failure waiting to happen.
  • Correct "I don't know" rate. Out of the must-refuse block, how many did it refuse?
  • Escalation rate. Too low can mean the bot is bluffing. Too high means your documents have gaps.

Rerun the full set after every change to documents, prompts or models, and keep the results. When someone asks why the bot said something odd, you'll have a baseline.

NIST's Generative AI Profile (NIST AI 600-1, July 26, 2024) lists confabulation — confident, false output — among the risks of generative AI and recommends pre-deployment testing (NIST, as of September 30, 2026). It builds on the voluntary AI Risk Management Framework 1.0, released January 26, 2023 (NIST AI RMF, as of September 30, 2026). Neither is a law. Both are good checklists for your own test plan.

When the bot is wrong: liability and handoff

Wrong answers will happen. Plan for them.

The best-known case is Canadian. In Moffatt v. Air Canada, 2024 BCCRT 149, decided February 14, 2024, British Columbia's Civil Resolution Tribunal heard a claim from a customer whose airline chatbot said he could apply for a bereavement fare after travel. The airline's policy said otherwise. Air Canada argued, in effect, that the chatbot was a separate legal entity. The tribunal called that "a remarkable submission" and wrote that the airline "is responsible for all the information on its website," whether it comes "from a static page or a chatbot." It ordered Air Canada to pay $812.02 in total, including $650.88 in damages (Moffatt v. Air Canada, 2024 BCCRT 149, paras. 27 and 44, as of September 30, 2026).

That is a small-claims decision in one Canadian province. It says nothing about US law, but the practical lesson travels anyway: your customer will treat the bot's answer as your answer. This article is not legal advice.

So design the exit first. The bot needs clear triggers to stop and pass the conversation to a person.

TriggerWhere it goesWhat goes with it
Customer asks for a humanLive agent queue, or ticket if after hoursFull transcript, customer ID, detected topic
No passage above the relevance thresholdAgent queue with "no source found" tagQuestion, top passages that were rejected
Two failed answers in a row ("that's not what I asked")Agent queue, priority bumpTranscript, both bot answers
Refund, cancellation or complaint keywordsBilling or retention teamOrder number, account status, transcript
Legal, medical, safety or financial advice requestHuman specialist or a fixed safe reply + ticketTranscript, flag for review
Angry language or threatsSenior agentTranscript, sentiment flag
Outside working hoursTicket with promised response timeEverything above, plus contact preference

The last column is where most bots fail. Customers forgive a bot that can't answer. They don't forgive repeating their order number to the agent it hands them to. Pass the transcript along with whatever the bot already identified about the customer and the issue.

Then read the logs and review unanswered questions weekly for the first month. Each one points to a missing or unclear document. Fix the document, add the question to the test set, rerun.

Build on a platform or custom

You have two broad routes, and both use the same retrieval idea.

Built-in bots in your helpdesk. Most major helpdesk and chat platforms now ship an AI agent that reads your help center. Setup is fast. The bot sits where your agents already work, and handoff inside the same tool is usually smooth. The limits show up later: you get the vendor's retrieval settings, the vendor's model choices and the vendor's reporting. Pricing is often per resolution or per conversation, so check the vendor's pricing page and model the cost at your real ticket volume.

Your own RAG setup. You pick the model, the search, the chunking and the logging. You can pull from sources the helpdesk can't see: product databases, internal wikis with a filtered export, order systems through an API. You own the test set, the transcripts and the prompts, and the cost is engineering time up front and someone to run it after launch.

A rough way to choose:

  • Help center under a few hundred articles, one channel, standard questions: start with the built-in bot. Use the checklist and test set anyway.
  • Answers depend on data outside the help center, or you need strict control over what gets indexed: a custom setup pays for itself in control.
  • High volume with per-resolution pricing: run the numbers for year two, not month one.

Support is rarely the only process worth handing to software. If you are mapping the rest, see other support workflows worth automating before you commit budget to a single bot.

FAQ

Do I need to fine-tune a model for a support bot?

No, for most support use cases retrieval is enough. Fine-tuning changes how a model writes. It is a poor way to teach facts that change every quarter, like prices or policies. With retrieval, you update a document and the next answer reflects it. Fine-tuning can help with tone or format later, once the retrieval side works.

How long does it take to prepare the knowledge base?

It depends on how many documents disagree with each other. A tidy help center with dated, owned pages may need a week of review. A help center with years of duplicates, old PDFs and conflicting FAQ pages can take much longer. Run the 12-point checklist first. The number of "no" answers gives you a realistic estimate.

Can a RAG bot use documents in several languages?

Yes, many embedding models handle multiple languages, and a question in Spanish can match an English passage. Quality varies by model and language pair. Test it. Add questions in each language to your test set, and check that answers cite the right regional policy. Label region-specific documents clearly so the search does not mix them.

What should the bot say when it doesn't know?

It should say so in one sentence and offer the next step. For example: "I don't have that information. I can connect you with an agent, or you can leave your email for a reply within one business day." Avoid guessing, apologizing at length or repeating the question. A clean refusal plus a working handoff keeps trust.

Is customer data sent to the model provider?

Usually yes, at least the question and the retrieved passages. Check each provider's data retention and training terms before launch. Keep confidential documents out of the index, and mask card numbers and similar fields before they reach the model. For account lookups, prefer a scripted API call that returns only the needed field.

Top articles ⭐

All categories
Best Web Development Companies in Denver (2026)
Denver’s web development teams offer the best of both worlds: West Coast creativity and Midwest dependability. They’re close enough to Silicon Valley to stay ahead on frameworks and tools, yet grounded enough to prioritize results over hype. Artyom Dovgopol Denver’s web dev scene surprised me. No buzzword rush — just…
October 31, 2025
13 min
923
All categories
Website design for conversion growth: key elements
Your website is a complex ecosystem of interconnected elements, each of which affects how users perceive you, your product, and brand. Let's take a closer look at what elements make websites successful and how to make them work for you. Artyom Dovgopol Web design is not art for art’s sake,…
May 30, 2025
11 min
0
All categories
User account development for business growth
A personal website account is that little island of personalization that can make users feel right at home. Want to know more about how personal accounts can benefit your business? We’ve gathered everything you need in this article – enjoy! Artyom Dovgopol A personal account is your user’s map to…
May 28, 2025
15 min
0
All categories
Website redesign strategy guide
The market is constantly shifting these days, with trends coming and going and consumer tastes in a state of constant flux. That’s not necessarily a bad thing — in fact, it’s one more reason to keep your product and your website up to date. In this article, we’ll walk you…
May 26, 2025
13 min
0
All categories
Rebranding: renewal strategy without losing customers
Market success requires adaptation. Whether prompted by economic crisis, climate change, or geopolitical shifts, we'll explain when rebranding is necessary and how to implement it strategically for optimal results. Artyom Dovgopol A successful rebrand doesn’t erase your story; it refines the way it’s told. Key takeaways 👌 Rebranding is a…
April 23, 2025
13 min
0
All categories
Website development cost 2026: pricing and factors
We've all heard about million-dollar websites and "$500 student specials". Let's see what web development really costs in 2026 and what drives those prices. Artyom Dovgopol Know what websites and cars have in common? You can buy a Toyota or a Mercedes. Both will get you there, but the comfort,…
January 23, 2025
6 min
0
Your application has been sent!

We will contact you soon to discuss the project

Close