A business AI chatbot is a conversational assistant that answers in plain language from your own content: documentation, product sheets, procedures, terms and conditions, a knowledge base. Done well, it answers accurately, cites its sources and knows how to say “I don’t know”. Done badly, it makes things up. The difference lies entirely in the architecture — and above all in one technique: RAG.

Scripted chatbot or AI chatbot?

The first chatbots followed decision trees: buttons, keywords, pre-written answers. They are still useful for very simple journeys (booking an appointment, tracking an order) but fail the moment a question goes off script.

An AI chatbot is built on a large language model. It understands freely worded questions, rephrases, summarises several documents and replies in the user’s language. Its natural weakness: a model on its own knows nothing about your business and may produce an answer that sounds right but is wrong. Hence the need to anchor it in your data.

RAG: making AI answer from your documents

RAG (retrieval-augmented generation) means searching your documents for the relevant passages before asking the model to write the answer. It works in three stages.

  1. Indexing. Your documents are split into passages, then turned into numerical representations (embeddings) that capture their meaning. These are stored in a search index.
  2. Retrieval. For each question, the system finds the passages closest in meaning, often combining semantic search with keyword search.
  3. Generation. The model receives the question and the retrieved passages, with a clear instruction: answer only from these sources, cite them, and say when the information is missing.

The result: answers that stay current (just update the documents), that can be checked (each answer points to its source) and that stay within what you have approved.

What makes a good RAG system

  • Source quality. A chatbot will never be better than the documents it is given. Contradictory or outdated procedures produce contradictory or outdated answers.
  • Chunking. Passages that are too short lose context; too long, and the key information is drowned out. Splitting should follow the documents’ own structure (headings, clauses, sections).
  • Hybrid search. References, product codes and proper names are best found by keyword; free-form questions, by meaning. Combining both noticeably improves relevance.
  • Metadata. Date, language, audience, confidentiality level: these let sources be filtered for each user.
  • Evaluation. A set of reference questions with expected answers lets you measure quality before going live and after every change.

Writing content that a chatbot can use

Documents written for people are not always easy for a retrieval system to use. A few habits make a real difference: one topic per section, with a heading that names it; explicit dates and versions; terms defined once and used consistently; tables rather than long paragraphs for conditions and lead times; and no critical information hidden in images or scanned PDFs. These habits also make the documents better for your staff.

Where to deploy an AI chatbot

On your website, it answers visitors around the clock — questions about your services, terms and lead times — and hands over to a person when a request becomes commercial or sensitive.

Internally, it becomes the team’s assistant: finding a procedure, a contract clause, a technical sheet, the right answer to a routine HR question. This is often where the time saved is greatest.

Inside a business application, it sits on the screen where people work: contextual help, searching a file’s history, assisted drafting.

Security and compliance: the rules to follow

A chatbot handles information and talks to people, so several rules apply.

  • Transparency. The EU AI Act requires people to be told they are interacting with an AI. The chatbot must say so clearly.
  • Access rights. An internal chatbot must never reveal a document to an employee who has no access to it. Rights are checked at retrieval time, not just on display.
  • Personal data. Conversations may contain personal data: a defined retention period, information for users, a provider bound by contract.
  • Resistance to misuse. Some users will try to push the chatbot out of its role. A solid system prompt, a scope limited to approved sources and regular testing reduce that risk.
  • Choice of provider. Prefer business offerings whose terms exclude using your data to train models. Our article on AI, the GDPR and the AI Act explains these obligations.

Launching an AI chatbot in five steps

  1. Define the scope. Which audience, which questions, which sources — and, above all, which questions the chatbot must not handle.
  2. Prepare the content. Gather, de-duplicate and update the documents. This is the most underestimated step, and the most decisive.
  3. Build and evaluate. Indexing, retrieval, instructions, then testing against real questions until answers are reliable and sourced.
  4. Integrate. On the website, intranet or application, with a hand-over to a person and, where useful, a link to your tools (CRM, booking).
  5. Monitor and improve. Review unanswered questions, fill gaps in the sources, refine the instructions. A chatbot improves mainly through its content.

Measuring a chatbot’s quality

A chatbot is steered with a handful of simple indicators, tracked from the pilot onwards:

Indicator What it measures How to improve it
Sourced answer rate Share of answers backed by a document Add sources, tighten the instructions
Unanswered questions Topics your content does not cover Write the missing documents
Hand-overs to a person Requests the chatbot should not handle Adjust the scope and the hand-over
User satisfaction Perceived usefulness of answers Review conversations rated poorly
Accuracy on the test set Quality on reference questions Revisit chunking and retrieval

Read them together: a chatbot that answers everything but rarely cites its sources is riskier than one that admits its limits.

From chatbot to agent

A chatbot answers; an AI agent acts. The line blurs as soon as the assistant can, say, open a ticket in your support tool, check an order’s status or offer an appointment slot. Those actions rely on connecting AI to your software, with the same guardrails as an agent: limited permissions and approval of binding actions.

Frequently asked questions

Can an AI chatbot make up answers?

A language model on its own, yes. With well-designed RAG, an instruction to answer only from the sources and an obligation to cite them, that risk falls sharply — and the chatbot must be able to say it cannot find the information.

How many documents do you need to start?

There is no minimum: a chatbot can start with a handful of well-written reference documents. A few reliable sources beat many contradictory ones.

Can the chatbot answer in several languages?

Yes. Current models understand and write most languages, even when your documents are in one language only. It is still wise to have standard answers reviewed in the most-used languages.

Do we need to host the model ourselves?

Rarely. Major providers’ business offerings, used through an API with contractual commitments, suit most needs. Hosting an open model makes sense only for very strict confidentiality requirements.

The Step By Step studio in Ajaccio builds AI chatbots and assistants that answer from your content and plug into your tools. Let’s talk about your project.