Knowledge sourcesBeta
Let the bot answer from your website and from documents you upload
This feature is in beta. It is usable, but details can still change.
A widget answers best from your own content. It has two sources: your website, which the setup assistant reads for you, and documents you upload. Both live in the Knowledge section.
How the bot uses its knowledge
With Knowledge retrieval (RAG) switched on, the bot searches the knowledge base before it answers and grounds its reply in what it finds. Visitors see this as a short "Searching knowledge base" step above the answer, with the number of sources it found.
Switching retrieval on or off applies immediately, without a publish. The knowledge base belongs to this one widget; no other widget and no other product can read it.
Upload documents
Upload PDFs, Word files and other documents in the Knowledge section. The section shows how many documents fit and how large each may be. Every upload is processed in the background and moves through these states:
| Status | Meaning |
|---|---|
| Queued | waiting to be processed |
| Extracting | the text is being read from the file |
| Processing | the text is being prepared for search |
| Ready | the bot can answer from it |
| No text found | the file contains no readable text, for example a scan without OCR |
| Failed | processing did not work; delete the document and upload it again |
Uploading is the right way for PDFs your website links to, such as forms and statutes: the website crawler skips PDFs.
Your website as a source
You do not have to upload anything to get started. Tell the setup assistant your website address; it reads the site, shows you what it found, and takes in the pages you pick.
In five steps
- Create the widget. Open Chatbot Widgets → New widget. The setup assistant opens on the right.
- Name the website.
bocholt.deis enough,https://www.bocholt.deworks just as well. - See what it finds. It fetches the homepage,
robots.txtand the sitemap if there is one, then shows you the pages grouped by section ("Citizen services: 41 pages", "News: 120 pages"). Nothing is indexed at this point. About half of all sites have no sitemap; the assistant then follows the links from the homepage instead. - Pick, then let it index. Tell it which sections matter, for example "citizen services and waste yes, press releases no". It reads those pages and puts them in the knowledge base, and tells you how many of the site's pages are covered.
- Try it. Ask a question in the test widget whose answer is on one of those pages.
The Knowledge section shows the same website source with every page it took in, and lets you remove it again.
What can be read
The crawler reads ordinary web pages, the way a visitor without JavaScript would see them.
| Read | Skipped |
|---|---|
| HTML pages on your site | PDFs, Word files, images, video |
| publicly reachable pages | anything behind a login, paywall or form |
| text, headings, lists, tables | content that only appears once JavaScript runs |
| pages on the domain you named | links to other domains, even if your sitemap lists them |
When a page cannot be read, the assistant says why instead of skipping it silently. The reasons and what to do about them are in Troubleshooting.
How many pages
| State | Page ceiling |
|---|---|
| domain not verified | 150 |
| domain verified | 2,500 |
The ceiling is a cost control. Every page taken in is processed, and that spends credits before the widget has answered a single visitor. Reading a 40,000-page site by accident has to be impossible.
For most sites 150 pages is plenty: across 59 municipal websites we measured, the typical sitemap holds about 100 URLs. The selection matters more than the number anyway. Thirty good pages answer more questions than 2,000 press releases spanning ten years. To raise the ceiling, verify your domain.
Keep it current
A page that has been taken in is a snapshot. When your website changes, the knowledge base does not follow on its own.
Ask the assistant to read the pages again when something substantive changed: new opening hours, new fees, a rebuilt section. Pages whose content did not change are recognised and cost almost nothing to re-read. For a page that changes daily, a knowledge base is the wrong tool.
What indexing costs
Taking pages and documents in is billed to your workspace credits, once, when they are processed. It is not charged per visitor question. A typical site of a few hundred pages costs a few euros. The ongoing cost comes later, from the answers; see Limits and billing.