Zum Inhalt springen
DeutschlandGPT

Knowledge sourcesBeta

Let the bot answer from your website and from documents you upload

This feature is in beta. It is usable, but details can still change.

A widget answers best from your own content. It has two sources: your website, which the setup assistant reads for you, and documents you upload. Both live in the Knowledge section.

How the bot uses its knowledge

With Knowledge retrieval (RAG) switched on, the bot searches the knowledge base before it answers and grounds its reply in what it finds. Visitors see this as a short "Searching knowledge base" step above the answer, with the number of sources it found.

Switching retrieval on or off applies immediately, without a publish. The knowledge base belongs to this one widget; no other widget and no other product can read it.

Upload documents

Upload PDFs, Word files and other documents in the Knowledge section. The section shows how many documents fit and how large each may be. Every upload is processed in the background and moves through these states:

StatusMeaning
Queuedwaiting to be processed
Extractingthe text is being read from the file
Processingthe text is being prepared for search
Readythe bot can answer from it
No text foundthe file contains no readable text, for example a scan without OCR
Failedprocessing did not work; delete the document and upload it again

Uploading is the right way for PDFs your website links to, such as forms and statutes: the website crawler skips PDFs.

Your website as a source

You do not have to upload anything to get started. Tell the setup assistant your website address; it reads the site, shows you what it found, and takes in the pages you pick.

In five steps

  1. Create the widget. Open Chatbot Widgets → New widget. The setup assistant opens on the right.
  2. Name the website. bocholt.de is enough, https://www.bocholt.de works just as well.
  3. See what it finds. It fetches the homepage, robots.txt and the sitemap if there is one, then shows you the pages grouped by section ("Citizen services: 41 pages", "News: 120 pages"). Nothing is indexed at this point. About half of all sites have no sitemap; the assistant then follows the links from the homepage instead.
  4. Pick, then let it index. Tell it which sections matter, for example "citizen services and waste yes, press releases no". It reads those pages and puts them in the knowledge base, and tells you how many of the site's pages are covered.
  5. Try it. Ask a question in the test widget whose answer is on one of those pages.

The Knowledge section shows the same website source with every page it took in, and lets you remove it again.

What can be read

The crawler reads ordinary web pages, the way a visitor without JavaScript would see them.

ReadSkipped
HTML pages on your sitePDFs, Word files, images, video
publicly reachable pagesanything behind a login, paywall or form
text, headings, lists, tablescontent that only appears once JavaScript runs
pages on the domain you namedlinks to other domains, even if your sitemap lists them

When a page cannot be read, the assistant says why instead of skipping it silently. The reasons and what to do about them are in Troubleshooting.

How many pages

StatePage ceiling
domain not verified150
domain verified2,500

The ceiling is a cost control. Every page taken in is processed, and that spends credits before the widget has answered a single visitor. Reading a 40,000-page site by accident has to be impossible.

For most sites 150 pages is plenty: across 59 municipal websites we measured, the typical sitemap holds about 100 URLs. The selection matters more than the number anyway. Thirty good pages answer more questions than 2,000 press releases spanning ten years. To raise the ceiling, verify your domain.

Keep it current

A page that has been taken in is a snapshot. When your website changes, the knowledge base does not follow on its own.

Ask the assistant to read the pages again when something substantive changed: new opening hours, new fees, a rebuilt section. Pages whose content did not change are recognised and cost almost nothing to re-read. For a page that changes daily, a knowledge base is the wrong tool.

What indexing costs

Taking pages and documents in is billed to your workspace credits, once, when they are processed. It is not charged per visitor question. A typical site of a few hundred pages costs a few euros. The ongoing cost comes later, from the answers; see Limits and billing.

Was this page helpful?