robots.txt and domain verificationBeta
What the widget crawler respects, and how verifying your domain raises the page limit and lets you override robots.txt
This feature is in beta. It is usable, but details can still change.
The crawler that reads your website for the knowledge base behaves like a polite guest. Verifying that the site is yours lets it do more.
We honour your robots.txt
Our crawler identifies itself as DeutschlandGPT-Widget-Bot and obeys your site's robots.txt. It fetches one page at a time, half a second apart, so a widget crawl never looks like a load test to a small server.
When pages are skipped because of robots.txt, the assistant says so and names the reason. That is often a surprise: plenty of sites block crawlers wholesale without anyone having decided to. Many municipal sites block AI crawlers by name, a rule aimed at strangers collecting training data rather than at you reading your own site. You have two options:
- Adjust your site's
robots.txt. It is the clean fix and applies to every crawler. - Verify your domain and override
robots.txtfor this bot specifically. Only possible for sites you can prove you own.
When robots.txt is too large to read
We read at most 512 KB of a robots.txt and at most 1,000 rules per user-agent group. If the part that applies to us is cut off, we do not know what the missing rules forbid, so we skip those pages rather than crawl against half a rule set. The assistant reports this before any crawl. It is our reading limit, not a block by your site. Verifying your domain turns robots.txt off for that source, after which the limit no longer matters.
Verify your domain
Two things depend on it: the page ceiling, which rises from 150 to 2,500, and the ability to override robots.txt. There are two ways and you only need one.
Option 1: a meta tag on the homepage. Your web team adds one line to the homepage's <head>:
<meta name="dgpt-site-verification" content="YOUR-TOKEN" />
Option 2: a DNS record. Whoever manages the domain adds a TXT record:
TXT dgpt-site-verification=YOUR-TOKEN
Ask the assistant to verify the domain and it gives you the token, which belongs to that one website source. Then tell it to check; it looks and unlocks the source. The meta tag is usually faster because any web editor can add it. The DNS record is the cleaner answer but often means a ticket to IT.
Already verified for sign-in
If your organisation already verified this domain for single sign-on, that counts here too and there is nothing further to do. It does not work the other way round: a website source verified here unlocks no sign-in features.
Verification is checked on every crawl
Before every crawl the verification is checked again. If the domain loses it, because the record was removed or the address changed, robots.txt is honoured again and the assistant tells you why the behaviour changed.