Skip to main content
Website Links is where your chatbot learns from pages on the web. Open your chatbot’s sidebar and expand Website Links to see the four ways to add links, and Links List to manage what is already there.
Give SiteGPT your site’s homepage (or any starting URL) and it crawls the site by following links.
  1. Open Website LinksScrape Website.
  2. Enter the starting URL.
  3. Set Max Pages to Scrape. The crawl stops at this many pages or at your remaining pages quota, whichever is lower.
  4. Set Recursion Depth, which is how many clicks away from the starting page the crawler may go.
  5. Optionally restrict the crawl with Allowed Domains.
  6. Start the scrape. Progress appears in Links List as pages are discovered and trained.
Best for: getting broad coverage of a site quickly.

Scrape options

The website, sitemap, and multiple-links flows share these options:

Editing a link’s settings

Open any link in Links List to change its scrape options or sync frequency. Saving re-imports that URL with the new settings; the page count adjusts to the content’s current length. To update the configuration of many links at once, select them in the list and use Update Config. A Failed badge means the page could not be fetched or processed. Open the row to see the error. Common causes:
Bot protection (Cloudflare and similar) can refuse automated requests. If you control the protection, allowlist SiteGPT’s requests or use Custom Headers with a token your firewall accepts.
Pages behind a login cannot be crawled. Export the content and upload it as a file, or paste it as a text snippet.
Check that the URL loads in a private browser window and update the link.
Imports stop when your plan’s pages quota is used up. Free quota by deleting content you do not need, or upgrade. See Pages quota.
After fixing the cause, select the failed links and use the Resync button (it shows the selected count).