Skip to main content
Content is what your chatbot learns from. The CLI calls it knowledge, so all commands on this page start with sitegpt knowledge. Each item of content, such as one web page or one file, is a document. This page shows the common path for each task. For every option, see the command reference. To connect apps such as Google Drive or GitHub, see Connect apps.

Before you start

  • A SiteGPT account on the Growth plan or above. See Plans and limits.
  • The CLI installed and logged in. See Install and log in. The default access from sitegpt login is enough for the tasks on this page.
  • Your chatbot ID. Run sitegpt chatbots list to find it.
  • Enough pages left for new content. Run sitegpt limits to check.

Add a sitemap

A sitemap is the most reliable way to add a website, because it lists the exact URLs.
  • --max-links sets how many URLs to add, from 1 to 1000. The default is 50. It is also capped by the pages you have left.
  • --wait waits until training is done. The default wait is 600 seconds.
  • To keep the sitemap up to date, see Keep a website or sitemap up to date.
Check that it works: The CLI prints Training settled: and the number of trained documents. If some documents failed, see Resync content.

Add a website

If the website has no sitemap, crawl it from a start URL:
  • The crawl follows links up to 3 levels deep by default. Change it with --depth (1 to 5).
  • --max-links works as for a sitemap. The default is 50.
  • --include-path and --exclude-path limit the crawl to some paths. You can repeat them.
Check that it works: The CLI prints Training settled: and the number of trained documents.

Add specific pages

Use this when you know the exact pages that you want:
--skip-existing skips pages that the chatbot already has. The --sync option of links add does not schedule these pages. To refresh them on a schedule, run documents update-config with --sync. See the command reference. To refresh them now, resync them. Check that it works: The CLI prints Training settled: and the number of trained documents.

Choose which part of a page to read

By default, SiteGPT reads only the main content of each page. To choose the parts yourself, add CSS selectors when you add pages:
These options also work with website add and sitemap add. To change them for a document that you already added, use documents update-config:
For the full list of read options, see Knowledge documents. Check that it works: Run sitegpt knowledge documents get --chatbot YOUR_CHATBOT_ID YOUR_DOCUMENT_ID --content. The parsed text has the parts that you want, and not the parts that you excluded.

Add YouTube videos

You can pass video, playlist, and channel URLs. The chatbot does not auto-sync videos. To pick up changes, resync them. Check that it works: Run sitegpt knowledge documents list --chatbot YOUR_CHATBOT_ID --source YOUTUBE. Each video has the status SUCCESS when training is done.

Upload files

Pass local file paths, not URLs. One command takes up to 10 files, up to 20 MB for each file, and up to 50 MB in total. Check that it works: Run sitegpt knowledge documents list --chatbot YOUR_CHATBOT_ID --source LOCAL_FILE. Each file has the status SUCCESS when training is done.

Add or replace the text snippet

Each chatbot has one text snippet. This command creates it, or replaces all of its text:
You can also pass the text directly: text update --chatbot YOUR_CHATBOT_ID "Refund policy: ...". The text can be up to 10,000 characters. Put all the text that you want to keep in each update. See Add a text snippet. Check that it works: The CLI prints Added text document and the document ID. If there was a text snippet before, it also prints Replaced previous text document.

Wait for training to finish

If you did not pass --wait when you added content, wait for training now:
The default timeout is 600 seconds. If training does not finish in time, the command fails with TRAINING_TIMEOUT. Training continues in the background. Check that it works: The CLI prints Training settled: and the number of trained documents.

Find documents

You can filter by text with --query, by source with --source, by status with --status, and by type with --type. For the allowed values, see Knowledge documents. To see what one document contains:
Check that it works: The list shows the ID, name, type, source, and status of each matching document. If nothing matches, the CLI prints No documents found.

Edit a document

Change the text that the chatbot learned from one document:
Check that it works: The CLI prints Document YOUR_DOCUMENT_ID queued for update. and the status.

Resync content

Resync makes the chatbot read content again to pick up changes. Resync one document:
Resync every document that failed:
--state failed selects documents with the status FAILED, CANCELLED, or BLOCKED_BY_QUOTA. You can also select documents with filters, such as --source WEBSITE --status FAILED. Check that it works: The CLI prints Documents queued for resync: with a number. Run sitegpt knowledge wait --chatbot YOUR_CHATBOT_ID to wait for the result.

Delete content

Delete one document:
Delete every document that failed:
To delete all content, use --all --yes. Before a bulk delete, add --dry-run to print the selection that the command would send. A dry run does not contact SiteGPT, so it does not show which documents match. Check that it works: The CLI prints Documents queued for deletion: with a number. The documents no longer show in knowledge documents list when the deletion is done.

Keep a website or sitemap up to date

SiteGPT has two schedules for this:
  • Auto-sync reads the same source again with the saved settings. It works only for website, sitemap, and GitHub sources. It does not delete pages.
  • Scan reads a sitemap again, adds new URLs, and deletes pages whose URLs are no longer in the sitemap. It works only for sitemap sources.
See Keep your content up to date. The easiest way is to set the schedule when you add the source:
website add accepts --sync. The values are DAILY, WEEKLY, MONTHLY, and NEVER. The default is NEVER. To change the schedule of a source that you already added:
1

Find the sync job ID

Each import of a website, sitemap, or GitHub source is a sync job. If the source already has a schedule, list the jobs:
The list shows only jobs that have a schedule. If the source has no schedule yet, run sitegpt knowledge documents get --chatbot YOUR_CHATBOT_ID YOUR_DOCUMENT_ID for one of its documents. The Ingest job part of the output shows the job ID.
2

Set the schedule

Set --sync, --scan, or both to DAILY, WEEKLY, or MONTHLY:
  • If a frequency is not available on your plan, SiteGPT sets it to NEVER and the response includes a warning. See Plans and limits.
  • If sync and scan use the same frequency, only the sync runs on that schedule.
  • --sync on another source type fails with SYNC_FREQUENCY_UNSUPPORTED_SOURCE. --scan on a source that is not a sitemap fails with SCAN_FREQUENCY_UNSUPPORTED_SOURCE.
  • To turn off auto-sync, do not use --sync NEVER. It fails with SYNC_JOB_DELETE_REQUIRED. See Stop auto-sync and scan.
To scan a sitemap now, even when its scan frequency is NEVER:
Check that it works: Run sitegpt knowledge sync-jobs get --chatbot YOUR_CHATBOT_ID YOUR_JOB_ID. The output shows the new Sync frequency and Scan frequency.

Stop auto-sync and scan

This sets both frequencies to NEVER. It does not delete the content that the chatbot already learned. sync-jobs disable does the same. Check that it works: The job no longer shows in sitegpt knowledge sync-jobs list --chatbot YOUR_CHATBOT_ID.

Add a custom response

A custom response is a question and the answer that you write for it. See Custom responses.
Check that it works: The CLI prints Added custom response with the ID and the state UPDATED. Ask the chatbot the question with sitegpt messages send --chatbot YOUR_CHATBOT_ID "What is your refund policy?".

Fix answers that visitors downvoted

1

List downvoted answers that have no custom response yet

2

Write the correct answer

OPEN means that the question has no answer from you yet. UPDATED means that it has one. Check that it works: Run sitegpt knowledge custom-responses get --chatbot YOUR_CHATBOT_ID YOUR_CUSTOM_RESPONSE_ID. The output shows State: UPDATED and your answer.

Delete a custom response

Check that it works: The custom response no longer shows in sitegpt knowledge custom-responses list --chatbot YOUR_CHATBOT_ID.