Knowledge & content

Scanning your website

The website scan is the fastest way to give a new agent something to work with. It runs during onboarding, and you can re-run it any time from Sources.

What the scan does

HelpShelf fetches your homepage and sitemap and extracts:

  • Product context — what your product is and what it does, written into your agent's summary
  • Brand colours — sampled from your site so the widget and help center match it
  • Detected tools — the analytics, chat, and support tools you already run
  • A sitemap snapshot — the shape of your site

The scan runs within a time budget of up to 30 seconds. It's deliberately shallow: it's a fast first pass to get you useful, not an exhaustive archive.

What it doesn't do

The scan reads a sample of your site, not every page. It won't find content behind a login, and it won't pull in your entire blog archive.

For a complete knowledge base, follow the scan with one of:

Re-running the scan

Re-scan after a significant site change — new pricing, a rebrand, a major feature launch. Existing content isn't duplicated; pages already imported are recognised and skipped.

Re-scanning does not overwrite articles you've written or edited by hand.

Review what it found

Scanned pages land in the standard trust tier, which means they rank below anything you write yourself. That's intentional — marketing copy is rarely the best answer to a support question.

After a scan, it's worth skimming Content and unpublishing anything that shouldn't inform answers: press releases, job listings, legal boilerplate. Fewer, better pages beat more pages.