Skip to Content
Knowledge Sources

Knowledge Sources

Knowledge sources decide what the agent can answer. Manage them under Knowledge → Links and Knowledge → Documents.

Where content comes from

  • Your site domain — when an agent is created, siteaiagent crawls the site starting from the home page and auto-discovers sitemaps from standard paths and robots.txt.
  • Sitemap links — add sitemap URLs and siteaiagent indexes the pages they list.
  • Page links — add individual page URLs from the Links screen.
  • Uploaded documents — upload .txt or .md files from the Documents screen.

Sign-in, sign-up, login, and registration pages are never indexed.

Keeping content current

Use Update content on the Links screen:

  • Incremental update only fetches newly added links and documents.
  • Full update re-fetches every indexed page to pick up content changes. Pages whose content has not changed are detected by a content hash and skipped.

Each link shows its status while an update runs: Pending, Updating, Failed, or Skipped.

What indexing stores

Expand a link to see the raw indexed content plus Retrieval metadata: a per-chunk summary, extracted facts, suggested questions, and search terms. siteaiagent enriches every indexed page automatically so the agent can find the right passage — it builds search metadata in both English and Chinese, extracts key facts, and prepares the questions each passage can answer. No configuration is needed.

Content hygiene

  • Remove outdated prices, policies, and guarantees.
  • Avoid internal-only content.
  • Prefer short, focused documents over large mixed files.
  • Run an update after important website or policy changes.