User Guide

Data Sources

Connect, configure, and manage the data sources that feed your search index.

Adding a data source#

  1. Click "Add Data Source" on the Data Sources page.
  2. Select the connector type (SharePoint, Confluence, Web Crawler, etc.).
  3. Enter the required configuration (URLs, credentials, etc.).
  4. Test the connection to verify it works.
  5. Save and trigger the first crawl.

Supported connectors#

ConnectorAuthKey features
SharePointAzure AD appFull crawl, delta sync, permission sync, multi-site
ConfluenceAPI tokenSpaces, pages, blog posts, attachments, labels
DatabricksAccess tokenColumn mapping, Unity Catalog
Genesys CloudOAuth2Call transcripts, speaker labels, delta sync
Web CrawlerNoneMulti-site BFS, depth/page limits, robots.txt
File SharesUsername/passwordFTP, FTPS, SMB — recursive traversal, file type filtering
Cloud StorageAccess key / service accountAzure Blob, S3, GCS — unified interface
SQL DatabaseConnection stringPostgreSQL, MySQL, SQL Server — configurable queries
SlackBot tokenComing soon — channels, threads, files
Google DriveService accountComing soon — files, folders, shared drives

Crawling#

Crawls fetch documents from source systems and index them. The first crawl is a full sync. Subsequent crawls use delta detection to only process changes (content hashes and timestamps).

Error recovery#

If individual items fail during a crawl, they're tracked separately with error category and retry count. View error details and retry failed items without re-crawling everything.

Auto-sync scheduling#

Each data source has an auto-sync interval selector on its card:

  • Every hour — for rapidly-changing sources.
  • Every 6 hours — the default, good for most content.
  • Every 12 hours / Daily / Every 2 days — for stable content like documentation.
Note
Deleting a data source also removes all indexed documents from that source.