User Guide
Data Sources
Connect, configure, and manage the data sources that feed your search index.
Adding a data source#
- Click "Add Data Source" on the Data Sources page.
- Select the connector type (SharePoint, Confluence, Web Crawler, etc.).
- Enter the required configuration (URLs, credentials, etc.).
- Test the connection to verify it works.
- Save and trigger the first crawl.
Supported connectors#
| Connector | Auth | Key features |
|---|---|---|
| SharePoint | Azure AD app | Full crawl, delta sync, permission sync, multi-site |
| Confluence | API token | Spaces, pages, blog posts, attachments, labels |
| Databricks | Access token | Column mapping, Unity Catalog |
| Genesys Cloud | OAuth2 | Call transcripts, speaker labels, delta sync |
| Web Crawler | None | Multi-site BFS, depth/page limits, robots.txt |
| File Shares | Username/password | FTP, FTPS, SMB — recursive traversal, file type filtering |
| Cloud Storage | Access key / service account | Azure Blob, S3, GCS — unified interface |
| SQL Database | Connection string | PostgreSQL, MySQL, SQL Server — configurable queries |
| Slack | Bot token | Coming soon — channels, threads, files |
| Google Drive | Service account | Coming soon — files, folders, shared drives |
Crawling#
Crawls fetch documents from source systems and index them. The first crawl is a full sync. Subsequent crawls use delta detection to only process changes (content hashes and timestamps).
Error recovery#
If individual items fail during a crawl, they're tracked separately with error category and retry count. View error details and retry failed items without re-crawling everything.
Auto-sync scheduling#
Each data source has an auto-sync interval selector on its card:
- Every hour — for rapidly-changing sources.
- Every 6 hours — the default, good for most content.
- Every 12 hours / Daily / Every 2 days — for stable content like documentation.
Note
Deleting a data source also removes all indexed documents from that source.