User Guide
Document Processing (OCR)
Extract text from scanned PDFs and image files using optical character recognition.
How it works#
The system uses a smart fallback approach:
- Library-based extraction runs first (fast, free — works for native PDFs, DOCX, PPTX).
- If extracted text is below the minimum threshold (default 50 characters), OCR triggers automatically.
- Image files (PNG, JPG, TIFF, BMP, WebP) are sent directly to OCR.
- If OCR fails, the system falls back to whatever text the library extracted — crawls are never blocked.
OCR providers#
| Provider | Best for | Configuration |
|---|---|---|
| Azure OpenAI Vision | Already using Azure OpenAI | None — reuses AI Assistant credentials |
| Azure Document Intelligence | High volume, lowest per-page cost | Separate endpoint and API key |
| Ollama Vision | On-premises, air-gapped | Requires a vision model pulled on Ollama |
Setup steps#
- Go to Settings → Document Processing.
- Toggle Enable OCR on.
- Select your preferred OCR provider.
- If using Azure Document Intelligence, enter the endpoint URL and API key.
- If using Ollama Vision, verify the model name and ensure it's pulled.
- Click Save, then Test Connection to verify the provider works.
- Re-crawl data sources to process existing scanned documents.
Advanced settings#
- Max Pages per Document — limits how many pages are sent to OCR (default 50). Prevents runaway costs on large documents.
- Min Text Threshold — character count below which a PDF is considered "scanned" and triggers OCR (default 50). Set to 0 to always run OCR.
Tip
After enabling OCR, re-crawl existing data sources to process previously skipped scanned documents and image files.