User Guide

Document Processing (OCR)

Extract text from scanned PDFs and image files using optical character recognition.

How it works#

The system uses a smart fallback approach:

  1. Library-based extraction runs first (fast, free — works for native PDFs, DOCX, PPTX).
  2. If extracted text is below the minimum threshold (default 50 characters), OCR triggers automatically.
  3. Image files (PNG, JPG, TIFF, BMP, WebP) are sent directly to OCR.
  4. If OCR fails, the system falls back to whatever text the library extracted — crawls are never blocked.

OCR providers#

ProviderBest forConfiguration
Azure OpenAI VisionAlready using Azure OpenAINone — reuses AI Assistant credentials
Azure Document IntelligenceHigh volume, lowest per-page costSeparate endpoint and API key
Ollama VisionOn-premises, air-gappedRequires a vision model pulled on Ollama

Setup steps#

  1. Go to Settings → Document Processing.
  2. Toggle Enable OCR on.
  3. Select your preferred OCR provider.
  4. If using Azure Document Intelligence, enter the endpoint URL and API key.
  5. If using Ollama Vision, verify the model name and ensure it's pulled.
  6. Click Save, then Test Connection to verify the provider works.
  7. Re-crawl data sources to process existing scanned documents.

Advanced settings#

  • Max Pages per Document — limits how many pages are sent to OCR (default 50). Prevents runaway costs on large documents.
  • Min Text Threshold — character count below which a PDF is considered "scanned" and triggers OCR (default 50). Set to 0 to always run OCR.
Tip
After enabling OCR, re-crawl existing data sources to process previously skipped scanned documents and image files.