User Guide

Search pipeline

Query reformulation, document expansion, and contextual chunking — tuning the retrieval pipeline.

Query reformulation#

Query reformulation rewrites your query before retrieval to bridge vocabulary gaps between how users phrase questions and how answers are written in documents.

StrategyHow it worksBest for
HyDEGenerates a hypothetical answer, then searches for documents similar to itFactual questions with predictable answer phrasing
Multi-QueryGenerates 3 alternative phrasings and merges resultsAmbiguous or domain-specific queries
Step-BackGenerates a broader version of the query for high-level contextNarrow questions that need broader context
Tip
Start with Multi-Query — it's the most generally useful strategy. Each reformulation adds one LLM call to the search pipeline.

Document expansion#

Document Expansion improves answer quality for queries targeting a specific document (policy number, ticket ID, case number). When 3 or more of the top 5 results come from the same document, the system pulls in all chunks from that document so the AI sees the full picture.

  • Enable/Disable — off by default. Turn on when your content includes document-specific queries.
  • Max Expanded Chunks — limits how many chunks are pulled per document (default 50).
When to use
Document expansion works best with document-specific queries like "What are the exclusions for policy ABC-12345?" or "Summarize ticket HELP-4521". For broad topic searches, standard chunk retrieval is usually sufficient.

Contextual chunking#

Contextual chunking prepends document-level context to each chunk before embedding. Without context, a chunk like "The deadline is March 15" is ambiguous — with context, the system knows it's about "Q1 budget review deadlines from the Finance team".

ModeHow it worksCost
Rule-BasedPrepends document title, source name, and metadata to each chunkFree — no API calls
LLMUses LLM to generate a contextual summary for each chunkOne LLM call per chunk
Tip
Rule-Based mode gives 80% of the benefit at zero cost. Use LLM mode only for content where chunks are frequently ambiguous without additional context.
Note
After changing the contextual chunking mode, re-crawl your data sources to apply the new strategy to existing documents.

Hybrid search balance#

The hybrid search balance controls the weight between semantic (vector) and keyword (full-text) search. A value between 0.0 (pure keyword) and 1.0 (pure semantic). The default 0.70 works well for mixed content.

  • Lower values (0.30–0.50) — when exact terms matter (product codes, legal clauses).
  • Higher values (0.70–0.90) — when meaning matters more than wording.