Hybrid Search

FTS + semantic + RRF fusion search engine.

Search Guide

This guide explains how to use Fortémi's hybrid search system effectively.

When to Use Which Mode

Your QueryRecommended ModeWhy
Known keywords, code snippets, exact phrasesFTSLexical matching finds precise terms
Conceptual questions, "how to..."SemanticEmbedding similarity captures meaning
General search, don't know what to expectHybrid (default)Best of both worlds via RRF fusion
CJK characters, emojiHybrid with language hintCombines bigram/trigram with semantic
Large result set, need precisionFTS with strict filtersFastest, most precise for known terms

Quick rule: If unsure, use hybrid (the default). Switch to FTS for precision or semantic for discovery.

Search Modes

Fortémi offers three search modes, each optimized for different use cases:

1. Hybrid Search (Default)

Combines lexical (BM25) and semantic (dense retrieval) results using Reciprocal Rank Fusion.

curl "http://localhost:3000/api/v1/search?q=retrieval+augmented+generation"

Best for: Most queries. Finds both exact matches and semantically related content.

2. Lexical Search (FTS)

Pure keyword matching using BM25 ranking via PostgreSQL full-text search.

curl "http://localhost:3000/api/v1/search?q=API+documentation&mode=fts"

Best for: Finding exact phrases, code snippets, or when you know the precise terminology.

Pure embedding similarity using dense retrieval.

curl "http://localhost:3000/api/v1/search?q=how+to+build+neural+networks&mode=semantic"

Best for: Conceptual queries, finding related content with different terminology.

Query Embedding Provider

Unscoped semantic and hybrid searches use the active embedding route from the provider registry (`MATRIC_EMBEDDING_PROVIDER`, or the default inference provider). The registry supplies the endpoint and credentials as well as the configured model and dimension.

When `set=<slug>` is present, Fortémi first resolves that set's embedding configuration. The query uses the set's provider, model, and effective dimension, including its MRL truncation dimension. OpenAI-compatible set profiles can select a registered compatible route with `provider_config.provider_id`; endpoints and credentials still come from the provider registry and are never accepted from the search request.

Without `set`, search and background indexing both resolve the archive's default embedding set/config before falling back to the global registry default. Both paths derive the same non-secret contract fingerprint from provider, model, dimension, normalization policy, and embedding-set identity.

The returned query vector must match the effective configured dimension. Provider construction failures, request failures, empty responses, and dimension mismatches fall back to FTS so search remains available, but the fallback is explicit:

{
  "results": [],
  "query": "how to build neural networks",
  "total": 0,
  "degraded": true,
  "degradation": {
    "code": "embedding_request_failed",
    "effective_mode": "fts"
  }
}

Successful semantic and hybrid responses contain `"degraded": false` and omit `degradation`. Degraded responses emit a redacted `search.embedding_degraded` audit/telemetry event and are never written to the search cache.

Understanding Results

Score Interpretation

Score RangeMeaning
0.8 - 1.0Highly relevant
0.6 - 0.8Moderately relevant
0.4 - 0.6Somewhat relevant
< 0.4Tangentially related

RRF Fusion

In hybrid mode, results are ranked using Reciprocal Rank Fusion:

score(d) = 1/(20 + rank_fts) + 1/(20 + rank_semantic)

Documents appearing high in both rankings score best. The k=20 constant (optimized from the original k=60 based on Elasticsearch BEIR benchmark analysis) emphasizes top-ranked results while preventing any single ranking from dominating.

Advanced Search Features

Adaptive RRF

The RRF k parameter automatically adapts based on query characteristics:

  • Short queries (1-2 tokens): k *= 0.7 (tighter fusion, more emphasis on top results)
  • Long queries (6+ tokens): k *= 1.3 (looser fusion, considers more results)
  • Quoted queries: k *= 0.6 (precision-focused, exact match emphasis)
  • Default: k=20 for balanced queries

This adaptive approach improves relevance by tailoring the fusion algorithm to query type.

Adaptive Weights

FTS and semantic weights automatically adjust based on query characteristics:

Query TypeFTS WeightSemantic WeightWhen to Use
Quoted phrases0.70.3"machine learning" - Exact phrase matching
Keywords (1-2 tokens)0.60.4rust, API - Short keyword queries
Balanced (3-5 tokens)0.50.5rust async programming - Medium queries
Conceptual (6+ tokens)0.350.65how do I implement semantic search - Natural language

Why this matters: Keyword queries benefit from lexical precision, while conceptual queries benefit from semantic understanding. The system automatically chooses the best balance.

Relative Score Fusion (RSF)

Alternative fusion algorithm that preserves score magnitude:

normalized_score = (score - min) / (max - min)
final_score = w_fts * norm_fts + w_sem * norm_sem

Differences from RRF:

  • RRF uses only rank position (1st, 2nd, 3rd...)
  • RSF preserves actual score values
  • RSF better captures large score differences
  • Weaviate reports +6% recall on FIQA benchmark vs RRF

When to use RSF: When score magnitudes matter (e.g., large quality gaps between results).

Result Deduplication

When documents are chunked for embedding, multiple chunks from the same document may appear in results. The system automatically:

1. Groups chunks by document ID 2. Keeps the best-scoring chunk per document 3. Adds metadata showing how many chunks matched 4. Re-sorts results after deduplication

Example response with chain info:

{
  "note_id": "uuid",
  "score": 0.85,
  "snippet": "...matching text...",
  "title": "Original Document Title",
  "chain_info": {
    "chain_id": "uuid",
    "original_title": "Original Document Title",
    "chunks_matched": 3,
    "best_chunk_sequence": 2,
    "total_chunks": 5
  }
}

This ensures clean results without duplicate entries for the same document.

Filtering

By Tags (SKOS Concepts)

Filter results to specific categories:

curl "http://localhost:3000/api/v1/search?q=machine+learning&tags=research,ai"

Strict vs. Soft Filtering

ModeBehaviorUse Case
Strict100% isolation, applied before searchMulti-tenancy, access control
SoftCombined with relevance scoringPreference-based filtering

Date Ranges

curl "http://localhost:3000/api/v1/search?q=meeting+notes&created_after=2024-01-01"

Query Tips

Natural Language Works

The semantic component understands natural language:

  • "how do I connect to the database" works as well as "database connection"
  • "problems with authentication" finds notes about "auth errors" or "login issues"

Combining Approaches

Start with hybrid search. If results are too broad, switch to FTS for precision. If missing related content, try semantic mode.

Phrase Matching

Use quotes for exact phrases in FTS mode:

curl "http://localhost:3000/api/v1/search?q=\"retrieval+augmented+generation\"&mode=fts"

Query Length Strategy

  • 1-2 words: System favors FTS (60/40) - good for precise terms
  • 3-5 words: Balanced (50/50) - hybrid works well
  • 6+ words: System favors semantic (35/65) - captures intent

The system handles this automatically, but you can override by selecting a specific mode.

Performance Characteristics

Collection SizeHybrid p95Notes
1,000 docs<50msNo HNSW needed
10,000 docs<200msHNSW kicks in
100,000 docs<500msO(log N) scaling

The HNSW vector index provides logarithmic query complexity, so performance degrades slowly as the knowledge base grows.

Search Result Cache

When Redis caching is enabled, Fortémi caches only requests that explicitly use `mode=fts`. The cache identity includes the archive, query, filter expression, and result limit. Requests using tags, strict filters, time constraints, diversity ranking, or an embedding set bypass the cache.

Semantic, hybrid, and default-mode searches also bypass the cache. Their results depend on the effective embedding provider, model, dimensions, and configuration; caching remains disabled for those modes until that complete lineage is part of the cache identity. If semantic query embedding degrades to FTS, the degraded result is therefore never stored as a semantic or hybrid entry.

Successful note and tag mutations invalidate all search entries. Remaining entries expire after `REDIS_CACHE_TTL` seconds, which defaults to 300.

HNSW Tuning

The system dynamically adjusts the HNSW `ef_search` parameter based on:

  • Corpus size: Larger collections get higher ef_search
  • Recall target: Choose between Fast (85%), Balanced (92%), High (96%), or Exhaustive (99%)

Formula: `ef = base_ef max(1.0, log2(corpus_size / 10000) scale_factor)`

This balances recall and latency based on your collection size.

API Reference

For complete search API documentation including all parameters, request/response schemas, and examples, see:

Fortémi supports full-text search across multiple languages and scripts.

Supported Languages

Language/ScriptSupport LevelConfiguration
EnglishFull stemming`matric_english` (default)
GermanFull stemming`matric_german`
FrenchFull stemming`matric_french`
SpanishFull stemming`matric_spanish`
PortugueseFull stemming`matric_portuguese`
RussianFull stemming`matric_russian`
Chinese, Japanese, KoreanBigram/Trigram`matric_simple` + pg_bigm
Emoji & SymbolsTrigram matchingpg_trgm
Other scriptsBasic tokenization`matric_simple`

Query Syntax

The search system supports boolean operators via `websearch_to_tsquery`:

# OR operator
curl "http://localhost:3000/api/v1/search?q=apple+OR+orange"

# NOT operator
curl "http://localhost:3000/api/v1/search?q=python+-snake"

# Phrase search
curl "http://localhost:3000/api/v1/search?q=%22machine+learning%22"
SyntaxExampleDescription
Simple`hello world`Match all words (AND)
OR`apple OR orange`Match either word
NOT`apple -orange`Exclude word
Phrase`"hello world"`Match exact phrase
Combined`"machine learning" OR AI`Phrase OR single word

Language Hints

Specify language for better stemming results:

# German search
curl "http://localhost:3000/api/v1/search?q=Haus&lang=de"

# Chinese search
curl "http://localhost:3000/api/v1/search?q=人工智能&lang=zh"

# Japanese search
curl "http://localhost:3000/api/v1/search?q=プログラミング&lang=ja"

Script Detection

The system automatically detects query script and routes to the appropriate search strategy:

Detected ScriptSearch Strategy
LatinFTS with matric_english
CJK (Han, Hiragana, Katakana, Hangul)Bigram (pg_bigm) or Trigram fallback
CyrillicFTS with matric_russian
Arabic, Hebrew, GreekFTS with matric_simple
EmojiTrigram matching (pg_trgm)
Mixed scriptsMulti-strategy search

Emoji search uses pg_trgm trigram matching with ILIKE substring fallback.

Supported emoji patterns:

PatternExampleResult
Single emoji🚀 🔥 ⭐✅ Found
Repeated same🔥🔥✅ Found
Adjacent different🚀🎉✅ Found
Emoji + textmeeting 📝✅ Found
Emoji with variation selector❤️✅ Found

Supported Unicode ranges:

  • Emoticons (😀-🙏)
  • Misc Symbols and Pictographs (🌀-🗿)
  • Transport and Map (🚀-🛿)
  • Misc Symbols (☀️, ⚡, ☔)
  • Dingbats (✅, ✨, ✔️)
  • Misc Symbols and Arrows (⭐, ⬆️, ⬇️)
# Single emoji
curl "http://localhost:3000/api/v1/search?q=🎉"

# Adjacent emojis
curl "http://localhost:3000/api/v1/search?q=🚀🎉"

# Emoji with text
curl "http://localhost:3000/api/v1/search?q=meeting+📝"

How it works: When the query contains emoji, the system uses two strategies: 1. `similarity()` function for fuzzy trigram matching 2. `ILIKE '%emoji%'` for exact substring matching (fallback)

The ILIKE fallback ensures emoji sequences are found even when trigram similarity is low.

CJK Search Requirements

Minimum 2 characters required for CJK (Chinese, Japanese, Korean) queries.

Query LengthResultWhy
1 character (中)0 resultsBelow n-gram minimum
2+ characters (中文)✅ FoundMeets bigram/trigram threshold

This is an industry-standard limitation shared by all major search engines:

  • PostgreSQL pg_trgm requires 3 characters (trigrams)
  • PostgreSQL pg_bigm requires 2 characters (bigrams)
  • Elasticsearch CJK analyzers recommend 2+ characters
  • Google, Baidu, Naver all require 2+ characters for meaningful results

Why single characters don't work: N-gram indexes create searchable tokens from character sequences. A single CJK character doesn't generate enough tokens for reliable matching against document content.

# Chinese: 2+ characters required
curl "http://localhost:3000/api/v1/search?q=中文"      # ✅ Works
curl "http://localhost:3000/api/v1/search?q=人工智能"  # ✅ Works

# Japanese with hiragana
curl "http://localhost:3000/api/v1/search?q=日本語"    # ✅ Works

# Korean
curl "http://localhost:3000/api/v1/search?q=한국어"    # ✅ Works

Feature Flags

Multilingual features can be enabled via environment variables:

VariableDefaultDescription
`FTS_WEBSEARCH_TO_TSQUERY`trueOR/NOT/phrase operators
`FTS_SCRIPT_DETECTION`falseAutomatic script routing
`FTS_TRIGRAM_FALLBACK`falseEmoji/symbol search
`FTS_BIGRAM_CJK`falseOptimized CJK search
`FTS_MULTILINGUAL_CONFIGS`falseLanguage-specific stemming

Enable all multilingual features:

export FTS_SCRIPT_DETECTION=true
export FTS_TRIGRAM_FALLBACK=true
export FTS_BIGRAM_CJK=true
export FTS_MULTILINGUAL_CONFIGS=true

How Search Indexing Works

Understanding the underlying technology helps set appropriate expectations for search behavior.

PostgreSQL Extensions

Fortémi uses three PostgreSQL extensions for full-text search:

ExtensionPurposeMinimum Query Length
tsvector/tsqueryStandard FTS with stemming1+ characters (Latin scripts)
pg_trgmTrigram similarity matching3 characters
pg_bigmBigram matching (CJK-optimized)2 characters

N-gram Tokenization Explained

N-gram indexes work by breaking text into overlapping character sequences:

Trigrams (3-character sequences):

"hello" → {"  h", " he", "hel", "ell", "llo", "lo ", "o  "}

Bigrams (2-character sequences):

"日本語" → {"日本", "本語"}

The search query is also tokenized, and matching occurs when enough n-grams overlap between query and document. This is why minimum character requirements exist—short queries don't generate enough tokens for reliable matching.

Script-Specific Search Strategies

The system automatically selects the optimal strategy based on detected script:

┌─────────────────┐     ┌──────────────────────────────────────┐
│   User Query    │────▶│         Script Detection             │
└─────────────────┘     └──────────────────────────────────────┘
                                        │
        ┌───────────────┬───────────────┼───────────────┬───────────────┐
        ▼               ▼               ▼               ▼               ▼
   ┌─────────┐    ┌─────────┐    ┌─────────┐    ┌─────────┐    ┌─────────┐
   │  Latin  │    │   CJK   │    │ Cyrillic│    │  Emoji  │    │  Mixed  │
   └────┬────┘    └────┬────┘    └────┬────┘    └────┬────┘    └────┬────┘
        │              │              │              │              │
        ▼              ▼              ▼              ▼              ▼
   FTS English    pg_bigm or     FTS Russian    pg_trgm +      pg_trgm
   (stemming)     pg_trgm        (stemming)      ILIKE         fallback

Index Types and Their Characteristics

Index TypeBest ForLimitations
GIN tsvectorWord-based FTSNo substring matching
GIN pg_trgmSimilarity search, LIKE/ILIKE3-char minimum
GIN pg_bigmCJK, short strings2-char minimum
HNSW (pgvector)Semantic similarityRequires embeddings

Default Configuration

The Docker bundle enables these features by default:

# docker-compose.bundle.yml
environment:
  - FTS_SCRIPT_DETECTION=true
  - FTS_TRIGRAM_FALLBACK=true
  - FTS_BIGRAM_CJK=true
  - FTS_MULTILINGUAL_CONFIGS=true

Performance Implications

Query TypeIndex UsedComplexityTypical Latency
English keywordsGIN tsvectorO(log n)<10ms
CJK 2+ charsGIN bigm/trgmO(log n)<20ms
EmojiGIN trgm + ILIKEO(n) for ILIKE<50ms
SemanticHNSWO(log n)<100ms

The ILIKE fallback for emoji is slower (linear scan) but ensures correctness. For large collections with heavy emoji usage, consider pre-filtering with tags.

Advanced Topics

Understanding Score Components

In hybrid mode with adaptive weights, the final score combines:

1. FTS score - BM25 relevance (term frequency, document length normalization) 2. Semantic score - Cosine similarity of embeddings (0-1 range) 3. Fusion - RRF or RSF combines the two rankings 4. Adaptive weighting - Query-dependent balance between FTS and semantic

Choosing Between RRF and RSF

Use RRF when:

  • You want proven, unsupervised fusion
  • Rank position matters more than score magnitude
  • You need consistent behavior across query types

Use RSF when:

  • Score differences are meaningful (e.g., 0.9 vs 0.3)
  • You want to preserve quality gaps between results
  • You need slightly better recall (Weaviate FIQA: +6%)

Chunked Document Handling

Large documents are automatically chunked for embedding. The search system: 1. Searches across all chunks 2. Finds the most relevant chunk per document 3. Returns deduplicated results with chunk metadata 4. Preserves the best snippet from the highest-scoring chunk

This ensures comprehensive coverage while maintaining clean results.

Search across multiple memories simultaneously with unified result ranking.

Search All Memories

curl -X POST http://localhost:3000/api/v1/search/federated \
  -H "Content-Type: application/json" \
  -d '{
    "query": "machine learning",
    "memories": ["all"]
  }'

Search Specific Memories

curl -X POST http://localhost:3000/api/v1/search/federated \
  -H "Content-Type: application/json" \
  -d '{
    "query": "project documentation",
    "memories": ["default", "work-notes", "research"]
  }'

Federated Search Response

{
  "results": [
    {
      "note_id": "550e8400-...",
      "memory": "work-notes",
      "score": 0.92,
      "title": "Project Documentation",
      "snippet": "...machine learning algorithms...",
      "tags": ["project", "ml"]
    },
    {
      "note_id": "660e8400-...",
      "memory": "research",
      "score": 0.85,
      "title": "ML Research Papers",
      "snippet": "...deep learning techniques...",
      "tags": ["research", "ml"]
    }
  ],
  "total": 2,
  "memories_searched": ["work-notes", "research"]
}

How It Works

1. Parallel Execution: Searches run concurrently across all specified memories 2. Score Normalization: Each memory's scores are normalized to [0,1] range 3. Unified Ranking: Results are merged and re-sorted by normalized score 4. Memory Attribution: Each result includes `memory` field showing source

Performance Considerations

  • Federated search latency = slowest memory search time
  • Use specific memory names instead of `["all"]` when possible
  • Consider memory size when searching many memories (large memories slow down federation)

Use Cases

  • Cross-project search: Find related work across all project memories
  • Multi-client search: Search across client memories for patterns
  • Comprehensive research: Discover connections across research and work notes

See the Multi-Memory Guide for comprehensive documentation.

Troubleshooting Poor Results

No Results Returned

Possible CauseDiagnosisFix
No embeddings generatedCheck `/api/v1/jobs` for pending embed jobsWait for jobs or trigger via `/api/v1/jobs`
`degraded=true`Query embedding provider failed or returned an incompatible vectorInspect `search.embedding_degraded` logs and verify provider/model/dimension configuration
Wrong search modeFTS won't find semantic matchesTry `mode=hybrid` or `mode=semantic`
Strict filter too narrowTag filter excludes all notesBroaden filter or check tag spelling
Language mismatchNon-English content with English stemmerAdd `lang` parameter or enable `FTS_SCRIPT_DETECTION`
CJK query too shortSingle-character CJK queryUse 2+ characters (e.g., 中文 not 中)
Features not enabledScript detection disabledEnable `FTS_SCRIPT_DETECTION=true`

Irrelevant Results

Possible CauseDiagnosisFix
Too many unrelated notesCheck if embedding set is too broadUse tag-filtered embedding set
Short query, broad matches1-2 word queries match everythingAdd more context words or use FTS mode
Stale embeddingsNotes updated but not re-embeddedTrigger re-embedding via job queue

Slow Search Performance

Possible CauseDiagnosisFix
Missing HNSW indexCheck `pg_indexes` for embedding indexRun migrations to create index
High ef_searchQuery accuracy too high for your needsLower `hnsw.ef_search` (default: 64)
Large corpus without MRLFull-dimension search on 100K+ docsUse MRL truncation (256-dim)

See Troubleshooting Guide for comprehensive diagnostics.


See also: Architecture | Best Practices | Configuration | Multi-Memory Guide | Glossary