Embedding Model Selection

Choosing embedding models: dimensions, MRL, and storage trade-offs.

Embedding Model Selection Guide

This guide helps you select the optimal embedding model for your use case, based on empirical research findings.

Key Insight: Domain Matters More Than Model Size

Research demonstrates counter-intuitive findings:

1. Bigger is not always better: MiniLM-v6 (22M params) outperforms BGE-Large (335M params) by 7.7-23.1% when combined with LLM re-ranking [REF-068].

2. Domain factor has 1.00 effect size vs model choice for specialized tasks [REF-069].

3. MRL enables flexible trade-offs: 12× storage reduction with <5% quality loss using dimension truncation [REF-067, REF-070].

Model Comparison

ModelParametersDimensionsMRL SupportBest For
nomic-embed-text-v1.5137M768✅General purpose, MRL
all-MiniLM-L6-v222M384❌Fast, LLM re-ranking
bge-large-en-v1.5335M1024❌High accuracy, no re-ranking
mxbai-embed-large-v1335M1024✅High accuracy + MRL
e5-mistral-7b7B4096❌Maximum quality
multilingual-e5-large560M1024❌Non-English content

Decision Tree

                    START
                      │
                      ▼
            ┌─────────────────┐
            │ Using LLM       │
            │ re-ranking?     │
            └────────┬────────┘
                     │
           ┌─────────┴─────────┐
           │                   │
           ▼                   ▼
          YES                  NO
           │                   │
           ▼                   ▼
    ┌──────────────┐   ┌──────────────┐
    │ MiniLM-v6    │   │ Need storage │
    │ (384-dim)    │   │ optimization?│
    └──────────────┘   └──────┬───────┘
                              │
                     ┌────────┴────────┐
                     │                 │
                     ▼                 ▼
                    YES                NO
                     │                 │
                     ▼                 ▼
              ┌──────────────┐   ┌──────────────┐
              │ MRL-enabled  │   │ Domain-      │
              │ (nomic,      │   │ specific?    │
              │ mxbai)       │   └──────┬───────┘
              └──────────────┘          │
                                 ┌──────┴──────┐
                                 │             │
                                 ▼             ▼
                                YES           NO
                                 │             │
                                 ▼             ▼
                          ┌──────────────┐ ┌──────────────┐
                          │ Fine-tune    │ │ bge-large or │
                          │ base model   │ │ e5-mistral   │
                          └──────────────┘ └──────────────┘

Use Case Recommendations

General Knowledge Base

  • Recommended: nomic-embed-text-v1.5 (768-dim)
  • With MRL: Truncate to 256-dim for 3× storage savings
  • Why: Good balance of quality and efficiency
{
  "embedding_config": "default",
  "truncate_dim": null
}

RAG with Claude/GPT Re-ranking

  • Recommended: all-MiniLM-L6-v2 (384-dim)
  • Why: Per REF-068, smaller models perform BETTER with LLM re-ranking
  • Latency: Fastest embedding generation
{
  "embedding_config_id": "minilm-config-id",
  "truncate_dim": null
}
  • Recommended: Fine-tune gte-large-en-v1.5 or e5-mistral
  • Why: Per REF-069, 88% retrieval improvement via fine-tuning
  • Data needed: ~6,000 synthetic query-document pairs
# Generate training data
POST /api/v1/fine-tuning/generate
{
  "name": "legal-training",
  "source": {"type": "embedding_set", "slug": "legal-docs"},
  "config": {"queries_per_doc": 4}
}

Multilingual / CJK Content

  • Recommended: multilingual-e5-large (1024-dim)
  • Alternative: intfloat/multilingual-e5-small for speed
  • Why: Trained on 100+ languages
{
  "name": "CJK Content",
  "embedding_config_id": "multilingual-e5-config-id"
}
  • Recommended: codesearchnet or codegen-350M-mono
  • Why: Trained on programming languages

Maximum Storage Efficiency

  • Recommended: nomic-embed-text-v1.5 with MRL @ 64-dim
  • Trade-off: 12× smaller, ~3% quality loss
  • Best for: Large corpora (>1M documents)
{
  "set_type": "full",
  "embedding_config_id": "nomic-config-id",
  "truncate_dim": 64
}

Matryoshka Representation Learning (MRL)

MRL-trained models encode information hierarchically, allowing embeddings to be truncated to smaller dimensions while preserving most quality.

Valid MRL Dimensions

Only specific dimensions produce quality results. Using arbitrary dimensions degrades quality significantly.

ModelValid Dimensions
nomic-embed-text768, 512, 256, 128, 64
mxbai-embed-large1024, 512, 256, 128, 64

Quality vs Size Trade-offs

DimensionStorage ReductionQuality Loss
64-dim12×~3-5%
128-dim6×~2%
256-dim3×~1%
Full1×0%

Two-Stage Retrieval

Use MRL for efficient coarse-to-fine search:

1. Stage 1 (Coarse): Search 64-dim index for 100 candidates 2. Stage 2 (Fine): Re-rank with full 768-dim similarity

Result: 128× MFLOP reduction with same Recall@1 [REF-067].

GET /api/v1/search?q=...&strategy=two_stage&coarse_dim=64&coarse_k=100

Anti-Patterns to Avoid

❌ Assuming bigger = better

REF-068 shows MiniLM-v6 (22M) beats BGE-Large (335M) when using LLM re-ranking.

❌ Using non-MRL models with dimension truncation

Standard models lose significant quality when truncated. Only use MRL-trained models for truncation.

// Wrong: truncating bge-large (non-MRL)
// This will produce poor quality embeddings
config.truncate_dim = Some(128); // ❌ DON'T DO THIS

❌ Fine-tuning when retrieval is already strong

Per REF-069, fine-tuning on DocsQA showed minimal gains because baseline was already good. Only fine-tune when baseline Recall@10 < 60%.

❌ Ignoring latency vs accuracy trade-offs

For real-time search, a fast 384-dim model may outperform a slow 4096-dim model in practice.

Performance Benchmarks

Storage Requirements (per 1M documents)

ModelDimensionsStorageWith MRL-128
MiniLM-v63841.5 GBN/A
nomic-embed7683.0 GB0.5 GB
bge-large10244.0 GBN/A
e5-mistral409616.0 GBN/A

Latency (Ollama, RTX 4090)

ModelPer DocBatch (100 docs)
MiniLM-v65ms150ms
nomic-embed12ms350ms
bge-large25ms750ms
e5-mistral200ms6000ms

API Examples

Create MRL-Enabled Set

POST /api/v1/embedding-sets
{
  "name": "Fast Search",
  "slug": "fast-search",
  "set_type": "full",
  "embedding_config_id": "nomic-config-id",
  "truncate_dim": 256,
  "auto_embed_rules": {
    "on_create": true,
    "on_update": true
  }
}

Search with Two-Stage Strategy

GET /api/v1/search?q=machine+learning&strategy=two_stage

Validate MRL Truncation

GET /api/v1/embedding-configs/nomic-embed-text

Response:
{
  "supports_mrl": true,
  "matryoshka_dims": [768, 512, 256, 128, 64],
  "default_truncate_dim": 256
}

References

Academic Papers

  • REF-067: Kusupati et al. (2022). Matryoshka Representation Learning. NeurIPS.
  • REF-068: Rao et al. (2025). Rethinking Hybrid Retrieval for RAG Systems. arXiv:2506.00049.
  • REF-069: Portes et al. (2025). Improving Retrieval and RAG with Embedding Finetuning. Databricks.
  • REF-070: Aarsen et al. (2024). Matryoshka Embedding Models. HuggingFace.

Industry Resources