Hardware Planning

GPU, memory, and storage sizing guidelines.

Hardware Planning Guide

This guide helps you choose the right hardware for running Fortémi with local LLM inference, based on your quality requirements and budget.

Overview

Fortémi supports two inference backends:

  • Ollama - Local inference (this guide's focus)
  • OpenAI - Cloud or hybrid deployment

This guide focuses on local hardware planning for Ollama-based deployments, with quality benchmarks to help you make informed decisions.

Quality Tiers

Model quality is measured on a 0-100% scale based on:

  • Accuracy - Correctness of generated content
  • Coherence - Logical flow and consistency
  • Instruction following - Adherence to prompts
  • Domain knowledge - Depth of understanding

Quality tiers are benchmarked against GPT-4o (97-99% baseline).


Tier 1: Budget (4-8GB VRAM)

Quality Score: 75-80%

Entry-level hardware for basic local inference. Suitable for personal knowledge bases, note-taking, and simple summarization tasks.

Hardware Examples

GPUVRAMPrice RangeNotes
RTX 3060 12GB12GB$250-350 (used)Best value for tier 1
RTX 4060 Ti 8GB8GB$350-400 (new)Modern, power efficient
RTX 2060 6GB6GB$150-200 (used)Minimum viable GPU
AMD RX 6600 XT8GB$200-250 (used)AMD option, lower CUDA support
Intel Arc A77016GB$250-300 (new)Experimental, good VRAM/price

Intel Arc/XPU hosts run inference outside the container: the bundle's Intel overlay routes generation to a host vLLM endpoint while embeddings stay on an embedding-capable provider. See Intel Arc / vLLM Deployment for setup.

TaskModelSizeQualityLatency (P95)
Embeddings`nomic-embed-text`0.5GB85%~50ms
Generation`phi3:mini`3.8GB75%~2-3s
Generation`llama3.2:3b`3.2GB78%~1-2s
Generation`qwen2.5:3b`3.2GB76%~1-2s

Performance Characteristics

  • Context window: 4K-32K tokens
  • Output quality: Basic summarization, simple Q&A
  • Throughput: 20-90 tok/s
  • Best for: Personal use, learning, prototyping

Limitations

  • Struggles with complex reasoning
  • Limited code generation accuracy
  • May produce hallucinations on edge cases
  • Not suitable for production workloads

Example Configuration

# .env
OLLAMA_URL=http://localhost:11434
OLLAMA_EMBEDDING_MODEL=nomic-embed-text
OLLAMA_GENERATION_MODEL=llama3.2:3b
OLLAMA_NUM_GPU=99

Tier 2: Mainstream (12-16GB VRAM)

Quality Score: 85-90%

Recommended tier for most users. Provides excellent quality-to-cost ratio for professional knowledge management.

Hardware Examples

GPUVRAMPrice RangeNotes
RTX 3060 Ti 16GB16GB$400-500 (used)Rare, excellent value
RTX 4060 Ti 16GB16GB$500-600 (new)Best mainstream choice
RTX 3080 10GB10GB$400-500 (used)High performance, limited VRAM
RTX 4070 12GB12GB$550-650 (new)Good balance
AMD RX 7700 XT12GB$400-500 (new)AMD option
TaskModelSizeQualityLatency (P95)
Embeddings`nomic-embed-text`0.5GB85%~50ms
Generation`llama3.1:8b`8.0GB87%~3-5s
Generation`qwen2.5:7b`7.6GB89%~2-4s
Generation`mistral:7b`7.2GB86%~2-3s
Code`qwen2.5-coder:7b`7.6GB88%~3-5s

Performance Characteristics

  • Context window: 32K-128K tokens
  • Output quality: Strong reasoning, accurate summaries
  • Throughput: 25-175 tok/s
  • Best for: Professional use, team knowledge bases

Capabilities

  • Excellent instruction following
  • Good code generation (with specialized models)
  • Handles multi-step reasoning
  • Suitable for production use cases

Example Configuration

# .env
OLLAMA_URL=http://localhost:11434
OLLAMA_EMBEDDING_MODEL=nomic-embed-text
OLLAMA_GENERATION_MODEL=qwen2.5:7b
OLLAMA_NUM_CTX=32768
OLLAMA_NUM_GPU=99

Tier 3: Performance (24GB VRAM)

Quality Score: 93-95%

High-performance tier for demanding workloads. Near-GPT-4-level quality for most tasks.

Hardware Examples

GPUVRAMPrice RangeNotes
RTX 309024GB$800-1000 (used)Best value for 24GB
RTX 409024GB$1500-2000 (new)Fastest consumer GPU
RTX A500024GB$1200-1500 (used)Workstation GPU, ECC
AMD MI2516GB$500-700 (used)Older, limited software support
TaskModelSizeQualityLatency (P95)
Embeddings`mxbai-embed-large`1GB90%~100ms
Generation`qwen2.5:14b`14.8GB94%~5-8s
Generation`deepseek-coder-v2:16b`15.7GB93%~4-6s
Generation`mixtral:8x7b`~22GB92%~6-10s

Performance Characteristics

  • Context window: 32K-128K tokens (some 256K+)
  • Output quality: Expert-level reasoning
  • Throughput: 8-240 tok/s (depends on model)
  • Best for: Enterprise knowledge bases, research

Capabilities

  • Advanced reasoning and planning
  • High-quality code generation
  • Minimal hallucinations
  • Excellent domain expertise

Example Configuration

# .env
OLLAMA_URL=http://localhost:11434
OLLAMA_EMBEDDING_MODEL=mxbai-embed-large
OLLAMA_GENERATION_MODEL=qwen2.5:14b
OLLAMA_NUM_CTX=65536
OLLAMA_NUM_GPU=99
OLLAMA_NUM_PARALLEL=2

Tier 4: Professional (48GB+ VRAM)

Quality Score: 95-97%

Workstation-class hardware for maximum local performance. Approaches cloud API quality.

Hardware Examples

GPUVRAMPrice RangeNotes
RTX 6000 Ada48GB$6000-7000 (new)Professional workstation
A600048GB$4000-5000 (used)Ampere generation
H100 PCIe80GB$25000+ (new)Data center GPU
Dual RTX 309048GB (2x24)$1600-2000 (used)Multi-GPU setup
Dual RTX 409048GB (2x24)$3000-4000 (new)Multi-GPU performance
TaskModelSizeQualityLatency (P95)
Embeddings`mxbai-embed-large`1GB90%~100ms
Generation`qwen2.5:32b`32.0GB96%~10-15s
Generation`llama3.1:70b`~40GB95%~15-20s
Code`deepseek-coder-v2:33b`33.0GB96%~12-18s

Performance Characteristics

  • Context window: 128K+ tokens
  • Output quality: Matches GPT-4 quality
  • Throughput: 5-60 tok/s (large models)
  • Best for: Enterprise, research institutions

Capabilities

  • State-of-the-art reasoning
  • Production-grade reliability
  • Handles complex multi-turn conversations
  • Suitable for critical applications

Example Configuration

# .env
OLLAMA_URL=http://localhost:11434
OLLAMA_EMBEDDING_MODEL=mxbai-embed-large
OLLAMA_GENERATION_MODEL=qwen2.5:32b
OLLAMA_NUM_CTX=131072
OLLAMA_NUM_GPU=99
OLLAMA_NUM_PARALLEL=1  # Large models need full VRAM

Tier 5: Cloud (Reference Baseline)

Quality Score: 97-99%

Cloud API providers for comparison. No hardware investment required.

Service Providers

ProviderModelQualityCost (Input/Output)Latency
OpenAIGPT-4o98%$2.50/$10/1M tokens~1-2s
OpenAIGPT-4o-mini92%$0.15/$0.60/1M tokens~500ms
AnthropicClaude 3.5 Sonnet99%$3.00/$15/1M tokens~1-3s
AnthropicClaude 3 Haiku90%$0.25/$1.25/1M tokens~500ms
GoogleGemini 1.5 Pro97%$1.25/$5/1M tokens~1-2s

Hybrid Deployment

Combine local and cloud for optimal cost/quality:

# Use local for embeddings (bulk operations)
OLLAMA_EMBEDDING_MODEL=nomic-embed-text

# Use cloud for generation (quality-critical)
OPENAI_GENERATION_MODEL=gpt-4o-mini

Cost Comparison

Example workload: 10,000 notes, 500 queries/month

DeploymentHardware CostMonthly CostQuality
Tier 1 Local$300$075-80%
Tier 2 Local$600$085-90%
Tier 3 Local$1500$093-95%
Cloud Only$0$50-20097-99%
Hybrid (Local embed + Cloud gen)$600$10-3095-98%

Model Selection by VRAM Chart

Quick reference for choosing models based on available VRAM:

VRAM   | Recommended Models              | Quality Tier
-------|----------------------------------|-------------
4GB    | llama3.2:1b, phi3:mini          | Tier 1 (75%)
6GB    | llama3.2:3b, qwen2.5:3b         | Tier 1 (78%)
8GB    | mistral:7b, qwen2.5:7b          | Tier 2 (86-89%)
12GB   | llama3.1:8b, qwen2.5-coder:7b   | Tier 2 (87-89%)
16GB   | llama3.1:8b, qwen2.5:14b (Q4)   | Tier 2-3 (90-92%)
24GB   | qwen2.5:14b, deepseek-r1:14b    | Tier 3 (93-95%)
40GB+  | qwen2.5:32b, llama3.1:70b (Q4)  | Tier 4 (95-97%)
80GB+  | llama3.1:70b, qwen2.5:72b       | Tier 4+ (96-97%)

Notes:

  • Q4 = 4-bit quantization (trades quality for VRAM)
  • Add 1-2GB for embedding model
  • Reserve 10-20% VRAM for context window

Upgrade Paths

From Tier 1 to Tier 2

Investment: $300-400 Quality gain: +10-12% ROI: High - Significant capability improvement

Recommended:

  • RTX 4060 Ti 16GB ($500-600 new)
  • RTX 3080 10GB ($400-500 used)

Model upgrade:

  • `llama3.2:3b` → `qwen2.5:7b`
  • Gain: Better reasoning, code generation

From Tier 2 to Tier 3

Investment: $800-1200 Quality gain: +5-8% ROI: Medium - Diminishing returns

Recommended:

  • RTX 3090 24GB ($800-1000 used)
  • RTX 4090 24GB ($1500-2000 new)

Model upgrade:

  • `qwen2.5:7b` → `qwen2.5:14b`
  • Gain: Advanced reasoning, minimal hallucinations

From Tier 3 to Tier 4

Investment: $2500-5000 Quality gain: +2-4% ROI: Low - Consider cloud hybrid instead

Alternative:

  • Hybrid deployment: Local Tier 3 + cloud for critical tasks
  • Cost: $10-30/month, Quality: 95-98%

Cost Calculator Concept

Total Cost of Ownership (3 Years)

Tier 1 (RTX 3060 12GB)

  • Hardware: $300
  • Power (150W @ $0.12/kWh, 8h/day): $158/year → $474
  • Total: $774
  • Cost per quality point: $10.32/point (75% quality)

Tier 2 (RTX 4060 Ti 16GB)

  • Hardware: $550
  • Power (165W @ $0.12/kWh, 8h/day): $173/year → $519
  • Total: $1069
  • Cost per quality point: $11.99/point (89% quality)

Tier 3 (RTX 4090 24GB)

  • Hardware: $1800
  • Power (450W @ $0.12/kWh, 8h/day): $473/year → $1419
  • Total: $3219
  • Cost per quality point: $34.24/point (94% quality)

Cloud (GPT-4o-mini)

  • Hardware: $0
  • Usage (10K notes, 500 queries/month): $50/month → $1800
  • Total: $1800
  • Cost per quality point: $19.57/point (92% quality)

Hybrid (Tier 2 + Cloud)

  • Hardware: $550
  • Power: $519 (3 years)
  • Usage (local embed, cloud gen): $20/month → $720
  • Total: $1789
  • Cost per quality point: $18.65/point (96% quality)

Break-Even Analysis

Local vs Cloud (GPT-4o-mini):

  • Tier 1: Breaks even in 6 months
  • Tier 2: Breaks even in 14 months
  • Tier 3: Breaks even in 32 months

Recommendation: Tier 2 local + cloud hybrid offers best ROI for most users.


Decision Matrix

Choose your tier based on these criteria:

Use CaseRecommended TierRationale
Personal notes, learningTier 1Adequate quality, low cost
Professional knowledge baseTier 2Best quality/cost ratio
Team collaboration (5-20 users)Tier 2-3Consistent performance
Enterprise (20+ users)Tier 3-4 or HybridReliability, scale
Research, experimentationTier 3Flexibility, no API limits
Privacy-critical dataTier 2-3 LocalNo cloud egress
Cost-sensitive, high qualityHybridOptimize per task

Performance Testing

Benchmark Your Setup

# Test embedding performance
time ollama pull nomic-embed-text
time curl http://localhost:11434/api/embeddings \
  -d '{"model": "nomic-embed-text", "prompt": "test text"}'

# Test generation performance
time ollama pull qwen2.5:7b
time curl http://localhost:11434/api/generate \
  -d '{"model": "qwen2.5:7b", "prompt": "Explain quantum computing", "stream": false}'

# Monitor VRAM usage
nvidia-smi -l 1

Expected Latencies

Model SizeFirst TokenFull Response (500 tokens)
3B~200ms2-5s
7-8B~300ms3-8s
14B~500ms5-15s
32B+~1000ms10-30s

System Requirements

Minimum (Tier 1)

  • GPU: 4GB+ VRAM
  • CPU: 4 cores
  • RAM: 8GB
  • Storage: 50GB SSD
  • OS: Linux (Ubuntu 22.04+), Windows 10+, macOS 12+
  • GPU: 12-16GB VRAM
  • CPU: 8 cores
  • RAM: 16GB
  • Storage: 100GB NVMe SSD
  • OS: Linux (for best performance)

Optimal (Tier 3+)

  • GPU: 24GB+ VRAM
  • CPU: 16+ cores
  • RAM: 32GB+
  • Storage: 500GB NVMe SSD
  • OS: Linux with CUDA 12+

Memory Capacity Planning

`MAX_MEMORIES` limits the number of live memories (active schemas in the database at once). This is not a cap on how many archives you can ever create — export a memory as a shard, delete it to free a slot, and re-import when needed. There is no limit on archived shards stored on disk.

Each empty live memory adds ~1MB overhead (41 tables + indexes). The real storage cost comes from data within each memory.

Storage Per Note

Average storage per note assuming 20% of notes have file attachments:

ComponentPer NotePer 1,000 Notes
Note metadata + content~11 KB11 MB
Embeddings (768-dim, 4 bytes/float)~3 KB3 MB
Attachments (avg 500KB, 20% rate)~100 KB100 MB
Thumbnails (100KB, 20% rate)~20 KB20 MB
Total average~134 KB~134 MB

Notes without attachments average ~14 KB each. Text-only knowledge bases can store 7x more notes in the same space.

Capacity Formula

available_storage = total_disk - OS - models - PostgreSQL_overhead
max_total_notes   = available_storage / 134 KB     (with attachments)
                  = available_storage /  14 KB     (text-only)
MAX_MEMORIES      = max_total_notes / target_notes_per_memory
TierRAMUsable StorageMAX_MEMORIESNotes/MemoryTotal Notes`MAX_MEMORIES` Setting
Tier 1 (Minimum)8 GB~7 GB10~5,000~50,000`10` (default)
Tier 2 (Standard)16 GB~70 GB50~20,000~1,000,000`50`
Tier 3 (Performance)32 GB~400 GB200~50,000~10,000,000`200`
Tier 4 (Professional)64 GB+~800 GB+500~50,000~25,000,000`500`

Assumptions: Usable storage accounts for OS (~3GB), Ollama models (~10-30GB depending on tier), and PostgreSQL WAL/overhead (~10%). Notes/memory targets keep search sub-second per memory.

Configuration

# .env — set based on your hardware tier
MAX_MEMORIES=10   # Tier 1 default (8GB RAM, 10GB disk)
MAX_MEMORIES=50   # Tier 2 (16GB RAM, 100GB disk)
MAX_MEMORIES=200  # Tier 3 (32GB RAM, 500GB disk)
MAX_MEMORIES=500  # Tier 4 (64GB+ RAM, 1TB+ disk)

See Configuration Reference for all memory-related settings and Multi-Memory Guide for architecture details.



Appendix: Model Quality Benchmarks

Quality scores based on comprehensive testing across multiple dimensions:

Testing Methodology

  • Task diversity: Summarization, Q&A, reasoning, code generation
  • Evaluation metrics: Accuracy, coherence, instruction following
  • Baseline: GPT-4o (97-99% quality)
  • Sample size: 100+ test cases per model

Tier 1 Models (75-80%)

ModelSizeQualityStrengthsWeaknesses
llama3.2:3b3.2B78%Fast, good general useLimited reasoning depth
phi3:mini3.8B75%Compact, efficientStruggles with long context
qwen2.5:3b3.2B76%Good instruction followingCode generation weaker

Tier 2 Models (85-90%)

ModelSizeQualityStrengthsWeaknesses
qwen2.5:7b7.6B89%Excellent reasoningSlightly slower
llama3.1:8b8.0B87%Well-rounded, reliableLarge model size
mistral:7b7.2B86%Fast, good qualityContext window smaller
qwen2.5-coder:7b7.6B88%Strong code generationLess general knowledge

Tier 3 Models (93-95%)

ModelSizeQualityStrengthsWeaknesses
qwen2.5:14b14.8B94%Near-GPT-4 reasoningSlower inference
deepseek-r1:14b14.8B93%Chain-of-thoughtHigher latency
deepseek-coder-v2:16b15.7B93%Excellent code qualityDomain-specific

Tier 4 Models (95-97%)

ModelSizeQualityStrengthsWeaknesses
qwen2.5:32b32.0B96%State-of-the-art localHigh VRAM requirement
llama3.1:70b70B95%Comprehensive knowledgeVery slow, 48GB+ VRAM

Updates

This guide is based on models available as of January 2025. Model capabilities and hardware prices change rapidly. Check the Ollama model library for the latest releases.

For custom hardware recommendations or enterprise deployment planning, consult the Operations Guide or contact support.