Vector Databases in Production: From Embeddings to Hybrid Search
Vector search started as a demo feature: embed a few documents, run a nearest-neighbor query, and show surprisingly relevant results. Production is different. You need predictable latency under load, fresh data, multi-tenant isolation, cost controls, and an evaluation loop that proves relevance. This article explains how modern vector databases work, where they fail, and how to design a system that survives real traffic.
Why vector search becomes a production problem
A vector database stores high-dimensional embeddings and retrieves the nearest vectors to a query embedding. That sounds simple until you add real-world constraints. The hard part is not the first million vectors. The hard part is keeping recall high while latency stays low as the corpus grows, updates arrive continuously, and tenants share infrastructure.
- Latency: Approximate nearest-neighbor indexes trade recall for speed. A configuration that returns answers in 10 ms on a laptop may collapse to 200 ms under concurrent writes and filtered queries.
- Freshness: Embeddings become stale when source documents change. You need an ingestion pipeline that re-embeds, upserts, and deletes vectors without rebuilding the entire index.
- Recall: Users do not care about index type. They care whether the right result appears in the top results. Recall must be measured against a golden set, not assumed from benchmark charts.
- Isolation: Multi-tenant systems must prevent one tenant from querying another tenant’s vectors, even if both share the same collection or index.
- Cost: Memory, storage, and GPU or CPU time all scale with vector count, dimensions, and replication. Without quantization and tiering, costs grow faster than value.
These constraints pull against each other. Higher recall often means more memory and slower queries. Stronger filtering can reduce the candidate set and improve speed, but it can also break the assumptions of the index. The right architecture starts by naming the trade-offs instead of hiding them.
Embeddings, distance metrics, and the meaning of similarity
An embedding is a dense numeric vector that represents text, images, audio, user behavior, or any object. Embedding models map similar items to nearby points in a high-dimensional space. The database does not understand semantics. It only computes distances between vectors. That means the quality of your search depends heavily on the embedding model, the preprocessing pipeline, and the distance metric.
Choosing a distance metric
- Cosine similarity: Measures the angle between vectors. It is common for text embeddings because it ignores magnitude and focuses on direction. It works well when vector length is not meaningful.
- Dot product: Measures both angle and magnitude. It is often used when the embedding model was trained with dot product or when magnitude carries signal, such as popularity or confidence.
- Euclidean distance: Measures straight-line distance. It is intuitive, but it can be dominated by vector magnitude unless vectors are normalized.
If vectors are normalized to unit length, cosine similarity and dot product become equivalent. Many teams normalize at ingestion time to simplify tuning. The important rule is consistency: use the same metric during index build, query, and evaluation. Mixing metrics silently degrades relevance.
Dimension also matters. A 1536-dimensional vector consumes more memory than a 384-dimensional vector. Larger models may improve recall, but they also increase index build time, query cost, and storage. In production, it is often better to start with a smaller model, measure recall, and only increase dimensions when evaluation proves a meaningful gain.
Index families: HNSW, IVF, PQ, DiskANN, and LSM-based designs
Vector indexes fall into two broad categories: graph-based and cluster-based. Each has different memory, latency, and recall characteristics. Most production databases combine multiple techniques to balance cost and performance.
HNSW
Hierarchical Navigable Small World is a graph index. It builds a multi-layer graph where upper layers provide long-range links and lower layers provide fine-grained connections. Queries start at an entry point, greedily move toward closer neighbors, and then search locally. HNSW usually delivers excellent recall and low latency, but it is memory-hungry because the graph edges and vectors must be accessible during traversal.
Key tuning parameters include M, which controls the number of connections per node, and efConstruction and efSearch, which control build quality and query effort. Higher values increase recall but also increase memory, build time, and query latency. HNSW is a strong default for read-heavy workloads that fit in memory.
IVF and IVF-PQ
Inverted File Index partitions vectors into clusters using k-means. At query time, the database searches only the closest clusters, which reduces the number of distance computations. IVF can be combined with Product Quantization, which compresses vectors into smaller codes. IVF-PQ dramatically reduces memory usage, but it lowers recall because quantized vectors are approximations.
IVF works well for very large corpora where memory is constrained. The main tuning knobs are the number of clusters and the number of clusters probed per query, often called nprobe. Higher nprobe improves recall at the cost of speed. IVF indexes also need periodic retraining as data distribution shifts.
DiskANN
DiskANN is designed for datasets that do not fit in RAM. It combines graph-based search with solid-state drive storage. The graph is stored on disk, and the system uses compressed vectors or careful caching to reduce I/O. DiskANN can deliver high recall with lower memory cost, but it depends on fast SSDs and careful capacity planning. It is a good fit for large, cost-sensitive corpora where in-memory HNSW is too expensive.
LSM-tree plus vector indexes
Some modern databases integrate vector indexes with an LSM-tree storage engine. Writes go to memory and are flushed to immutable segments. Each segment may have its own vector index. Queries fan out across segments, merge results, and handle deletes through tombstones. This design supports high write throughput and fresh data, but it complicates recall because each segment is approximate. Compaction merges segments and rebuilds indexes, which can create latency spikes if not managed.
Schema design for vector collections
A vector collection is more than a vector field. Production schemas need identity, metadata, tenant isolation, and lifecycle fields. A typical record includes:
- Primary key: A stable identifier from the source system. Avoid auto-incrementing IDs that change when data is re-ingested.
- Vector field: The embedding itself, with a fixed dimension and metric configured at collection creation.
- Metadata fields: Structured attributes such as category, language, timestamp, author, access level, and source. These fields power filtering and hybrid search.
- Tenant identifier: A mandatory field for multi-tenant isolation. It should be indexed and used in every query path.
- Version or timestamp: Useful for freshness, rollback, and conflict resolution when multiple writers update the same logical entity.
Keep metadata small and queryable. Large text fields belong in the original document store, not in the vector record. Use the vector database to retrieve candidate IDs, then fetch full content from a primary store. This separation reduces memory pressure and keeps the vector index focused on similarity.
Filtering: the hardest part of vector search
Pure vector search ignores business rules. Real users ask for results within a date range, in a language, from a tenant, or with a specific permission level. Filtering is where many vector systems fail.
- Pre-filtering: Apply filters before vector search. This is correct but can be slow if the filter is not selective. If only 1 percent of vectors match, the index may waste effort exploring irrelevant regions.
- Post-filtering: Search vectors first, then discard non-matching results. This is fast but can return too few results if the filter is selective. Increasing the search breadth helps, but it also increases latency.
- In-filtering: The index is aware of filters and traverses only valid nodes. This is the best approach when the database supports it, but not all vector indexes handle complex boolean filters efficiently.
A practical pattern is to combine strategies. Use a selective pre-filter for tenant and permission boundaries, then in-filter for common attributes, and finally post-filter for rare conditions. Always test with realistic filter combinations. A query that works on unfiltered data can behave very differently when 99 percent of vectors are excluded.
Scaling: sharding, replication, and tiered storage
Scaling vector search requires decisions about how to partition data, how to replicate it, and how to move cold data out of expensive memory.
Sharding
Sharding splits vectors across nodes. A common approach is hash-based sharding on the primary key or tenant ID. This distributes load evenly, but queries that span many tenants must fan out to all shards and merge results. Tenant-aware sharding keeps each tenant on a subset of shards, which improves isolation and reduces cross-shard queries, but it can create hotspots if tenants are large or uneven.
Replication
Replication improves availability and read throughput. Read replicas can serve queries while the primary handles writes. For vector indexes, replication is not just copying rows. Each replica must maintain its own index, and index build time can lag behind data ingestion. Plan for replica warm-up and avoid routing traffic to replicas that are still building.
Tiered storage
Not all vectors need the same latency. Recent or popular content can live in an in-memory HNSW index. Older or cold content can move to a disk-based or quantized index. A tiered design queries the hot tier first and falls back to the cold tier when needed. This reduces cost while keeping common queries fast. The trade-off is complexity: you must manage promotion, demotion, and deduplication across tiers.
Hybrid search: vectors plus keywords plus business rules
Vector search alone is not enough for many applications. It can miss exact matches, rare terms, product codes, and acronyms. Hybrid search combines dense vector retrieval with lexical retrieval and structured filters.
- Lexical retrieval: Use BM25 or another keyword scorer to find exact terms and rare tokens. This is fast, interpretable, and strong for precision.
- Vector retrieval: Use embeddings to capture semantic similarity, synonyms, and intent. This improves recall for natural language queries.
- Fusion: Merge result lists using weighted scores, reciprocal rank fusion, or a learned model. Reciprocal rank fusion is simple and robust because it uses ranks instead of raw scores, which are hard to compare across retrievers.
- Reranking: Take the top candidates from fusion and rerank them with a cross-encoder or a business-specific model. Reranking is more expensive per candidate, so keep the candidate set small, often 50 to 200 items.
A hybrid pipeline might look like this: parse the query, apply mandatory filters, run lexical search and vector search in parallel, fuse the results, rerank the top candidates, and then apply final business rules such as inventory, permissions, or diversity. Each stage should be measurable so you can see where relevance is lost.
Evaluation: measuring relevance, not just recall
You cannot tune what you do not measure. Vector search evaluation needs offline datasets and online feedback. Offline evaluation uses a golden set of queries with known relevant documents.
- Recall@k: The fraction of relevant documents found in the top k results. High recall is necessary but not sufficient.
- Mean Reciprocal Rank: Measures how high the first relevant result appears. It is useful for question answering and navigational queries.
- Normalized Discounted Cumulative Gain: Rewards relevant results that appear near the top and supports graded relevance.
- Online metrics: Click-through rate, dwell time, conversion rate, and task success. These reveal whether offline gains translate to user value.
Build a golden set from real user queries, not from synthetic examples. Include hard cases: misspellings, multi-intent queries, negations, and rare terms. Re-run evaluation whenever you change the embedding model, index parameters, chunking strategy, or fusion weights. Track results by segment, such as language, tenant, and query type, because aggregate metrics can hide regressions.
Observability and operations
Vector databases need the same observability as any critical data system, plus metrics specific to similarity search.
- Latency: Track p50, p95, and p99 for vector search, filtered search, and hybrid search separately. Tail latency often comes from index rebuilds or cold storage access.
- Recall: Monitor recall against a canary query set. A drop can indicate index corruption, stale embeddings, or parameter drift.
- Throughput: Measure queries per second and writes per second. Watch for backpressure during bulk ingestion.
- Index health: Track build time, segment count, graph connectivity, and memory usage. For HNSW, monitor graph degree distribution. For IVF, monitor cluster balance.
- Freshness: Measure the lag between source update, embedding generation, and index visibility. Users notice stale results quickly.
Run load tests with realistic query mixes. A workload that is 90 percent unfiltered vector search behaves very differently from one that is 50 percent filtered hybrid search. Include delete and update operations. Many systems perform well on append-only data and degrade when deletes create tombstones or fragmented segments.
Cost and performance tuning
Cost control in vector search comes from reducing memory, storage, and compute without destroying recall. The biggest levers are dimensionality, quantization, index type, and replication.
- Dimensionality reduction: Use a smaller embedding model or apply techniques such as PCA when the model supports it. Test recall carefully; some semantic distinctions disappear when dimensions shrink.
- Quantization: Scalar quantization, product quantization, and binary quantization reduce memory. They lower recall, but a reranking step can recover much of the loss. Quantized vectors are often used for candidate generation, while full vectors rerank the top results.
- Index selection: Use HNSW for hot data and latency-sensitive queries. Use IVF-PQ or DiskANN for large cold datasets. Match the index to the access pattern instead of forcing one index to serve everything.
- Replication strategy: Replicate only what you need. Extra replicas improve read throughput and availability, but they multiply memory and storage costs.
- Batch ingestion: Build indexes in batches rather than updating on every single write. This improves throughput and reduces fragmentation, at the cost of freshness. Tune batch size to meet your freshness target.
Always benchmark with the same hardware and data distribution as production. A configuration that works on 1 million vectors may fail at 100 million. Measure memory per vector, query latency at target QPS, and recall against your golden set.
Security, privacy, and multi-tenancy
Vector databases often store sensitive data. Embeddings can leak information, and metadata may contain personal data. Security must cover access control, encryption, and isolation.
- Tenant isolation: Enforce tenant filters at the query layer and, where possible, at the storage layer. Never rely on application code alone to add tenant predicates.
- Access control: Use role-based or attribute-based access control. Limit who can read, write, delete, and rebuild indexes. Separate ingestion credentials from query credentials.
- Encryption: Encrypt data at rest and in transit. If the database supports client-side encryption, evaluate whether it affects index quality or query performance.
- Privacy: Treat embeddings as personal data when they are derived from personal data. Support deletion requests by removing vectors, metadata, and any derived caches. Remember that backups and replicas also need deletion propagation.
- Audit logging: Log queries, filters, and result counts without storing raw sensitive queries when possible. Audit logs are essential for incident response and compliance.
Multi-tenancy is not just a filter. It affects sharding, resource quotas, noisy neighbor prevention, and index isolation. Large tenants may need dedicated collections or clusters, while small tenants can share infrastructure with strict quotas.
Production checklist
- Define the retrieval objective: semantic search, recommendation, anomaly detection, or deduplication. The objective determines the embedding model and metrics.
- Normalize vectors if using cosine similarity or dot product, and keep the metric consistent across ingestion, query, and evaluation.
- Choose an index family based on memory, latency, and recall requirements. Start with HNSW for in-memory workloads and IVF-PQ or DiskANN for large or cost-sensitive datasets.
- Design metadata for filtering from day one. Include tenant, permissions, timestamps, and source identifiers.
- Build a golden query set and measure recall, MRR, and nDCG before and after every change.
- Implement hybrid search with lexical retrieval, vector retrieval, fusion, and reranking. Tune fusion weights with evaluation, not intuition.
- Plan for deletes, updates, and re-embedding. Use stable primary keys and version fields.
- Monitor latency, recall, throughput, freshness, and index health. Alert on tail latency and recall drops.
- Control cost with quantization, tiered storage, and right-sized replication. Re-evaluate quarterly as data and usage grow.
- Enforce security and privacy at the database layer. Test tenant isolation with adversarial queries.
Conclusion
Vector databases are powerful, but production success comes from treating them as part of a retrieval system, not as a magic box. The embedding model, index configuration, filter strategy, fusion logic, and evaluation loop all shape relevance. Start with clear metrics, choose an index that matches your latency and cost constraints, and build observability before you scale. When those pieces are in place, vector search becomes a reliable capability instead of a demo that only works on clean data.

