Skip to content
AI & Automation

Vector Databases for AI: How They Work and When to Use Them

How vector databases work and how to choose one: approximate nearest neighbour indexes such as HNSW, filtering, hybrid search, scaling, cost, and when Postgres with pgvector is enough.

Quick answer

A vector database stores embeddings with metadata and quickly finds the vectors nearest to a query, using approximate nearest neighbour indexes such as HNSW or IVF. It powers semantic search and RAG retrieval. If your data already lives in Postgres and scale is moderate, pgvector is often enough; dedicated vector databases suit large scale or advanced vector features; search engines suit teams that need strong keyword and hybrid search. Choose by filtering needs, scale, hybrid search, operations and cost, tested on your own data.

Where This Fits

Vectors come from embedding models. Retrieval quality usually improves with hybrid search and reranking. The full pipeline is in the RAG guide, and product search use is in ecommerce semantic search.

How Vector Search Works

Each item (a document chunk, product or image) is stored as a vector: a list of numbers produced by an embedding model. A query is embedded the same way, and the database returns the stored vectors with the smallest distance (cosine, dot product or Euclidean). Comparing against every vector is exact but slow at scale, so databases build approximate nearest neighbour (ANN) indexes that find near-best matches quickly.

Index typeHow it worksTrade-offs
Flat (exact)Compare with every vectorExact, slow at scale
HNSWMulti-layer proximity graphFast and accurate; more memory, slower builds
IVFCluster vectors, search nearest clustersLess memory; needs training and tuning
QuantizationCompress vectorsLower memory and cost; some accuracy loss

Filtering and Multi-Tenancy

Real queries filter: this customer's documents, this product line, content the user may see. With approximate indexes, filters applied after the index scan can return too few results. Systems handle this differently: pre-filtering, filtered index traversal or scanning further. pgvector added iterative index scans in version 0.8.0 for this reason. For multi-tenant products, decide between shared indexes with tenant filters and separate indexes per tenant based on isolation needs and scale.

Choosing the Type of System

The best vector store is often the one that fits the data you already run.

Postgres and pgvector

pgvector adds vector types and indexes to Postgres. You get transactions, joins with business tables, SQL filtering and your existing backups and operations. It supports HNSW and IVFFlat indexes, half-precision vectors (halfvec) for smaller indexes and iterative scans for filtered queries. It suits many RAG and semantic search systems; at very large scale or with heavy query loads, tuning and dedicated infrastructure become more important.

Choosing a vector store for your AI application?

ZSpace Labs can benchmark pgvector, dedicated vector databases and search engines on your own data and queries before you commit.

Start a Project

Dedicated Vector Databases and Search Engines

Dedicated vector databases such as Qdrant, Pinecone, Weaviate and Milvus focus on vector workloads: scaling, filtering, quantization and often hybrid search with sparse vectors. Search engines such as Elasticsearch and OpenSearch combine mature keyword search with vector search and rank fusion, which suits content-heavy retrieval. Each adds a system to operate or a managed service to pay for.

Selection Criteria

  • Where your source data already lives
  • Number of vectors now and in two years, and dimensions
  • Filtering complexity and permission requirements
  • Need for keyword or hybrid search
  • Latency and throughput targets
  • Managed service versus self-hosting, and data residency
  • Team familiarity and operational tooling
  • Total cost including replicas and re-indexing

Operations and Cost

Plan for backups, re-indexing when embedding models change, index rebuild times, memory usage (HNSW indexes can be large), replicas for availability and monitoring of recall and latency. Reduce cost with smaller embedding dimensions where quality allows, quantization, and removing stale vectors.

Advantages and Limitations

AdvantagesLimitations
Fast semantic search at scaleApproximate results; recall must be measured
Find similar items without exact termsWeak on exact codes without hybrid search
Metadata filtering and multi-tenancyFiltering and ANN interact in tricky ways
Managed options reduce operationsAnother system and cost to manage

How to Choose Step by Step

  • 1. Write down data size, filters and latency needs
  • 2. Shortlist pgvector, one dedicated database and one search engine
  • 3. Load a realistic sample with your embeddings and metadata
  • 4. Run your real queries and measure recall, latency and filtered results
  • 5. Estimate cost at expected scale
  • 6. Decide, documenting re-indexing and backup plans

Example: Vector Search in Postgres With pgvector

For teams already on Postgres, a minimal setup looks like this. Check the pgvector documentation for current syntax and tuning parameters.

Example: pgvector table, HNSW index and filtered query (illustrative SQL)
CREATE EXTENSION IF NOT EXISTS vector;

CREATE TABLE doc_chunks (
  id bigserial PRIMARY KEY,
  tenant_id uuid NOT NULL,
  source_id text NOT NULL,
  content text NOT NULL,
  embedding vector(1024) NOT NULL
);

CREATE INDEX ON doc_chunks USING hnsw (embedding vector_cosine_ops);
CREATE INDEX ON doc_chunks (tenant_id);

-- with filtered queries, consider iterative index scans (pgvector 0.8.0+)
SET hnsw.iterative_scan = relaxed_order;

SELECT id, source_id, content
FROM doc_chunks
WHERE tenant_id = $1
ORDER BY embedding <=> $2   -- cosine distance to the query embedding
LIMIT 20;

Sizing and Capacity Planning

Estimate vectors (chunks per document times documents, including growth), dimensions and precision to size storage and memory. Graph indexes such as HNSW perform best when they fit in memory. Reduce footprint with shorter embeddings where quality allows, half-precision or quantized vectors, and removing stale content. Plan capacity for re-indexing, which can temporarily double storage, and for query peaks. Load-test with realistic filters, not just unfiltered nearest-neighbour queries.

Worked Example

An illustrative scenario, not a client case: a SaaS company adds document Q&A for customers. Its data is already in Postgres, with a few million chunks across tenants. pgvector with an HNSW index, tenant filters and iterative scans meets latency targets in testing, avoiding a new system. The team documents a threshold at which it would revisit a dedicated vector database.

Common Mistakes

  • Choosing from benchmarks that do not match your filters
  • Ignoring filtered-query recall
  • No plan for re-embedding
  • Vector-only retrieval for content full of identifiers
  • Over-provisioning dimensions and replicas

Need retrieval infrastructure that scales sensibly?

Talk to ZSpace Labs about RAG and vector search development and database and backend architecture.

Start a Project

Conclusion

Vector databases make semantic retrieval fast, but the right choice depends on your data, filters and team. Start with what you already run, measure filtered recall and add specialised systems only when needed. Related: vector embeddings, hybrid search and RAG.

FAQ

Common questions

A database designed to store vector embeddings and find the vectors most similar to a query vector quickly, usually with approximate nearest neighbour indexes, along with metadata filtering.

Get in touch

Have a project in mind?

Whether you're building a new digital product, improving an existing website, or looking to automate part of your business — let's talk.