Insights

Are Vector Databases Still a Separate Category? History and Market Data

9 min read#vector-db#pgvector#approximate-nearest-neighbor#hnsw#database-market

Who this is forEngineers and architects choosing a vector store for RAG or semantic search who want historical and market context before picking a database.

Teams building retrieval-augmented generation (RAG) or semantic search face an early architecture decision: buy a dedicated vector database, or add vector search to a database they already run. The choice affects cost, operational load, and how easily the system can change later. Vector search is not new. Its core algorithms date back to 1998, and the wave of standalone vector database companies arrived between 2019 and 2021. This article traces that history, examines DB-Engines rankings and funding data, and explains where the evidence suggests the category is heading. By the end, you should know which market signals are worth trusting, why Postgres and Elasticsearch are taking share from dedicated products, and the conditions under which a dedicated vector database still makes sense.

Summary

Vector search started with a 1998 academic paper on locality-sensitive hashing (LSH). It became a widely used library with FAISS in 2017, saw dedicated database companies appear between 2019 and 2021 (Pinecone, Weaviate, Milvus, and Qdrant), and was inflated by the RAG boom in 2023. From 2024 to 2026, vector search has been repositioned as an index or extension of existing databases rather than a separate product category. The top system in the DB-Engines vector category is actually the search engine Elasticsearch, with 96.02 points. The leading dedicated vector database, Pinecone, has 8.33 points, roughly one-twelfth of that. Thirty years of approximate nearest neighbor (ANN) research became a category in about 1.5 years, and then was absorbed back into a feature of relational databases and search engines.

Evolution diagram

Four stages of vector DB evolution: algorithms, libraries, dedicated DBs, and absorption into existing systems

Key data

1. Algorithm lineage (primary academic sources)

Year Algorithm Key paper / authors Significance
1998 LSH (Locality-Sensitive Hashing) Indyk & Motwani, STOC 1998 — “Approximate nearest neighbors: towards removing the curse of dimensionality” First algorithm with sublinear time for high-dimensional ANN
2003 IVF (Inverted File, for visual search) Sivic & Zisserman, ICCV 2003 — “Video Google” Applies the inverted index from text IR to visual vectors
2011 PQ (Product Quantization) Jégou·Douze·Schmid, IEEE PAMI 2011 Validated at a scale of 2 billion vectors; a standard for memory compression
2014~2016 HNSW (Hierarchical Navigable Small World) Malkov & Yashunin, Information Systems 2014/2018 Default index in most current vector databases

2. Product launch timeline

Year Event
March 2017 FAISS open-sourced (Meta/FAIR) — Douze, Guzhva, Deng et al. library
2019 Pinecone founded, Weaviate first release (SeMI Technologies), Milvus open-sourced (Zilliz)
April 20, 2021 pgvector 0.1.0 first release (Andrew Kane)
2021 Qdrant founded
November 2022 ChatGPT released, triggering an explosion in RAG demand
2023 Chroma released; Pinecone Series B, $100M at a $750M valuation (April); pgvector 0.5 with HNSW support (August)
2024 Stonebraker & Pavlo paper arguing the relational model wins: vectors should be an index and extension of RDBMS, not a separate category
December 2025 Milvus GitHub stars pass 40,000
2026 pgvector 0.9 (IVFFlat improvements, sparse vector support), accelerating the Postgres integration trend

3. DB-Engines vector DBMS ranking (May 2026, 25 systems tracked)

Rank System Score Original category
1 Elasticsearch 96.02 Search engine (vector feature added)
2 OpenSearch 20.42 Search engine
3 Couchbase 9.32 Document DB
4 Pinecone 8.33 Dedicated vector DB
5 Kdb 7.33 Time series
6 Milvus 6.33 Dedicated vector DB
7 DolphinDB 5.20 Time series
8 Qdrant 5.16 Dedicated vector DB
9 Weaviate 4.24 Dedicated vector DB
10 MarkLogic 3.95 Multi-model

→ The top system in the vector category is not a dedicated vector database. Elasticsearch’s score is about 12 times Pinecone’s, and about 4 times the combined score of the dedicated vector databases (8.33 + 6.33 + 5.16 + 4.24 = 24.06).

4. The 20 DB model categories DB-Engines tracks

Relational · Key-value · Document · Time Series · Graph · Search engines · Vector · Object oriented · RDF · Wide column · Multivalue · Spatial · Native XML · Event Stores · Columnar · Content stores · Navigational · Time Series Observability · Multi-Model · NoSQL

→ Vector is one of 20 categories. Its 25 tracked systems are one-fifth to one-tenth the size of Relational (150+), Document, and Key-value.

5. Funding of dedicated vector database companies (primary sources)

Company Cumulative funding Latest round Notes
Pinecone $138M Series B, $100M at a $750M valuation (April 2023, led by a16z) Revenue $26.6M (2024) → $14M (2025); headcount 200 → 127 (Crunchbase/Latka)
Zilliz (Milvus) $115M Series B-II, $60M (August 2022) Milvus 40k+ GitHub stars (December 2025)
Weaviate $68M Series B, $50M (2023) Over 1 million Docker pulls per month

→ Pinecone revenue down 47% YoY, a sign the bubble is deflating. Over the same period, pgvector and Elasticsearch expanded adoption.

6. Andy Pavlo and Michael Stonebraker’s position (CMU and MIT, 2024)

  • “Vector databases should end up as a type index and an extension to a lot of good stuff that already exists within database systems”
  • The June 2024 Stonebraker & Pavlo paper argues that the relational model has absorbed every challenger over 30 years, and that vectors will follow the same path: not a separate category, but a new index type in RDBMS.

Insights

1. “Vector DB” is not a new category; it is 30 years of ANN research turned into products

ChatGPT’s release in late 2022 made “vector DB” appear in the market as if it were a new technology. The algorithm lineage, however, starts with LSH in 1998. IVF (2003), PQ (2011), and HNSW (2014) were techniques already validated at billion scale in computer vision and information retrieval labs, and FAISS (2017) packaged them as a library. The founding of Pinecone, Weaviate, and Milvus in 2019 was the step of packaging a library as a database product. The 2022 RAG boom made the category visible all at once. Distribution channels, not technical novelty, were the key to forming the category.

2. The current leader is a search engine, not a dedicated vector database

The top two systems in the DB-Engines vector ranking are Elasticsearch and OpenSearch. Both were originally text search engines, and in 2022 and 2023 they added dense vectors with HNSW as formal features. Elasticsearch’s score of 96.02 is about 12 times that of Pinecone (8.33), the leading dedicated vector database. The reason is straightforward: adding a vector index column to a search infrastructure already in operation costs far less than running a separate database.

→ The real winners of the “vector DB market” are existing search and RDBMS vendors, not vector database companies.

3. Pinecone revenue down 47% YoY: category value is set by operating cost, not infrastructure

Pinecone became a symbol of the market after its Series B at a $750M valuation in 2023. By 2025, however, its revenue had fallen to $14M from $26.6M the previous year, roughly half, and its headcount dropped from 200 to 127 (Latka). The explanation is simple: vector search is not expensive enough to justify its own infrastructure. Creating one more index in Postgres or Elasticsearch costs far less than a separate SaaS subscription, and there is no operational burden of moving data to another system. As RAG becomes common, the value proposition of managing vector data separately weakens.

4. The pgvector integration trend: vectors are absorbed as one index type inside RDBMS

Andrew Kane built pgvector alone in April 2021 as a PostgreSQL extension. By 2026 it has become the central vehicle through which Postgres is integrated as the default data layer for AI applications. Major Postgres hosting vendors, including Supabase, Neon, and Tiger Data, all support it by default, as do the AWS, GCP, and Azure RDS services. At a scale of 1M vectors, pgvector HNSW equals or outperforms dedicated databases such as Qdrant, according to Supabase’s own benchmark. This is the same picture Stonebraker and Pavlo predicted in their 2024 paper.

Vector databases have not settled into a new category. They are being absorbed as a new index type in RDBMS or a dense field in search engines.

5. Comparison with other DB categories: the smallest, but fastest-growing, newer category

DB-Engines tracks 20 DB model categories. Compared with Relational (150+ systems), Document (60+), and Key-value (50+), Vector has 25 systems, making it one of the smallest newer categories. Its category score, however, grew fastest between 2022 and 2024.

This implies two things:

  • The workload scale clearly justifies a separate category. RAG, image search, recommendation, and anomaly detection are large enough that DB-Engines created a dedicated category for them.
  • But market value does not stay with dedicated companies. Elasticsearch and Postgres collect most of the score.

Vectors are quickly following the path Time Series took when it was absorbed into PostgreSQL and TimescaleDB. The independent vector database market that dedicated companies targeted in 2019 is closing. The category survives, but share within it is moving to existing database vendors.

6. Implications for practitioners in Korea

  • New RAG and AI search projects: Evaluate Postgres with pgvector first. At a scale of 1 million to 10 million vectors, it is sufficient. Adopting a separate vector database has weak return on investment relative to its operational burden.
  • Existing Elasticsearch or OpenSearch use cases: Extend them directly with a dense_vector field and an HNSW index. No separate infrastructure is needed.
  • SaaS products such as Pinecone: Limit them to early prototypes and traffic-spike cases, where zero-ops value is high. Review cost and lock-in before full production adoption.
  • Consulting and training content: Teaching vector databases as a standalone subject matches market demand less well than teaching how to add vector indexes to Postgres or Elasticsearch. This mirrors the n8n × SAP pattern: a capability added to an existing stack, not a separate tool.

Bottom line

The evidence supports a clear conclusion: vector search has become a capability of the databases and search engines teams already run, rather than a market that requires a separate product. The category is real, but most of its score sits with Elasticsearch and Postgres, and Pinecone’s revenue decline suggests that standalone vector services are under pressure. For a new RAG project, start with Postgres and pgvector, and use Elasticsearch if you already operate it. Consider a dedicated vector database mainly for early prototypes or workloads with sudden traffic swings, where zero-ops value is high, and review cost and lock-in before committing to it for production.

Sources

Primary academic sources

  • Indyk, P. & Motwani, R. (1998). “Approximate nearest neighbors: towards removing the curse of dimensionality.” STOC 1998. Citation link: https://people.csail.mit.edu/indyk/helsinki-2.pdf
  • Sivic, J. & Zisserman, A. (2003). “Video Google: A Text Retrieval Approach to Object Matching in Videos.” ICCV 2003.
  • Jégou, H., Douze, M., Schmid, C. (2011). “Product Quantization for Nearest Neighbor Search.” IEEE PAMI. PDF: https://arxiv.org/pdf/1102.3828
  • Malkov, Y. & Yashunin, D. (2014/2018). “Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs.” Information Systems. PDF: https://arxiv.org/pdf/1603.09320

Market and ranking sources

Product and vendor primary sources

Academic commentary

Benchmarks (secondary sources, for reference only)

Frequently asked questions

Is Elasticsearch the top vector database according to DB-Engines?
Yes. In the May 2026 DB-Engines Vector DBMS ranking, Elasticsearch scored 96.02, about 12 times Pinecone's 8.33. Pinecone was the top dedicated vector database but ranked fourth overall.
Should a new RAG project start with a dedicated vector database?
The note recommends evaluating Postgres with pgvector first, which it considers sufficient for roughly 1 million to 10 million vectors. Dedicated vector databases make more sense for early prototypes or workloads with sudden traffic swings, where zero-ops value is high.

Want the full system? The Claude Code & Codex Skills guidebook collects the skills and subagents behind this blog, from $19.