Weaviate: The Vector Database Powering AI Search

Weaviate: The Vector Database Powering AI Search

Type "charged twice for one order" into a typical help-center search box and you will probably get zero results. The answer exists. It sits in an article titled "Duplicate payment refunds," but none of your words appear in that title, so a keyword search engine has nothing to match.

That gap between what people type and what documents say is the reason vector databases exist. Weaviate is one of the better-known ones. It is open source, it started in Amsterdam in 2019, and it now runs search behind chatbots, product catalogs, support portals, and internal knowledge tools. Its core idea is searching by meaning instead of by exact words.

This guide explains how the Weaviate vector database works in plain terms, where it earns its keep, and where it gets awkward. The awkward parts deserve the most attention. Missing data, results that disagree with each other, tight response-time budgets, odd edge cases, and memory bills that climb with every million records are what decide whether a search project ships or stalls.

What Weaviate actually does

A regular database is good at exact questions. "Show me orders over $500 from March" has one correct answer, and SQL finds it fast. It struggles with fuzzy questions like "find support tickets that sound like this angry email," because there is no column for "sounds like."

Every record in the Weaviate vector database holds two things side by side. One is the record itself: the text, price, category, date, or any other fields you define. The other is a long list of numbers called a vector, which describes what the record means.

Embeddings, explained without the math

Those number lists come from an embedding model, a type of AI model trained to turn text, images, or audio into coordinates. Think of a giant map where every sentence gets a pin. Sentences with similar meaning land near each other. "Charged twice" and "duplicate payment" end up as neighbors even though they share no words. The word "bank" in "river bank" and in "bank loan" lands in two different neighborhoods, because the model reads the surrounding words.

These AI embeddings usually contain somewhere between 384 and 3,072 numbers each, depending on the model. OpenAI's text-embedding-3-small, for example, produces 1,536 by default. You never need to read the numbers yourself. What matters is that closeness on the map means closeness in meaning.

When someone searches, Weaviate turns the question into a vector using the same model, then finds the stored vectors nearest to it. That is vector search in one sentence: find the closest pins.

Where the embeddings come from

You have two options. You can create embeddings in your own code and hand Weaviate the finished numbers. Or you can connect an embedding provider through a module (a plug-in), and Weaviate calls that provider every time you add a record or run a query. Modules exist for OpenAI, Cohere, Google, Voyage AI, Jina AI, Hugging Face, AWS, and others. Weaviate also sells its own hosted service, Weaviate Embeddings.

The module route means less code. Bringing your own vectors gives you more control and makes it easier to switch providers later. Teams building a quick prototype usually start with modules. Teams with strict data rules, such as a hospital that cannot send patient notes to an outside API, often generate embeddings on their own servers.

PRO TIP

Pick your embedding model before you load serious data. Every stored vector is tied to the model that created it, and switching later means re-embedding every record. At a few million records, that is a real bill and a lost week.

How a single search travels through the system

Following one query from start to finish makes the moving parts easier to see. Say an online furniture store runs on Weaviate and a shopper types "cozy chair for a small reading nook."

1.      The text goes to the embedding model, which returns a vector for the phrase.

2.      Weaviate looks for the closest product vectors using an index called HNSW (Hierarchical Navigable Small World). It works like a layered road map. The top layer has a few highways linking distant regions, and lower layers add local streets. The search starts on the highways, gets close quickly, then drops to side streets for the last stretch. That is how it can search millions of items in milliseconds without comparing the query against every single one.

3.      Filters get applied, such as "in stock" or "under $300." Weaviate keeps a separate, traditional index of your fields for this, so meaning and hard rules work together.

4.      If you asked for hybrid search, a keyword search also runs. It uses BM25, a long-standing ranking formula that rewards rare words appearing often in a document. The two result lists are then merged.

5.      Optional steps run last. A reranker model can re-sort the top results with a slower, more careful read, and a generative module can pass them to a large language model that writes an answer. That final pattern is called retrieval-augmented generation, or RAG.

 

One detail catches people off guard. HNSW is approximate. It finds very close neighbors, and usually the closest, but it does not promise the mathematically perfect top ten every time. A setting called ef controls how many candidates get examined during a search, which lets you trade a little speed for a little more accuracy. For product search, finding 98 of the true top 100 is usually fine. For legal discovery, where one missed document can matter, teams raise those settings and test carefully.

Why hybrid search matters more than demos suggest

Pure meaning-based search has a blind spot. It handles ideas well and exact strings badly. Search for part number "HX-4471B" and an embedding model may cheerfully return "HX-4417B," because the two look almost identical to it. A customer who needs that exact part will not be pleased.

Hybrid search runs keyword matching and meaning matching together and blends the scores. In Weaviate, a setting called alpha controls the blend. An alpha of 1 uses only vectors, 0 uses only keywords, and the default of 0.75 leans toward meaning while still giving exact words some pull.

Kind of query

Example

Alpha to start testing

Product codes, SKUs, error codes

"ERR_CONN_RESET 502"

0.2 to 0.4

Natural language mixed with names

"Priya's onboarding checklist"

0.5

Descriptive, conversational

"chair for a small reading corner"

0.75 to 0.9

Vague intent or mixed languages

"like the blue lamp but cheaper"

0.9 or higher

 

These are starting points. Your own query logs should settle the final numbers.

This mix is where Weaviate works well as a semantic search database for real traffic. Plenty of searches blend both kinds of intent, and running two separate systems to cover them means double the infrastructure plus a messy merge step in your own code.

The market numbers, and why they disagree

Vector databases went from research tooling to standard AI infrastructure in roughly three years. Analysts agree the market is growing fast. They do not agree on its size.

Research firm

Base figure

Forecast

Annual growth

MarketsandMarkets (2025 report)

$2.65B in 2025

$8.95B by 2030

27.5%

The Business Research Company (2026 report)

$3.02B in 2025

$8.71B by 2030

23.6%

Grand View Research

$1.7B in 2023

$7.3B by 2030

23.7% (2024 to 2030)

Global Market Insights

$2.55B in 2025

Not stated in summary

22.3% (2026 to 2034)

The spread comes from definitions. Some firms count only purpose-built products like Weaviate, Pinecone, and Qdrant. Others include vector features inside general databases such as PostgreSQL, MongoDB, and Oracle, and some add consulting services on top. Any single figure is best read as a rough direction.

Demand is easier to pin down. Menlo Ventures' November 2024 survey of 600 companies found that 51% of enterprises were using RAG, up from 31% a year earlier. Every RAG system needs somewhere to store and search embeddings, which is the job Weaviate was built for. The company raised a $50 million Series B in April 2023, led by Index Ventures with Battery Ventures participating.

FIELD NOTE

Market reports are useful for pitch decks and not much else. For a build decision, your own numbers matter more: how many records you have, how fast they grow, how many searches per second you expect, and how quickly results must come back.

The hard parts: what happens when real data arrives

Demos run on clean data. Production never does. The problems in this section tend to show up in month two, right after the prototype impressed everyone.

Data gaps

Real records are incomplete. A product has a title and no description. A support article has a heading and a screenshot but almost no text. An employee directory entry lists a name and nothing else. Each of these causes a different problem.

▪       Thin text produces weak vectors. An embedding built from "Blue Chair" carries far less meaning than one built from a full paragraph, so those records either vanish from results or show up for anything vaguely chair-shaped. A common fix is to build each vector from several fields at once, such as title, category, description, and tags. Weaviate lets you choose which properties feed the vectorizer.

▪       Missing fields break filters without any warning. If 30% of your products have no brand value, a filter for brand "Acme" silently leaves out any Acme products that were never tagged. Weaviate can index empty values if you switch on the indexNullState setting, which is off by default, so you can at least find and count the gaps.

▪       General models miss specialist vocabulary. "MI" means myocardial infarction in a cardiology note and Michigan in a shipping log. For specialist fields, test a domain-tuned model or give keyword matching more weight.

▪       Edited records need fresh vectors. When someone changes a record's text, its AI embeddings have to be regenerated. Weaviate's vectorizer modules handle this on update. If you bring your own vectors, that job is yours, and forgetting it means search keeps matching the old wording.

Long documents create a quieter gap. Embedding models have input limits, so a 40-page policy gets split into chunks. A chunk that says "this does not apply to contractors" loses its meaning once it is cut away from the heading that explains what "this" is. Many teams now add the document title and section heading to the start of every chunk before embedding it. That small change often improves results more than switching to a pricier model.

Conflicting signals

A search result is the product of several opinions about relevance, and those opinions often disagree.

The keyword half of a hybrid query might rank a document first because it repeats the exact phrase, while the vector half ranks it twentieth because its overall topic is off. Weaviate merges the two lists with a fusion method. The default, relative score fusion, rescales each list's scores to a 0-to-1 range and adds them using your alpha weight. The older ranked fusion ignores raw scores and looks only at positions. Relative scoring keeps more information, but one extreme score can flatten everything else in its list. If results feel lopsided, trying the other method takes one line of code.

Business rules add a second layer. Relevance says the out-of-stock chair is the best match, while the merchandising team wants shoppers to see things they can buy. Weaviate 1.39, released in August 2026, made the Boost API generally available for this situation. A filter removes results; a boost only moves them up or down. You can favor in-stock and recently released items at 30% weight and let relevance keep the remaining 70%, so the perfect but unavailable chair slips a few places instead of disappearing.

Order matters when you stack these tools. A reranker runs after boosting and gets the final word, which means it can undo your business rules. Weaviate's release notes suggest using one or the other unless you want that layering on purpose.

WATCH OUT

The same release made MMR (maximal marginal relevance) generally available. MMR stops the first page from showing five near-copies of the same passage. Its balance setting defaults to 0.0, which means maximum variety and almost no weight on relevance after the first pick. Set balance yourself every time. Values between 0.3 and 0.7 are a reasonable range to test.

Real-time decisions

Most searches happen inside a live product with a person waiting. A shopping site might give search 100 milliseconds. A fraud check might get less. A few Weaviate behaviors matter a lot here.

Filters change the speed math. When a filter is very strict, say 200 matching items out of 5 million, walking the HNSW graph wastes time because most neighbors fail the filter. Weaviate switches to a direct scan of just the allowed items once the filtered set drops below a threshold called flatSearchCutoff, set to 40,000 objects by default. It also offers a filter strategy called ACORN, which moves through the graph more efficiently when filters and meaning point in different directions. When filtered queries are slow, these two settings are the first place to look.

New data is not always searchable the instant it arrives. By default, a new object is added to the vector index as part of the write. With asynchronous indexing switched on, writes return faster and the index catches up in the background, which helps large import jobs. The cost is a short window where a new item exists but does not yet appear in vector search results. A news archive can live with that. A marketplace where sellers expect their listing to appear the moment they hit publish should test it first.

Clusters with several copies of the data add a consistency choice. Each read or write can wait for confirmation from ONE node, a QUORUM (a majority), or ALL nodes. ONE is fastest and can briefly return slightly stale data. ALL is safest and slowest, and a single sluggish node then drags down every request. Most production setups pick QUORUM and accept a little extra delay in exchange for predictable answers.

When a query is slow and nobody knows why, the query profiling added in version 1.37 (still in preview) breaks the timing down shard by shard, so you can measure instead of guess.

Exceptions and edge cases

▪       Switching embedding models. Vectors from two different models cannot be compared, even when they have the same length. Mixing them in one index returns nonsense with no error message. Weaviate's named vectors let a single object hold several vectors, so you can add the new model's vectors next to the old ones, test, and move queries over once you are satisfied.

▪       Very short queries. A one-word search like "refund" gives the embedding model almost no context. Hybrid search with a lower alpha usually beats pure vector search for these.

▪       Negation. Embedding models are poor at understanding "not." A query for "laptops without touchscreens" can return touchscreen laptops, because the vector still leans toward both "laptop" and "touchscreen." Structured filters handle negation far better.

▪       Bulk deletes. Removing an object from an HNSW graph leaves a tombstone, a marker saying the entry is gone, and a background cleanup reconnects the graph later. Deleting millions of records at once can slow queries and use extra memory until that cleanup finishes.

▪       Settings you cannot reverse. Some choices are locked at creation time. In 1.39, the bit width of rotational quantization cannot be changed after you enable it, and that method pads dimensions to a multiple of 64. A 1,000-dimension model is stored as 1,024.

▪       Content in many languages. Many embedding models place "zapatos rojos" right next to "red shoes." Keyword search does not, and languages written without spaces, such as Japanese, Korean, and Chinese, need a matching tokenizer for BM25 to work at all. Weaviate ships tokenizers for these, but you have to choose them field by field.

Behavior under pressure and at scale

Founders should read this part twice, because this is where the costs live.

HNSW graphs perform best when the vectors sit in memory. Here is the arithmetic. One 1,536-number vector stored at full precision takes 6,144 bytes. Ten million of them take about 61 GB of RAM before you count the graph links, the operating system, or any spare room. That figure doubles if you keep two copies for safety.

Compression, called quantization, is how most teams bring that number down. It stores each number with fewer bits. Weaviate offers product, binary, scalar, and rotational quantization. Its 1.39 release notes put 4-bit rotational quantization, currently in preview, at 784 bytes per 1,536-dimension vector, about 7.84 times smaller than full precision.

Number of vectors

Full precision (1,536 dims)

4-bit rotational quantization

1 million

About 6.1 GB

About 0.8 GB

10 million

About 61 GB

About 7.8 GB

100 million

About 614 GB

About 78 GB

Raw vector storage only. Graph links, replicas, and headroom come on top.

Compression costs some accuracy, so Weaviate keeps the full vectors on disk and rescores. It finds candidates using the compressed copies, then rechecks the top few against the originals. Each query pays a small extra disk read, and in return most of the lost accuracy comes back.

Several other pressure points appear as data and traffic grow:

▪       Huge workloads that can wait. Version 1.36 introduced HFresh, an index type in technical preview that keeps most data on disk instead of in memory. Weaviate pitches it at applications that can accept responses in the hundreds of milliseconds rather than tens. Internal archives and batch analysis fit that profile, while search-as-you-type does not.

▪       Many customers on one cluster. SaaS products often give each customer an isolated slice of data, which Weaviate calls a tenant. Tenants can be marked inactive to free memory, or offloaded to cloud storage on self-hosted and enterprise setups (added in 1.26). The trade-off is wake-up time, since the first query after reactivation has to load that tenant back.

▪       Big imports. Loading tens of millions of objects one by one is slow and fragile. Batching groups them, and server-side batching, generally available since 1.36, lets the server set the pace so clients stop flooding it.

▪       Restarts. A large HNSW index used to rebuild itself on startup by replaying a log of every change ever made, which got slower as the log grew. Since 1.39, Weaviate takes snapshots automatically, loads a compact image at startup, and deletes old logs. Disk use rises briefly while a new snapshot is written, so keep spare space.

▪       Uneven shards. Weaviate splits collections into shards spread across machines. A query fans out to every shard and waits for all of them, so one overloaded machine slows every query that touches it. Even data distribution matters more than raw machine count.

PRO TIP

For a first memory estimate, multiply vectors by dimensions by 4 bytes, then roughly double it to cover the graph and overhead, as Weaviate's resource-planning guidance suggests. Multiply again by the number of copies you plan to keep. Then check how much quantization would cut.

How Weaviate compares with the alternatives

Weaviate is one of several good options, and the best fit depends on what you already run. The table below compares it with four common choices as of September 2026. Vendors in this space ship features every few months, so confirm details in current documentation before you commit.

 

Weaviate

Pinecone

Qdrant

Milvus

pgvector

What it is

Open-source database, self-hosted or managed cloud

Fully managed service

Open-source database, self-hosted or managed

Open-source database, managed via Zilliz

Extension for PostgreSQL

License

BSD-3-Clause

Proprietary

Apache 2.0

Apache 2.0

PostgreSQL License

Built-in embedding

Yes, many provider modules plus its own service

Yes, hosted models

Yes, FastEmbed and cloud inference

Yes, embedding functions

No, bring your own

Keyword plus vector

Native BM25 hybrid

Sparse and dense vectors

Sparse vectors with fusion

Native full-text search

Hand-built with Postgres full-text

Many customers

Native tenants with offloading

Namespaces

Payload-based partitioning

Partition keys

Rows and SQL rules

Strongest fit

Hybrid search, RAG, multi-tenant SaaS

Teams who want no servers to run

Speed-focused self-hosting

Very large self-hosted datasets

Small to mid data already in Postgres

The bigger practical split is between a dedicated vector database and adding vectors to a database you already run. pgvector inside PostgreSQL is simple when you have up to a few million vectors and your team knows Postgres. You keep one system, one backup routine, and ordinary SQL joins. The strain shows up with large indexes, heavy filtering, and hybrid ranking, which you end up building by hand. The Weaviate vector database starts paying for itself when search sits at the center of the product, when you want hybrid ranking and multi-tenancy without writing them yourself, or when data outgrows one Postgres server.

Who gets real value from it

For startup founders, the draw is speed to a working product. A RAG chatbot over your help docs, a "similar items" widget, or smarter search inside a SaaS app can go from idea to prototype in days, because embedding, storage, hybrid ranking, and answer generation all live in one place.

For content teams and marketers, the use is usually internal. A semantic search database built over years of blog posts, briefs, and research lets a writer ask "have we covered onboarding churn for fintech?" and get back the three relevant articles, even if none of them uses the word churn. It also surfaces near-duplicate pieces before they get published.

For office teams, think of the shared drive that nobody can search: HR policies, IT how-tos, old sales decks. Weaviate normally sits behind a tool someone builds, such as a Slack bot or an intranet search page, rather than being something staff use directly. Version 1.37 also previewed a built-in MCP server. MCP, the Model Context Protocol, is an open standard many AI assistants use to connect to data, so tools that support it can query Weaviate without custom glue code.

For developers, the practical pull is the set of official clients for Python, TypeScript, Go, and Java, a local Docker setup that mirrors production, and gRPC for fast queries. Version 1.39 added an experimental plain HTTP search API as well, handy for shell scripts and serverless functions where a full client library is overkill.

When it is the wrong tool

▪       Your data is small, in the tens of thousands of records, and already lives in PostgreSQL. pgvector will probably be enough.

▪       Nearly every search is an exact lookup by order ID, email address, or SKU. A regular database handles that more cheaply.

▪       Nobody on the team can own it. Running a distributed database yourself means upgrades, backups, and monitoring. If that is a stretch, use Weaviate Cloud or a fully managed competitor.

▪       You must explain every ranking to an auditor. Vector similarity is hard to justify line by line, whereas keyword rules with logged scores are easier to defend.

Getting started without losing a month

1.      Collect 50 to 100 real queries from support tickets, site search logs, or user interviews, and write down the result a person would expect for each. This test set is the most useful thing you will build.

2.      Start a free sandbox on Weaviate Cloud, or run Weaviate locally with Docker.

3.      Load a representative slice of 10,000 to 50,000 records using a vectorizer module, so you are not writing embedding code yet.

4.      Run every test query three ways: keyword only, vector only, and hybrid at a few alpha values. Count how often the expected answer lands in the top five.

5.      Make production decisions only after that: embedding model, quantization, number of copies, and whether to self-host.

 

That query set is what separates teams that tune with evidence from teams that tune on gut feel. Without it, every change seems like an improvement and nobody can prove otherwise.

PRO TIP

Budget for two bills. Weaviate Cloud or your own servers charge for storage and compute. Your embedding provider charges per token each time you add or edit a record, and each time a search query gets embedded. At high traffic, generating AI embeddings for queries can cost more than the database. Caching embeddings for repeated queries helps.

Key takeaways

1.  Weaviate stores each record next to a vector describing its meaning, so searches can match intent instead of exact wording.

2.  Hybrid search is the safer default for real users, especially when queries include codes, names, or jargon.

3.  Most production problems trace back to data: thin text, missing fields, poor chunking, and stale vectors.

4.  Memory drives cost. Quantization can shrink vector storage several times over, and rescoring wins back most of the accuracy.

5.  Build a small set of real test queries before you tune anything.

Wrapping up

Weaviate solves a specific problem well: finding things by what they mean. For a support portal, a product catalog, or a chatbot answering from your own documents, a good semantic search database changes what users experience. The shopper who types "cozy chair for a small reading nook" sees chairs, and the customer who was charged twice finds the refund article.

The hard engineering is rarely in the setup. It hides in the details this guide spent most of its time on: incomplete records, signals pulling in different directions, latency budgets, settings you cannot reverse, and memory math at scale. Teams that plan for those from the first week tend to ship on time. Teams that expect the Weaviate vector database to sort everything out alone tend to spend month three explaining strange results.

Start small, measure with real queries, and let the numbers choose your settings.

Nainesh Pandya

Nainesh Pandya

Nainesh is the marketing expert helping our clients and customers achieve success in terms of outreach and visibility. From understanding the complexities of value-chain and the impact of future technologies, Nainesh’s incredible understanding of digital marketing and online outreach helps create high-impact strategies.

Build Your Agile Team

We provide you with a top-performing extended team for all your development needs in any technology.

Hourly
$20
It Includes
Duration
Hourly Basis
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
25 Hours (MIN)
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile
Monthly
$2600
It Includes
Duration
160 Hours
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
1 Month
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile
Team
$13200
It Includes
Team Members
1 (PM), 1 (QA), 4 (Developers)
Communication
Phone, Skype, Slack, Chat, Email
Hiring Period
1 Month
Project Trackers
Daily Reports, Basecamp, Jira, Redmime, etc
Methodology
Agile

Frequently Asked Questions

Is Weaviate free to use?
The open-source version is free under the BSD-3-Clause license, and you can run it on your own servers or laptop. Weaviate Cloud, the managed service, offers a free sandbox for testing and paid plans for production. Your embedding provider bills separately unless you run your own model.
Do I need machine learning skills to use Weaviate?
No. With a vectorizer module, you send plain text and Weaviate handles the conversion into AI embeddings for you. Understanding the basics in this guide still helps, especially hybrid search and chunking, because those choices affect result quality more than most model settings.
How is Weaviate different from Elasticsearch?
Elasticsearch started as a keyword search engine and added vector features later. Weaviate started with vector search and added keyword ranking later. Both support hybrid search today. Teams already running Elasticsearch often extend it, while teams building AI features from scratch often find Weaviate's modules and RAG features quicker to set up.
Can Weaviate handle millions of records?
Yes. Weaviate splits collections into shards across multiple machines, so capacity grows as you add nodes. Budget is usually the tighter limit, since memory needs rise with the number and size of vectors. Plan quantization and sharding early instead of retrofitting them under load.
Is my data safe if I use a vectorizer module?
A vectorizer module sends your text to the provider you configure, such as OpenAI or Cohere, and that provider's data policies apply. If records cannot leave your environment, run an embedding model on your own servers and send Weaviate the finished vectors. Weaviate also supports role-based access control, so you can limit who reads or changes each collection.