AI Models & Platforms

Databricks Brings Full-Text and Vector Search to Lakebase Postgres

mm
Add Unite.AI to your preferred sources on Google

Databricks on September 28, 2026, introduced Lakebase Search, a built-in search engine for its Lakebase Postgres database delivered through two extensions: lakebasevector for approximate nearest neighbor search and lakebasetext for BM25 full-text search. Both extensions are generally available on AWS and Azure.

The extensions let developers run semantic, keyword, and hybrid search directly inside Postgres alongside operational data. Databricks said traditional OLTP systems were not built for the search demands of AI agents, which require low-latency, high-accuracy retrieval and often execute massive parallel searches, and that solving this until now meant attaching a standalone search engine to the primary database with an ETL pipeline. The company said it built Lakebase Search alongside feedback from hundreds of beta customers.

Benchmark Results and the Conexiom Deployment

Databricks said lakebase_vector delivers twice the throughput of the next-best system on the VectorDBBench 100M benchmark, which uses the LAION dataset, and that it is four times cheaper than a cloud Postgres vendor using pgvector, before additional savings from autoscaling. The company reported a P99 latency of 71 milliseconds at 97% recall, meaning the engine successfully retrieved the true nearest neighbors 97% of the time, and noted that pgvector and DiskANN were tested only on a single large instance.

Databricks also reported a customer result from Conexiom, which runs BM25 hybrid search over more than 100 million rows at half the compute footprint of its earlier pgvector setup. Conexiom’s infrastructure costs fell threefold and throughput rose fivefold relative to pgvector, the company said.

“Lakebase Search gives us a whole new level of scalability over pgvector, and unlocks BM25 in the same serverless database,” said Jordan Voves, an AI/ML architect at Conexiom. “We use Lakebase to connect data to our agents at scale.”

The Pgvector Limits Behind the Design

Databricks said pgvector is the most-installed extension in Lakebase Postgres, and it described three recurring pain points observed from customers running it at scale.

The first is cost that scales with data volume rather than usage. pgvector keeps its HNSW index in database memory, and because HNSW search relies on random-access graph traversal, performance falls by a factor of 10 to 50 once the index spills to disk and queries turn into chains of random reads. A 768-dimensional float32 vector takes about 3.3 kilobytes of memory once graph links and Postgres overhead are included, so an index over 100 million rows requires roughly 330 gigabytes of RAM to stay resident, provisioned in full whether queries touch it or not.

The second is index maintenance. When a build spilled to disk, a pgvector index took nearly 50 hours to construct on a standard cloud instance, Databricks said, and writes suffer from the same bottleneck because inserting a vector requires random-access traversal and modification of several graph layers. Since HNSW has no global rebalancing, recovering search quality means running a full REINDEX, an operation that locks the table and stops production writes.

The third is that a single query cannot be parallelized. A pgvector query is executed by one Postgres backend process, leaving the HNSW index scan without any parallelization. Raising recall requires visiting more graph nodes, which adds random memory reads and distance comparisons, inflating latency and cutting queries per second, so scaling throughput means adding database connections or read replicas.

How Lakebase_vector Is Built

Lakebase Postgres separates storage from compute: permanent data sits in low-cost cloud object storage, while RAM and local NVMe serve as short-lived caches that hold the active working set. On top of that foundation, Databricks combined two techniques.

Hierarchical IVF clustering groups vectors into clusters stored as contiguous blocks. A query scores the cluster centroids in memory, then reads only the handful of blocks that look promising as large sequential reads instead of making many random hops. Binary quantization, using the RaBitQ method, compresses each vector to about one bit per dimension, roughly 32 times smaller than float32, so queries scan compact codes to shortlist candidates and rerank only that shortlist against full-precision vectors.

Because the design is stateless, it scales to zero, and at rest users pay only for storage. Databricks reported a measured P90 of 1.13 seconds for the first query after scale-to-zero on a 100-million-vector, 768-dimensional dataset, and said 100 million vectors can be served on one Lakebase Compute Unit.

Index construction trains the cluster centroids a single time on a small random sample; each vector is then assigned to its nearest centroid, quantized, and written to the appropriate block as an independent operation, so the work fans out across however many cores are available. Databricks said its LTAP architecture offloads index builds from the primary database to distributed engines such as Spark, bringing build times down to minutes, and said more is coming on that capability. Because predicates are applied during the block scan itself, filtered queries avoid over-fetching candidates and keep recall high, and a single query parallelizes across CPU cores.

BM25 Text Search and Hybrid Queries

Databricks said standard Postgres tsvector search lacks corpus-wide relevance context. lakebase_text scores terms using global inverse document frequency, giving more weight to rare, high-intent terms and less to common filler words. It also checks score upper bounds as it traverses the index, discarding whole posting blocks that cannot affect the top-K outcome, which the company said makes it faster than tsvector with GIN indexes.

Combining the two extensions enables native hybrid search inside Postgres: a single query can apply ordinary SQL filter predicates, join live operational tables, and merge semantic vector scoring with BM25 keyword relevance. Databricks said the system scales from one row to one billion vectors and from one query per second to thousands without manual reprovisioning, describing the design as intended for AI agents whose workflows can trigger thousands of concurrent retrieval requests within seconds.

Databricks positions Lakebase Search for users who want operational and search data consolidated in a single database, while Databricks AI Search remains its managed engine for retrieval that works out of the box. Lakebase Search is generally available on AWS and Azure; existing Lakebase users can enable the extensions, and new users can sign up for Lakebase.

Theo Nash is an AI-generated specialist at Unite.AI, covering AI infrastructure, compute, and the hardware systems that power modern artificial intelligence. His work focuses on the technical foundations behind large-scale AI workloads, including data centers, accelerators, networking, and the software stacks that tie them together.

With an analytical and engineering-driven perspective, Theo examines how advances in GPUs, custom silicon, memory architectures, and distributed systems enable new generations of AI models. He pays particular attention to performance trade-offs, energy efficiency, scalability, and the practical constraints that shape real-world deployment of AI infrastructure.

Articles authored by Theo Nash are AI-generated and reviewed by Unite.AI’s editorial team to ensure technical accuracy, clarity, and responsible coverage of the rapidly evolving AI compute landscape.