Scale a Multi-Tenant RAG Platform with YugabyteDB

In the first three parts of this series, What Is RAG? From PostgreSQL pgrag to Distributed RAG with YugabyteDB pg_dist_rag, we progressively built up a distributed Retrieval-Augmented Generation (RAG) architecture with YugabyteDB.

We started by learning what RAG is and how YugabyteDB’s pg_dist_rag differs architecturally from earlier PostgreSQL RAG projects.

Then on part two, Build a Distributed RAG Pipeline from YSQL with pg_dist_rag, we built a complete distributed preprocessing pipeline:

Documents
|
v
pg_dist_rag
|
v
Work Queue
|
v
RAG Workers
|
v
pgvector

In Part 3, Build a Multi-Tenant Fraud Investigation Assistant with YugabyteDB pg_dist_rag, we made the architecture multi-tenant and used a fictional fraud investigation application to demonstrate how the same vector knowledge base could contain data from multiple financial institutions while retrieval remained tenant-aware.

Now comes the architecture question that eventually appears in every successful multi-tenant platform…

What happens when:

  • 3 tenants becomes 30 tenants becomes 300 tenants becomes 3,000 tenants?

This final tip focuses less on another pg_dist_rag function and more on the decisions that determine how well a multi-tenant RAG platform will scale.

Those decisions include:

  • ● Data isolation
  • ● Security isolation
  • ● Vector index design
  • ● Worker scaling
  • ● Noisy-neighbor protection
  • ● Metadata strategy
  • ● Geographic placement
  • ● Data residency
  • ● Operational complexity

YugabyteDB provides several different isolation boundaries, and choosing the right one can be just as important as choosing the embedding model.

There Is More Than One Kind of Multi-Tenancy

In Part 3, our architecture looked like this:

MultiTenant-RAG-Vectors

All three tenants could share:

  • ● YugabyteDB cluster
  • ● YSQL database
  • ● pgvector table
  • ● HNSW index
  • ● RAG worker infrastructure

while tenant_id identified ownership of each vector.

That provides a high degree of consolidation.

But it is not the only possible architecture.

As the platform grows, we need to decide where the tenant boundary should live.

Four Useful Isolation Patterns

I would think about multi-tenant RAG using four increasingly isolated patterns.

Pattern Tenant Boundary Consolidation Isolation
Shared Vector Index tenant_id Highest Logical / row level
Separate Vector Index Table / HNSW index High Data structure
Separate Database YSQL database Moderate Database + CPU
Separate Cluster YugabyteDB cluster Lowest Strongest infrastructure boundary

There is no universally correct choice.

The right design depends on the workload, security requirements, tenant size, regulatory requirements, and how much operational isolation each tenant requires.

Pattern 1: One Shared Vector Index

This is what we built in Part 3.

Fraud)kb-2-HNSW
The backing table contains:
  • ● chunk_text
  • ● embeddings
  • ● document_id
  • ● tenant_id
  • ● metadata_filters

pg_dist_rag carries the optional tenant_id supplied when a source is registered into the generated vector rows. It also stores JSONB metadata that can be combined with vector similarity searches.

A query might look like:

				
					SELECT
    chunk_text,
    metadata_filters,
    embeddings <=> :query_embedding AS distance
FROM public.fraud_kb
WHERE tenant_id = :tenant_id
ORDER BY embeddings <=> :query_embedding
LIMIT 10;
				
			

This architecture maximizes consolidation:

  • ● One table.
  • ● One HNSW index.
  • ● One pipeline.
  • ● Potentially thousands of tenants.

But tenant_id Is Not a Security Boundary by Itself

This point from Part 3 becomes even more important as the platform scales.

Consider:

				
					WHERE tenant_id = :tenant_id
				
			

If the application forgets that predicate, the database has not automatically prevented cross-tenant retrieval.

For shared-table architectures, database-enforced Row-Level Security (RLS) can provide an additional defense.

YugabyteDB supports PostgreSQL-style RLS policies that restrict which rows users can retrieve or modify.

For example, a simplified application-controlled pattern might use:

				
					ALTER TABLE public.fraud_kb
ENABLE ROW LEVEL SECURITY;

ALTER TABLE public.fraud_kb
FORCE ROW LEVEL SECURITY;
				
			

Then:

				
					CREATE POLICY fraud_kb_tenant_policy
ON public.fraud_kb
FOR SELECT
USING (
    tenant_id =
    NULLIF(
        current_setting('app.tenant_id', true),
        ''
    )::uuid
);
				
			

The application establishes the tenant context before retrieval:

				
					BEGIN;

SET LOCAL app.tenant_id =
    '11111111-1111-4111-8111-111111111111';

SELECT
    chunk_text,
    embeddings <=> :query_embedding AS distance
FROM public.fraud_kb
ORDER BY embeddings <=> :query_embedding
LIMIT 10;

COMMIT;
				
			

Now the tenant restriction can be enforced independently of whether the application remembered to put tenant_id in the query.

RLS Does Not Replace Authentication

The application must still authenticate the user and establish a trustworthy tenant identity. A custom GUC such as app.tenant_id is appropriate only when the application controls the database session and sets the tenant context from trusted authentication data. Users with direct database access should not be allowed to set an arbitrary tenant context. Also remember that superusers, table owners unless RLS is forced, and roles with BYPASSRLS can bypass row-level policies.

Using a Dynamic Tenant GUC with RLS?

If your application uses a custom GUC such as app.tenant_id to establish tenant context, see
Safely Use Dynamic Tenant GUCs with RLS and YSQL Connection Manager.
That tip shows how to use SET LOCAL with a fail-closed RLS policy that safely handles missing or reset tenant values, including when YSQL Connection Manager is enabled.

For stricter implementations, tenant identity can instead be associated with database roles or another trusted authorization mechanism.

There Is an Important HNSW Detail

There is another reason to think carefully about a very large shared vector index.

YugabyteDB supports HNSW approximate nearest-neighbor search using the ybhnsw index access method.

It also supports combining ANN vector search with a SQL WHERE clause.

However, YugabyteDB’s pgvector documentation makes an important point:

  • The WHERE filter is applied after the HNSW index scan.

That matters for multi-tenancy.

Imagine

  • 10,000,000 total vectors but Tenant A owns only 5,000

The HNSW search discovers approximate nearest candidates across the index and then the tenant predicate removes rows that do not belong to Tenant A.

As the number of tenants increases, heavily filtered searches should therefore be tested carefully for both:

  • Recall and Latency
Shared HNSW Indexes Need Realistic Testing

A shared vector table may be operationally simple, but do not benchmark only unfiltered vector searches. Test realistic tenant and metadata predicates using the actual tenant-size distribution expected in production.

One tuning control is:

				
					SET hnsw.ef_search = 200;
				
			

Increasing hnsw.ef_search considers more HNSW candidates during retrieval, trading additional work for potentially better recall. YugabyteDB exposes hnsw.ef_search as a query-time HNSW tuning parameter.

But tuning is not always the answer.

At some point, separating tenants into different indexes may be simpler.

Pattern 2: Separate Vector Indexes

Instead of:

fraud_kb

Tenant A + B + C + D

we could create:

fraud_kb_bank_a

fraud_kb_bank_b

fraud_kb_bank_c

Each pg_dist_rag vector index has its own backing table and HNSW index.

The architecture becomes:

multi-tenant-multi-hnsw

There are several advantages:

  • ● Tenant queries never compete with another tenant’s vectors inside the same ANN index.
  • ● An index can be rebuilt independently.
  • ● Different tenants can potentially use different document sources or embedding configurations.
  • ● Deleting an entire tenant becomes operationally simpler.
  • ● And very small tenants are no longer sharing an enormous HNSW graph with very large tenants.

The cost is more database objects and more lifecycle management…

  • ● 10 tenants = 10 vector tables
  • ● 1,000 tenants = 1,000 vector tables

That may or may not be desirable.

Shared Index or Separate Index?

A useful rule of thumb is:

Requirement Consider
Many small tenants with similar workloads Shared vector index
Very different tenant sizes Separate indexes for large tenants
Strong per-tenant vector isolation desired Separate vector indexes
Minimal schema objects Shared vector index
Independent rebuild / lifecycle Separate vector indexes

Hybrid designs are entirely reasonable.

Perhaps 950 small tenants share one index while the 50 largest customers receive dedicated indexes.

Pattern 3: Separate YSQL Databases

The next isolation boundary is particularly interesting in YugabyteDB 2026.1.

Instead of one database containing every tenant:

  • ● rag_platform
    • ○ Tenant A
    • ○ Tenant B
    • ○ Tenant C

we could use multiple database:

  • ● rag_bank_a
  • ● rag_bank_b
  • ● rag_bank_c

or perhaps group tenants into several databases:

  • ● rag_small_tenants
  • ● rag_enterprise_tenants
  • ● rag_regulated_tenants

Why might that matter?

Because YugabyteDB 2026.1 introduced the Resource Governor for Multitenancy as an Early Access feature.

The Resource Governor treats each YSQL database as a tenant and uses Linux cgroups to provide CPU isolation between databases.

During CPU contention, active databases receive fair CPU weighting so one database cannot monopolize the node. Optional CPU caps can also be configured.

That changes our isolation model.

Multi-Tenant-DB-per-Tenant
Important Resource Governor Boundary

Resource Governor isolation is per YSQL database, not per tenant_id row. If Bank A and Bank B share the same database and vector table, Resource Governor does not provide separate CPU buckets for those two tenants.

This gives architects an interesting choice.

tenant_id provides very high consolidation.

A database boundary provides stronger workload isolation.

Resource Governor Example

After the required Linux cgroup setup, Resource Governor is enabled using YugabyteDB flags.

For example:

  • ● enable_qos=true
  • ● qos_max_db_cpu_percent=25
  • ● qos_system_high_cpu_reserved_percent=5
  • ● qos_max_db_count=20

The YugabyteDB documentation describes qos_max_db_cpu_percent as a per-database maximum CPU percentage and qos_system_high_cpu_reserved_percent as CPU reserved for high-priority database system work.

A yugabyted example might resemble:

				
					./bin/yugabyted start \
  --tserver_flags="enable_qos=true,qos_max_db_cpu_percent=25,qos_system_high_cpu_reserved_percent=5" \
  --master_flags="enable_qos=true,qos_max_db_cpu_percent=25,qos_system_high_cpu_reserved_percent=5,qos_max_db_count=20"
				
			

Resource Governor is currently Early Access and has important limitations.

It governs CPU, not:

  • ● memory
  • ● disk I/O
  • ● network

and it provides performance isolation rather than additional security isolation.

Still, for a consolidated RAG platform, CPU isolation can be extremely useful.

Consider the Noisy-Neighbor Problem

Imagine:

  • Bank A
    5,000 vectors
    10 searches/minute
  •  
  • Bank B
    2,000,000 vectors
    500 searches/minute
  •  
  • Bank C
    50,000 vectors
    20 searches/minute

Then Bank B launches a new AI application.

Suddenly:

  • 5,000 searches/minute

If all three tenants share one database, that demand is simply part of the same YSQL database workload.

If Bank B lives in a separate Resource-Governor-managed database, YugabyteDB has a database-level CPU boundary available during contention.

That can turn this:

  • One customer gets busy -> Everybody slows down

into:

  • One customer gets busy -> Database CPU boundaries -> Other tenants retain CPU access

That is a compelling workload-consolidation story.

Pattern 4: Separate Clusters

Some tenants simply should not share infrastructure.

Reasons might include:

  • ● Regulatory requirements
  • ● Contractual isolation
  • ● Extreme workload size
  • ● Independent upgrade schedules
  • ● Dedicated operational ownership
  • ● Very different geographic requirements
  • ● Strict data residency boundaries

In those situations:

Tenant-per-Cluster

may be the appropriate architecture.

You lose some consolidation efficiency, but gain the strongest infrastructure boundary.

The goal should not be:

  • Put every customer into the same database at any cost.

The goal should be:

  • Use the smallest isolation boundary that satisfies the tenant’s requirements.

Scale the Workers Independently

So far we have focused on query-time architecture.

But RAG has another scaling dimension:

  • Ingestion

Fortunately, pg_dist_rag workers are deliberately decoupled from the YugabyteDB database processes.

Workers claim different documents from the lease-based dist_rag.work_queue, and adding workers increases preprocessing parallelism without changing the YugabyteDB cluster configuration.

The flow looks like this:

YB-Work-Queue

You can therefore have something like:

  • ● 10 TEXT workers
  • ● 3 PDF workers
  • ● 0 workers overnight

and scale based on the type and volume of content being processed.

Workers can also run on separate infrastructure from YugabyteDB, allowing preprocessing compute and database compute to scale independently.

What Happens If a Worker Fails?

The work queue uses lease-based locking.

The current worker configuration exposes TASK_LEASE_DURATION with a documented default of 600 seconds.

Workers also have polling and error-backoff controls.

Pipeline state includes statuses such as:
  • ● QUEUED
  • ● PROCESSING
  • ● COMPLETED
  • ● FAILED
  • ● RETRY

and pg_dist_rag exposes pipeline views that allow failed documents to be identified and retried from SQL.

That means worker failure does not require us to invent a separate distributed-task framework.

The database remains the source of truth for pipeline state.

Watch the Pipeline, Not Just the Workers

At scale, do not rely solely on:

  • Is the Python process running?

The more useful question is:

  • Is the pipeline making progress?

Use:

				
					SELECT
    index_name,
    document_name,
    pipeline_status,
    chunks_processed,
    embeddings_persisted,
    current_step,
    last_error_message
FROM dist_rag.vector_index_pipeline_details;
				
			

And:

				
					SELECT
    index_name,
    document_name,
    calls,
    total_chunks_processed,
    total_embeddings_persisted,
    completion_rate_percent
FROM dist_rag.pipeline_stats;
				
			

Those views expose per-document and aggregate pipeline state directly from YugabyteDB.

Metadata Design Matters More as You Scale

In Part 3 we used:

				
					{
  "institution": "Bank A",
  "knowledge_type": "historical_case"
}
				
			

At enterprise scale, metadata becomes part of the retrieval architecture.

You might add fields such as:

				
					{
  "institution": "Bank A",
  "knowledge_type": "historical_case",
  "payment_rail": "ACH",
  "fraud_type": "account_takeover",
  "country": "US",
  "classification": "internal",
  "effective_year": 2026
}
				
			

Then retrieval can combine:

  • tenant identity + semantic similarity + document classification + business metadata

For example:

				
					SELECT
    chunk_text,
    embeddings <=> :query_embedding AS distance
FROM public.fraud_kb
WHERE tenant_id = :tenant_id
  AND metadata_filters @>
      '{
         "payment_rail":"ACH",
         "knowledge_type":"historical_case"
       }'::jsonb
ORDER BY embeddings <=> :query_embedding
LIMIT 10;
				
			

This is one of the reasons putting vectors in a SQL database is compelling.

The vector does not have to exist independently of the rest of the application’s data model.

Think Carefully About Data Residency

YugabyteDB supports row-level geo-partitioning and tablespaces to control where YSQL data is placed geographically. These capabilities can be used to lower latency and satisfy data residency requirements.

However, there is an important architectural distinction for RAG.

Having:

				
					WHERE tenant_id = Bank A
				
			

does not automatically mean:

  • Bank A vectors stay in Region A.

Tenant identity and physical data placement are separate concepts.

Multi-Tenancy Is Not Data Residency

A tenant_id controls logical ownership of a vector row. It does not by itself control where that row is physically stored. If tenants have different geographic or regulatory placement requirements, design the table, database, or cluster placement strategy accordingly and validate the resulting placement.

For example, tenants with strict regional requirements may be better candidates for:

  • ● separate indexes
  • ● separate databases
  • ● separate clusters

rather than one global shared vector table.

A Practical Tiered Architecture

Real platforms do not have to choose one multi-tenancy pattern for every customer.

A hybrid architecture can be more effective.

For example:

RAG-Practical-Tiered-Architecture

That could translate into:

Tenant Type Possible Architecture Why
Many Small Tenants Shared table + tenant_id Maximum consolidation
Large Tenant Dedicated vector index Independent HNSW and lifecycle
Performance-Sensitive Tenant Dedicated YSQL database Database-level CPU isolation with Resource Governor
Highly Regulated Tenant Dedicated database or cluster Stronger operational and placement boundaries

That is often a more realistic consolidation strategy than forcing every customer into exactly the same model.

The Architecture Decision Becomes a Spectrum

Rather than asking:

  • Shared or dedicated?

think of isolation as a spectrum:

Move a tenant farther down that spectrum only when its requirements justify the additional infrastructure.

Final Takeaway

The first three tips in this series concentrated on RAG itself. This final tip is really about platform architecture and how the pieces fit together as the number and size of tenants grow.

pg_dist_rag provides the distributed preprocessing pipeline, while pgvector provides vector storage and similarity search. YSQL adds relational filtering and security capabilities, all running on top of YugabyteDB’s distributed database architecture.

Those capabilities can then be combined with different isolation boundaries… from a shared table using tenant_id, to separate vector indexes, separate YSQL databases, or dedicated YugabyteDB clusters… depending on each tenant’s performance, security, regulatory, and operational requirements.

And those components can be combined at different isolation boundaries depending on what each tenant needs.

For small tenants, that could mean:

  • One shared vector table + tenant_id + metadata_filters + RLS

For larger tenants:

  • Dedicated HNSW index

For tenants that need performance isolation:

  • Dedicated YSQL database + Resource Governor

And for tenants that demand infrastructure or geographic isolation:

  • Dedicated cluster

The important point is that multi-tenancy does not need to be an all-or-nothing architecture.

A YugabyteDB RAG platform can consolidate aggressively where it makes sense and introduce stronger boundaries where workloads, security, performance, or regulatory requirements demand them.

That is what turns the simple RAG pipeline we built in Part 2 into something much more interesting:

  • a distributed, multi-tenant AI data platform.

Series Recap

PartTitleFocus
1What Is RAG? From PostgreSQL pgrag to Distributed RAG with YugabyteDB pg_dist_ragRAG fundamentals, embeddings, vector search, and why YugabyteDB uses a distributed preprocessing architecture.
2Build a Distributed RAG Pipeline from YSQL with pg_dist_ragBuild an end-to-end pipeline from S3 through parsing, chunking, embedding generation, pgvector, and HNSW search.
3Build a Multi-Tenant Fraud Investigation Assistant with pg_dist_ragCombine tenant_id, metadata filtering, and vector similarity to provide institution-specific fraud investigation context.
4Scale a Multi-Tenant RAG Platform with YugabyteDBChoose the right isolation boundary as the platform grows, from shared indexes to separate databases and dedicated clusters.

References

Have Fun!

Our daughter’s dog, Maple, loves her toys. Here she is playing with three of them all at once!