RAG architecture guide
RAG vs vector database: how the pieces fit.
Retrieval-augmented generation (RAG) is a workflow that gives a language model retrieved evidence before it generates an answer. A vector database stores and searches numerical representations of content. It can supply the retrieval step within that workflow.
The short answer
RAG describes the workflow; vector search supplies a component.
An embedding is a numerical representation produced by a model. Vector search compares those representations to find candidates that may be relevant even when the wording differs. The application must still decide which evidence it may use and how to use it.
| Question | Vector database | RAG workflow |
|---|---|---|
| What does it do? | Stores vectors and searches for similar ones, often with metadata filters. | Retrieves evidence, supplies context to a model, and generates a response. |
| What comes back? | Matches, identifiers, scores, and any associated text or metadata. | An application response informed by retrieved context. |
| Where do sources come from? | The ingestion process supplies the records. | The system needs a defined source collection and a process for preparing it. |
| What needs separate evaluation? | Search quality, latency, filtering behavior, and operating limits. | Evidence coverage, source support, permissions, freshness, and final-answer quality. |
A vector database product may bundle ingestion, embeddings, reranking, or answer generation. Check the actual service boundary: the product category alone does not establish which responsibilities are covered.
A fictional example
Follow a changed refund policy through the system.
Imagine a support application with an old 30-day refund policy and an approved replacement allowing 14 days. Both passages discuss the same topic, so similarity alone cannot decide which version the application should use.
- Prepare the evidence. Preserve each policy version, its effective date, source identity, and applicable audience before indexing passages.
- Retrieve within scope. Resolve the permitted policy version and access boundary, then search its passages. Keep the original question and requested filters available for inspection.
- Build context. Return the relevant passage with its source reference. If no approved evidence supports the question, handle that explicitly.
- Generate and check. Give the evidence to the application’s model. Verify that the answer reflects the applicable policy and that its citation supports the statement.
The database can help enforce filters, but someone must define and maintain those filters correctly. Replacing the database does not by itself resolve stale policy selection or an answer that contradicts a retrieved passage.
Choosing the retrieval layer
Does RAG require a dedicated vector database?
No. Retrieval can use keyword search, structured queries, vector similarity, or a combination. Choose using representative questions, the shape of the evidence, and measured operating requirements.
- Keyword or structured retrieval: useful when exact identifiers, names, dates, or relational conditions determine the relevant evidence. Evaluate paraphrased questions as well as exact matches.
- A database with vector support: useful when your existing data platform can meet the workload. PostgreSQL with pgvector is one option; its documentation describes combining vector and full-text search.
- A dedicated vector service: worth evaluating when its search capabilities, capacity, latency, filtering, and operating model fit your workload.
- A maintained retrieval service: worth evaluating when acquiring and reviewing the agreed sources is part of the work you need covered. Verify the source scope separately from the search engine.
There is no universal document-count threshold for switching. Compare your actual chunk count, vector dimensions, update rate, query concurrency, filters, and latency target, then include operations and evaluation in the cost review.
A practical acceptance test
Find the failing stage before changing the stack.
Record the expected source passages for a small set of real questions. Add a changed-source question, an unauthorized request, an exact identifier, a paraphrase, and a question with no supported answer.
| Observation | Inspect first |
|---|---|
| The needed source is absent | Collection coverage and ingestion completeness. |
| The source exists but is not retrieved | Passage boundaries, query handling, filters, embeddings, and ranking. |
| The returned source is obsolete or forbidden | Version selection, update activation, and authorization rules. |
| The passage is correct but the answer is wrong | Context assembly, prompts, model behavior, and answer evaluation. |
Use the retrieval metrics example to understand missing evidence and ranking. Measure the final answer separately; a high retrieval score does not establish correctness.
Where CorpusMesh fits
A maintained knowledge boundary for your application.
CorpusMesh maintains agreed external source collections and returns source-backed passages through its private-beta retrieval API. Your team keeps its model, prompts, answer generation, and user experience.
Inspect the RAG API contract, the reviewed source trace, and the source maintenance process. These show different parts of the boundary; the fixed public trace is not a live result for your workload.
For a provider-selection checklist and cost review, continue to the RAG as a service guide. Beta access requires approval for a named application and an agreed source scope.
Further reading
Check the underlying architecture and implementation references.
- Pinecone: retrieval-augmented generation, for the ingestion, retrieval, context, and generation stages.
- pgvector: hybrid search, for combining PostgreSQL full-text search and vector retrieval.
- AWS: choosing a vector database for RAG, for workload-based evaluation of vector stores.
Define the source boundary
Evaluate the evidence your application needs.
Bring your sources, update requirements, and representative questions.