Architecture and buying guide
RAG as a service: choose the knowledge work you need managed.
Retrieval-augmented generation (RAG) gives a language model relevant source material before it produces an answer. RAG as a service moves some of that retrieval pipeline to a provider. The useful comparison is who maintains the sources, who retrieves the evidence, and who owns the final answer.
What is managed
Separate retrieval from answer generation.
A RAG workflow collects source material, prepares searchable passages, retrieves context for a question, and gives that context to a model. A service may manage several of these stages; its contract determines where its responsibility ends.
A retrieval API returns passages for your application to use. An answer service also calls a model and returns generated text. With either approach, evaluate the final application: a relevant passage alone does not prove that the answer is correct.
Compare approaches
Start with the source maintenance problem.
These are responsibility patterns, not a ranking of vendors. Products can combine them. Verify the features and exclusions in the specific service you evaluate.
| Approach | Managed work | Questions your team still owns |
|---|---|---|
| Managed vector database | Stores and searches vector representations of your content. | Who acquires sources, prepares passages, updates them, checks permissions, and evaluates retrieval? |
| Document RAG service | Ingests supplied files or connected repositories and prepares them for retrieval; some services generate answers. | Which connectors and source permissions are supported? Who verifies completeness and update delays? |
| Maintained knowledge API | Maintains an agreed source collection, versions, and retrieval evidence for a bounded workflow. | Does that collection cover your authorities, jurisdictions, languages, and permitted use? |
For an internal handbook assistant, document connectors and permission synchronization may decide the fit. For an application that cites changing external rules or curricula, source coverage, version history, and change review deserve their own acceptance criteria.
Evaluation checklist
Ask for evidence you can inspect.
Use the same questions and expected sources for every candidate. Include a valid request, a question outside the collection, an unauthorized filter, and a changed-source scenario.
| Check | Define first | Evidence to request |
|---|---|---|
| Source coverage | Name the authorities, documents, languages, jurisdictions, and excluded material. | Ask for an inventory with source URLs and the permitted reuse boundary. |
| Updates | Separate source-check frequency from the time a reviewed version becomes retrievable. | Ask what happens when a document changes, disappears, or is withdrawn. |
| Citations | Follow a returned passage to the exact source and version used. | Check the citation anchor, captured source identity, and version metadata. |
| Access | Test that a request cannot widen the permitted corpus or filters. | Include forbidden-source and cross-jurisdiction requests in acceptance tests. |
| Retrieval quality | Use questions from your workflow, including unsupported questions. | Inspect relevant passages, missing evidence, abstentions, and regressions separately. |
| Operating limits | Match request volume, result depth, concurrency, and response size to your application. | Read the API contract, retry behavior, quotas, and recovery process. |
Measure retrieval and generation separately. Retrieval recall asks whether the expected evidence was found; citation checks ask whether the reference supports the passage. Evaluate your model’s answers against those passages and test when the application should decline to answer.
The CorpusMesh acceptance worksheet provides a starting point. Its published benchmark covers one fixed EU AI Act workload and does not predict results for a different collection.
Cost and fit
Compare the same workload over the same period.
Record initial source preparation, recurring change review, indexing, retrieval requests, storage, evaluation, and recovery. Add integration time and your application’s model costs. Confirm which costs a provider includes and which remain yours.
A low per-query price does not tell you the cost of maintaining the sources. A fixed service fee does not tell you whether the collection meets your coverage or latency requirements. Compare both against a written workload and acceptance test.
Keeping the pipeline internal can be appropriate when the corpus is your proprietary advantage, information must remain inside a boundary the provider cannot enter, or you need processing the service cannot support. See the build-versus-buy worksheet for a like-for-like cost review.
A concrete example
CorpusMesh maintains knowledge; your application runs its model.
CorpusMesh focuses on approved external sources. Its private-beta retrieval contract returns ranked passages, citations, source versions, and an explicit retrieval decision. Your team retains the model, prompts, user experience, and final decisions.
Private beta The interfaces shown here define the intended product contract. Public access is not yet available.
For the published EU AI Act example, start with the source coverage and exclusions. Follow a passage through the reviewed source trace, inspect the RAG API request and response, then review source maintenance and the current beta scope and pricing.
The public trace is a reviewed example pinned to its source version. It is not a live query or a guarantee that your workflow will meet its acceptance criteria. Access is reviewed manually for an approved organization and named application.
Plan your evaluation
Bring the sources and questions your product depends on.
Describe your coverage, update requirements, and expected retrieval volume.