Architecture and buying guide

RAG as a service: choose the knowledge work you need managed.

Retrieval-augmented generation (RAG) gives a language model relevant source material before it produces an answer. RAG as a service moves some of that retrieval pipeline to a provider. The useful comparison is who maintains the sources, who retrieves the evidence, and who owns the final answer.

What is managed

Separate retrieval from answer generation.

A RAG workflow collects source material, prepares searchable passages, retrieves context for a question, and gives that context to a model. A service may manage several of these stages; its contract determines where its responsibility ends.

Preparesources · passages · index
Retrievequestion · filters · evidence
Generatemodel · prompt · answer

A retrieval API returns passages for your application to use. An answer service also calls a model and returns generated text. With either approach, evaluate the final application: a relevant passage alone does not prove that the answer is correct.

Compare approaches

Start with the source maintenance problem.

These are responsibility patterns, not a ranking of vendors. Products can combine them. Verify the features and exclusions in the specific service you evaluate.

Three ways to provide context to an AI application
ApproachManaged workQuestions your team still owns
Managed vector databaseStores and searches vector representations of your content.Who acquires sources, prepares passages, updates them, checks permissions, and evaluates retrieval?
Document RAG serviceIngests supplied files or connected repositories and prepares them for retrieval; some services generate answers.Which connectors and source permissions are supported? Who verifies completeness and update delays?
Maintained knowledge APIMaintains an agreed source collection, versions, and retrieval evidence for a bounded workflow.Does that collection cover your authorities, jurisdictions, languages, and permitted use?

For an internal handbook assistant, document connectors and permission synchronization may decide the fit. For an application that cites changing external rules or curricula, source coverage, version history, and change review deserve their own acceptance criteria.

Evaluation checklist

Ask for evidence you can inspect.

Use the same questions and expected sources for every candidate. Include a valid request, a question outside the collection, an unauthorized filter, and a changed-source scenario.

Checklist for reviewing a managed RAG service
CheckDefine firstEvidence to request
Source coverageName the authorities, documents, languages, jurisdictions, and excluded material.Ask for an inventory with source URLs and the permitted reuse boundary.
UpdatesSeparate source-check frequency from the time a reviewed version becomes retrievable.Ask what happens when a document changes, disappears, or is withdrawn.
CitationsFollow a returned passage to the exact source and version used.Check the citation anchor, captured source identity, and version metadata.
AccessTest that a request cannot widen the permitted corpus or filters.Include forbidden-source and cross-jurisdiction requests in acceptance tests.
Retrieval qualityUse questions from your workflow, including unsupported questions.Inspect relevant passages, missing evidence, abstentions, and regressions separately.
Operating limitsMatch request volume, result depth, concurrency, and response size to your application.Read the API contract, retry behavior, quotas, and recovery process.

Measure retrieval and generation separately. Retrieval recall asks whether the expected evidence was found; citation checks ask whether the reference supports the passage. Evaluate your model’s answers against those passages and test when the application should decline to answer.

The CorpusMesh acceptance worksheet provides a starting point. Its published benchmark covers one fixed EU AI Act workload and does not predict results for a different collection.

Cost and fit

Compare the same workload over the same period.

Record initial source preparation, recurring change review, indexing, retrieval requests, storage, evaluation, and recovery. Add integration time and your application’s model costs. Confirm which costs a provider includes and which remain yours.

A low per-query price does not tell you the cost of maintaining the sources. A fixed service fee does not tell you whether the collection meets your coverage or latency requirements. Compare both against a written workload and acceptance test.

Keeping the pipeline internal can be appropriate when the corpus is your proprietary advantage, information must remain inside a boundary the provider cannot enter, or you need processing the service cannot support. See the build-versus-buy worksheet for a like-for-like cost review.

A concrete example

CorpusMesh maintains knowledge; your application runs its model.

CorpusMesh focuses on approved external sources. Its private-beta retrieval contract returns ranked passages, citations, source versions, and an explicit retrieval decision. Your team retains the model, prompts, user experience, and final decisions.

Private beta The interfaces shown here define the intended product contract. Public access is not yet available.

For the published EU AI Act example, start with the source coverage and exclusions. Follow a passage through the reviewed source trace, inspect the RAG API request and response, then review source maintenance and the current beta scope and pricing.

The public trace is a reviewed example pinned to its source version. It is not a live query or a guarantee that your workflow will meet its acceptance criteria. Access is reviewed manually for an approved organization and named application.

Plan your evaluation

Bring the sources and questions your product depends on.

Describe your coverage, update requirements, and expected retrieval volume.

Discuss your knowledge workflow