{"id":2926,"date":"2026-10-08T15:12:18","date_gmt":"2026-10-08T15:12:18","guid":{"rendered":"https:\/\/www.exam-topics.info\/blog\/rag-in-microsoft-foundry-design-the-retrieval-before-the-answer\/"},"modified":"2026-10-08T15:12:18","modified_gmt":"2026-10-08T15:12:18","slug":"rag-in-microsoft-foundry-design-the-retrieval-before-the-answer","status":"publish","type":"post","link":"https:\/\/www.exam-topics.info\/blog\/rag-in-microsoft-foundry-design-the-retrieval-before-the-answer\/","title":{"rendered":"RAG in Microsoft Foundry: Design the Retrieval Before the Answer"},"content":{"rendered":"<p>Consider an engineering assistant asked whether a software release can proceed under a company&#8217;s change policy. The policy library contains the current approval rules, a superseded version from last year and a technical document describing release gates for one legacy product. A language model can write a confident answer based on any of them. Retrieval-augmented generation (RAG) is supposed to supply the right evidence, but it only works reliably when the retrieval system can distinguish authoritative, relevant and permitted documents. The design challenge is not merely connecting Microsoft Foundry to a search index; it is determining which information deserves to enter the answer.<\/p>\n<p>Microsoft Foundry supports workflows grounded in enterprise data, commonly using Azure AI Search and agent tools. Some designs retrieve a small set of passages with a conventional search request. Others use a more agentic retrieval process that can reformulate a complex question into subqueries. Those patterns solve related but different problems. A simple retrieval system may be transparent and efficient for narrow questions; multi-query retrieval can help with compound questions, but may introduce latency, cost and additional opportunities to fetch irrelevant content.<\/p>\n<h3>Begin with a data contract and document authority<\/h3>\n<p>Before creating embeddings, decide which repository is authoritative, how documents are updated and which versions should be excluded. A change policy may reside in SharePoint, a document-management platform or a controlled internal repository. Files copied into an index become a separate data lifecycle: publication, replacement, revocation and deletion must be reflected there. Otherwise an agent can answer from a policy that was removed from its original location months earlier. Retrieval freshness is a governance requirement, not a tuning parameter at the end of the project.<\/p>\n<p>Each indexed record should carry meaningful metadata: source system, stable document ID, version or effective date, owner, document type, classification and access scope. Chunk IDs need to connect back to their originating document so investigators can trace an answer. If a source contains an effective-from date and a superseded-on date, preserve them. Do not ask a model to infer which of two almost identical policies is newer based solely on wording. Authoritative metadata should guide filtering and ranking.<\/p>\n<p>The same principle applies to access control. If two employees can ask identical questions but only one has permission to see a confidential operating procedure, retrieval must return different authorized evidence sets. Filtering the final answer after unauthorized chunks have already entered the model context is weaker than enforcing permissions at retrieval. Build the entitlement mapping before choosing a vector index configuration. For security-sensitive use cases, resource authorization is part of the RAG correctness definition.<\/p>\n<h3>Index structure determines the questions you can answer<\/h3>\n<p>Chunking is often introduced as a mechanical step of splitting a document into fixed lengths. That approach can separate a rule from its exception. A contract clause may say that a purchase requires approval, while the next paragraph lists exemptions. If those pieces are retrieved independently, the model can give an incomplete answer despite using apparently relevant text. Preserve headings, section identifiers, table context and related clauses where practical, while still keeping passages small enough to rank and assemble efficiently.<\/p>\n<p>Different source types call for different ingestion strategies. A troubleshooting runbook with commands may benefit from section-aware chunks. A price schedule may require normalized structured records rather than free-text chunks. A long incident postmortem can be split by chronology or subsystem, with metadata identifying the incident and revision. OCR and extracted tables deserve careful verification, because an incorrectly parsed column can reverse the meaning of a limit or threshold. If a source cannot be reliably parsed, the system should not silently treat its extracted text as authoritative.<\/p>\n<p>Embedding similarity helps locate conceptually related passages, but exact identifiers, product codes and regulation numbers often benefit from lexical matching. Hybrid retrieval combines text and vector signals and may improve recall when users mix precise terms with natural-language descriptions. The best configuration is workload-specific. A corpus with thousands of repeated SKU descriptions is different from a small set of deeply structured policies, and an index that performs well on one may behave poorly on the other.<\/p>\n<h3>Retrieve and rerank for the actual user task<\/h3>\n<p>A question such as \u201cCan we ship this change?\u201d is underspecified until the assistant knows which product, environment, risk category and approval status apply. A conventional retriever might use a single query and receive general release-policy chunks. A more deliberate pipeline identifies the missing parameters, either asks the user or retrieves additional context from authorized systems, and then searches for the relevant policy sections. Agentic retrieval may decompose the question into \u201cproduction deployment gates,\u201d \u201csecurity approval exceptions,\u201d and \u201crollback requirement\u201d subqueries. That can help if the answer truly spans several documents.<\/p>\n<p>More retrieval is not always better. Expanding a question into ten speculative subqueries may flood context with irrelevant passages and give the model more chances to be distracted. Limit query breadth according to the question, deduplicate passages and rerank candidates based on relevance and authority. A good search result set answers the question with enough evidence while minimizing contradictory or low-value material. It is not simply the largest number of chunks the token budget will accept.<\/p>\n<p>When two documents conflict, the agent should not blend them into a synthetic compromise. It needs version and policy-precedence rules or a clear statement that the authoritative source is uncertain. In a regulated environment, the workflow might require escalation rather than inference. A retrieved document describes what an organization has written; only a separate control system can determine whether a change has actually been approved.<\/p>\n<h3>Assemble context that preserves meaning<\/h3>\n<p>The prompt-construction step should make source boundaries visible. Include stable references, titles and dates alongside retrieved passages so the model can attribute claims. Avoid presenting document text in the same channel or format as the application&#8217;s controlling instructions. Documents may contain sentences like \u201cignore earlier rules\u201d as examples or malicious content; those are data, not operational commands. An agent should interpret the passage&#8217;s topic while refusing to grant it authority over tool use.<\/p>\n<p>Source attribution also supports an answer style appropriate for enterprise use. The assistant can state a policy requirement, cite the precise section and explain how the requirement relates to the facts of the user&#8217;s case. If the relevant source is absent, it should communicate the gap rather than borrow from general model knowledge without warning. For decisions with business consequences, distinguish sourced facts, assumptions and recommendations. A citation to a retrieved passage is meaningful only when that passage actually supports the sentence that carries it.<\/p>\n<p>Keep answer evidence separate from mutable business state. RAG can find a runbook that says a release requires a rollback plan; it cannot prove that the current deployment has an approved rollback plan unless the workflow checks the relevant change record. Combining retrieved policy with a structured change-management API can produce a genuinely actionable answer. That architecture resembles broader <a href=\"https:\/\/www.exam-topics.info\/ab-100\">agentic solution design<\/a>: grounding helps interpretation, while deterministic systems establish current state and enforce conditions.<\/p>\n<h3>Evaluate retrieval failures, not only polished responses<\/h3>\n<p>Test questions should represent the corpus&#8217;s most demanding cases: an exact SKU, two documents that differ only by date, a policy with an exception in another section, a revoked document, and a sensitive file that the test user is not allowed to read. Measure whether the right evidence appeared among retrieved results and whether unauthorized or outdated content was excluded. An answer can accidentally be correct despite poor retrieval, so the pipeline needs independent retrieval tests as well as answer review.<\/p>\n<p>Useful metrics include retrieval recall on labeled source passages, precision of relevant results, version accuracy, permissions violations, citation support and time to first useful evidence. Compare results by query category. An index optimized for natural-language policy questions may underperform on log codes or command syntax. Use examples from genuine support and operations requests rather than a small set of artificial questions that happen to match the index&#8217;s wording.<\/p>\n<p>Groundedness evaluation can help identify when the model adds unsupported claims, but it is not a substitute for validating whether the cited evidence is current. A grounded answer to an obsolete policy remains operationally wrong. Document freshness and access filtering need deterministic checks where possible, and human owners should review the highest-impact failure cases.<\/p>\n<h3>Control latency, cost and operational drift<\/h3>\n<p>RAG cost depends on indexing, embedding generation, search operations, reranking, model context size and repeated queries. A complex agentic retriever may outperform simple retrieval on demanding questions but be unnecessary for short definitions. Route retrieval approaches by task complexity and authorization needs. Cache only where access rules and content freshness permit; a cached summary produced for a privileged user must not be served to a less privileged one. Every performance optimization should preserve entitlement and version guarantees.<\/p>\n<p>Monitoring should reveal when ingestion stalls, indexing fails, document access rules change or a retrieval query begins returning a different mix of sources after a model or ranking update. Record which document versions contributed to an answer without collecting more personal information than necessary. A production incident may require reconstructing why the agent relied on a superseded procedure, so traceability deserves consideration at design time.<\/p>\n<p>Teams preparing for <a href=\"https:\/\/www.exam-topics.info\/ai-300\">AI platform operations<\/a> or for the related <a href=\"https:\/\/www.exam-topics.info\/aws-certified-generative-ai-developer-professional-aip-c01\">AWS generative AI development<\/a> track will encounter analogous tradeoffs: data freshness, retrieval quality, guardrails, model context and operating cost are not uniquely Microsoft problems. In Foundry, the implementation choices are platform-specific, but the quality criteria are broadly applicable.<\/p>\n<h3>A release-policy example from question to decision<\/h3>\n<p>Suppose a user asks whether a database migration can deploy tonight. The system first identifies the application and environment, then retrieves the currently effective release policy and the relevant migration runbook. It checks the user&#8217;s access before revealing restricted material. The index returns passages about change windows, rollback plans and additional review for schema changes; a structured change-management call separately verifies whether the current deployment has required approvals. The answer explains which conditions are satisfied, which remain open and which source establishes each rule.<\/p>\n<p>If the most recent policy is missing from the index, the assistant should not silently rely on an older version. It can disclose that the relevant policy cannot be verified and direct the user to an authorized process. That behavior may feel less impressive than an instant yes-or-no answer, but it is evidence of a dependable architecture. RAG succeeds when the right information reaches the model under the right authority and the resulting answer respects what the sources can actually establish.<\/p>\n","protected":false},"excerpt":{"rendered":"<p>Consider an engineering assistant asked whether a software release can proceed under a company&#8217;s change policy. The policy library contains the current approval rules, a [&hellip;]<\/p>\n","protected":false},"author":1,"featured_media":0,"comment_status":"","ping_status":"","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[1],"tags":[],"class_list":["post-2926","post","type-post","status-publish","format-standard","hentry","category-uncategorized"],"_links":{"self":[{"href":"https:\/\/www.exam-topics.info\/blog\/wp-json\/wp\/v2\/posts\/2926","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.exam-topics.info\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.exam-topics.info\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.exam-topics.info\/blog\/wp-json\/wp\/v2\/users\/1"}],"replies":[{"embeddable":true,"href":"https:\/\/www.exam-topics.info\/blog\/wp-json\/wp\/v2\/comments?post=2926"}],"version-history":[{"count":0,"href":"https:\/\/www.exam-topics.info\/blog\/wp-json\/wp\/v2\/posts\/2926\/revisions"}],"wp:attachment":[{"href":"https:\/\/www.exam-topics.info\/blog\/wp-json\/wp\/v2\/media?parent=2926"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.exam-topics.info\/blog\/wp-json\/wp\/v2\/categories?post=2926"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.exam-topics.info\/blog\/wp-json\/wp\/v2\/tags?post=2926"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}