Microsoft AI-103: Building Agents in Foundry

An agent is more than a chatbot with a longer prompt. It combines a model with instructions, state, tools, knowledge and a runtime that allows the model to decide what action should happen next. Microsoft Foundry Agent Service brings those pieces together, but AI-103 expects developers to understand how to design the agent itself rather than treating the managed platform as a black box.

Within the Microsoft AI certification family, this is a major shift toward application architecture. The candidate must be able to define an agent’s role, give it appropriate tools, connect knowledge, constrain permissions and evaluate behavior across multiple turns.

Start with a role and a bounded responsibility

A useful agent begins with a clear job. “Help with everything” is not a production requirement. A support agent may be responsible for diagnosing product issues, searching documentation and preparing a response. An operations agent may investigate an alert, collect evidence and recommend a remediation without being allowed to execute the change.

The narrower role makes instructions, tools and evaluation much easier to design. It also limits blast radius. If an agent’s purpose is read-only analysis, it does not need write access simply because a downstream API supports it.

Good instructions define objectives, decision rules, escalation conditions and what the agent should do when it lacks enough information. They should not attempt to encode every possible path. Deterministic rules belong in application logic where possible; the agent should be used for the decisions that genuinely require interpretation.

Tool schemas are part of the agent’s reasoning environment

Tools convert an agent from a language interface into an actor. They can search, query databases, call APIs, run code or trigger business workflows. Tool design therefore affects both capability and safety.

A strong tool has a clear name, a narrow purpose and a schema that makes invalid actions difficult. If two tools overlap heavily, the model has a harder selection problem. If one tool exposes dozens of unrelated operations, permissions and error handling become more difficult to reason about.

Descriptions matter because the model uses them when deciding whether and how to call a tool. The application should also validate arguments deterministically. A model-generated tool call is a proposal; normal software validation still applies before an external system is changed.

Knowledge and tools solve different problems

Knowledge sources help the agent answer questions from private or changing information. Tools let the agent perform an action or fetch structured runtime data. These capabilities can overlap, but they are not interchangeable.

A policy handbook may belong in retrieval so the agent can cite the applicable rule. A current account balance should normally come from an authorized system call rather than a document index. Product documentation can be retrieved, while a ticket update should go through a tool with explicit permissions.

The architecture becomes clearer when every connection is classified as evidence or action. Evidence informs reasoning. Actions change or query the environment. The security controls around actions usually need to be stricter.

Conversation state needs a deliberate design

Agents often operate across multiple turns, which means state accumulates. The system may need conversation history, task progress, tool results, user preferences or external state. Simply replaying an unlimited conversation into every request is expensive and can make important details harder to identify.

State can be summarized, stored externally or separated into durable and temporary components. A long-running task may keep a concise plan and retrieve details as needed rather than carrying every previous token. The right approach depends on whether exact past wording matters and how long the task is expected to run.

State is also a privacy boundary. Applications should retain only what is needed, apply the correct access controls and avoid mixing one user’s context into another user’s session.

Identity should be designed before the agent is given power

Foundry agents can operate with Microsoft Entra identity and Azure RBAC. This allows a system to access Azure resources without embedding reusable secrets in prompts or application configuration. The principle is straightforward: the agent should receive only the permissions required for its role.

Read and write capabilities should be separated where possible. An investigation agent may need to read logs but not delete resources. A deployment agent may be able to modify a narrow resource group but not the entire subscription. If a tool can act on behalf of a user, the application must understand the difference between the agent’s service identity and delegated user authority.

This is the same least-privilege discipline used elsewhere in Azure, and it is one reason AI engineering increasingly overlaps with cloud security and platform engineering.

Approval flows are part of agent architecture

Not every action should be autonomous. A well-designed agent can gather evidence, prepare a plan and then pause before an irreversible step. Approval boundaries work best when they are tied to consequence rather than inserted after every tool call.

The approval request should include context: what action is planned, which resource or record will change, why the agent wants to perform it and what the expected result is. This makes the human checkpoint meaningful.

Approval is especially important when an agent can communicate externally, move money, change access rights or modify production systems. The broader AI and generative AI certifications landscape increasingly treats these governance decisions as core engineering skills rather than separate compliance concerns.

Agent reliability depends on explicit failure behavior

Tools fail. APIs return rate limits. Search results can be empty. A model may choose a poor tool or produce malformed arguments. The agent needs a recovery strategy instead of assuming every step works on the first attempt.

Some failures justify a retry with backoff. Others require a different tool, a revised query or human escalation. The architecture should distinguish transient failures from semantic failures. Repeating the same wrong action three times is not resilience.

Stopping conditions matter too. An agent that can keep calling tools indefinitely can consume cost and create cascading side effects. Limits on turns, tool calls or elapsed time provide a deterministic safety net around probabilistic behavior.

Testing must cover the trajectory, not just the response

An agent can reach the correct final answer through an unacceptable path. It may call a restricted tool, expose sensitive context or require far more steps than necessary. Evaluations should therefore capture the sequence of actions as well as the final state.

Useful test cases include successful paths, missing information, conflicting evidence, tool errors, denied permissions and malicious content returned by tools. Tracing can reveal which model calls and tool decisions caused a failure, while evaluation metrics can track task completion, safety and efficiency over time.

Production monitoring then extends those tests to real traffic. Latency, token use, tool error rate, approval frequency and evaluation outcomes help teams understand whether an agent is behaving as designed.

Single-agent simplicity should be the default

It is easy to assume that a sophisticated workload needs several specialized agents. Often one well-designed agent with a focused tool set is easier to operate. Multi-agent systems become useful when tasks can genuinely be decomposed, parallelism creates value or different workers need distinct permissions and instructions.

That distinction is important for AI-103. Foundry provides rich agent capabilities, but the exam skill is not maximizing the number of components. It is matching the architecture to the requirement.

Developers who are building toward this role can also use the site’s Azure AI engineering career path as broader context. The practical competency is consistent: build agents that are capable enough to solve the task, constrained enough to be trusted and observable enough to improve when reality exposes a weakness.

Agent instructions should separate policy from strategy

Instructions often become difficult to maintain when hard policy, workflow hints and writing preferences are mixed into one long prompt. A better design separates rules the agent must never violate from strategy it can adapt as the task changes. Deterministic application checks should enforce non-negotiable constraints whenever possible.

This separation improves testing. Policy failures can be evaluated independently from task-planning quality. Developers can refine how the agent searches or summarizes without accidentally weakening permission or approval behavior.

It also makes versioning meaningful. When the agent changes, the team can identify whether the change affected instructions, tools, knowledge or runtime configuration. That traceability matters when a production regression appears after a deployment.

Production agents need lifecycle management

An agent evolves after release. Tool APIs change, data sources move, prompts improve and models are updated. Versioning makes it possible to compare behavior and roll back when a change creates a regression.

Publishing should therefore be treated like application deployment. Stable endpoints, controlled versions, test environments and release gates reduce the risk of editing a production agent directly. The operational owner should know which agent version handled a request and which configuration was active at the time.

That lifecycle discipline is part of what turns Foundry Agent Service from a prototyping environment into a production platform.

Tool discovery should not overwhelm the agent

Large enterprise agents can accumulate dozens of integrations. Exposing every tool to every request increases context and creates more ambiguous choices. Toolboxes, routing or staged discovery can keep the active capability set focused.

The design goal is not maximum tool count. It is giving the agent the right capabilities at the point they are needed, with descriptions and permissions precise enough that tool selection remains reliable.