Microsoft AI-901: Foundry SDK Basics

The Foundry SDK matters in AI-901 because Microsoft now expects more than portal-only familiarity. Candidates should understand the basic pattern of building a lightweight client that connects to a Foundry project or deployed model, sends an instruction or request, receives a response, and integrates that response into ordinary application logic.

The exam is not a professional software-development assessment, but it does expect conceptual familiarity with Python syntax, REST APIs, SDKs, CLIs, Azure resources, and the difference between configuring something in the portal and consuming it programmatically.

Think of the SDK as the application bridge to Foundry

The SDK gives application code a supported way to work with Foundry capabilities. The portal is useful for deploying models, testing prompts, configuring projects, and exploring agents. The SDK is how a lightweight application can use those resources repeatedly without a person clicking through the portal.

This distinction appears throughout cloud exams: management surfaces and runtime interfaces solve different problems. A developer may configure a model deployment once and invoke it thousands of times through code. AI-901 scenarios may therefore ask you to recognize when an SDK client is the correct interface.

Keep project configuration separate from application code

Applications should not hard-code environment-specific values throughout the codebase. Project endpoints, deployment names, resource identifiers, and similar settings are easier to manage when they come from configuration. This makes the same client easier to run across development, test, and production environments.

The same principle reduces accidental mistakes during exam reasoning. A deployment is not the same thing as the source model, and a project or endpoint is not the same thing as an application. Treat each as a separate layer with a specific role.

Use secure authentication instead of embedding secrets

Authentication is part of the client pattern. In Azure, application code should use an appropriate identity mechanism rather than storing credentials in source code. The exact credential class can vary with the environment, but the principle is stable: authenticate the caller securely and grant only the permissions the application needs.

Hard-coded keys create operational and security problems because they can leak through repositories, logs, or copied samples. AI-901 does not require deep identity engineering, but it does expect candidates to understand that an AI client still follows normal cloud-security practices.

Send clear system and user instructions

A lightweight generative client normally separates durable behavior from the user’s immediate request. System-level instructions define role, boundaries, or output expectations, while user content carries the current question or task. Keeping those responsibilities distinct makes prompts easier to reason about and test.

The programmatic interface does not make prompt design less important. The SDK simply delivers instructions to the model. Poorly structured prompts, ambiguous context, or untrusted input can still produce weak or unsafe results, so prompt design remains part of the application.

Handle the response as data, not magic

Model output arrives through a software interface and should be treated like any other external result. The application needs to read the response, check whether the request succeeded, extract the useful content, and decide how to present or process it.

If downstream code expects structured fields, do not assume every free-form answer will match the required format. A stable application may request a structured result, validate it, and reject or retry responses that do not meet the contract.

Agents add state and tool use to the client pattern

An agent client is conceptually similar to a model client, but the application may be interacting with an agent that has instructions, tools, knowledge, or conversation state. The agent can decide to use additional capabilities before producing a final answer.

This makes observability more important because a single user request may trigger several internal steps. For AI-901, the core point is recognizing that a lightweight client can invoke an agent rather than implementing all orchestration logic directly in the application.

Plan for transient failures and rate limits

Cloud calls can fail temporarily. Network errors, service throttling, unavailable dependencies, or malformed requests should not cause an application to behave unpredictably. Lightweight clients should use sensible retry and timeout behavior and should surface clear failures when recovery is not appropriate.

Retries must also be safe. If a request triggers an external action, blindly repeating it could create duplicate work. This is one reason tool-using agents need careful boundaries and idempotent designs where possible.

Log enough to troubleshoot without leaking content

Useful logs capture request identifiers, timing, model or deployment choices, error categories, and application state that helps explain what happened. Logging every full prompt and response is not always appropriate because those messages may contain sensitive data.

Good AI observability balances diagnostics with privacy. That principle connects the SDK basics topic with the security and responsible-AI material in AI-901: implementation decisions determine what data exists outside the model interaction itself.

Separate model quality from client bugs

When an application returns a bad result, first determine whether the request reached the correct deployment with the intended parameters and prompt. A surprising answer can come from a coding error, stale configuration, truncated context, or model behavior. Treating every problem as ‘the model is wrong’ slows troubleshooting.

Capturing the final request shape and response metadata in a safe form helps isolate issues. A reproducible test case is much more useful than a screenshot of a surprising answer with no context.

Use evaluation before promoting a new client or prompt

Changes to prompts, models, parameters, or tool definitions can alter behavior. A small set of representative test cases helps detect regressions before users do. Even a fundamentals-level application benefits from explicit expected outcomes rather than informal spot checks.

Evaluation also makes model upgrades easier. If the app has a known test set, the team can compare candidate deployments instead of choosing by reputation or model size alone.

Recognize where the portal ends and code begins

Use the portal for exploration, resource configuration, deployment, and interactive testing. Use the SDK when the requirement is to integrate that capability into software. Some exam questions become straightforward once you distinguish a human management workflow from an application runtime workflow.

The scope of Microsoft AI certifications puts AI-901 on the foundational path, yet its objectives include hands-on implementation concepts. Candidates should recognize simple SDK calls, prompt configuration and model-selection trade-offs rather than treat the credential as a purely conceptual survey.

Exam focus: follow the client lifecycle

A simple mental model is authenticate, connect, configure the request, send it, receive the response, validate the result, and handle errors. For agents, add state, tools, and potentially multi-step execution. For speech, vision, or information extraction, the same application lifecycle wraps a different AI capability.

Do not overcomplicate fundamentals questions with production architecture that the scenario does not ask for. The broader AI and generative AI certification path includes deeper engineering credentials; AI-901 mainly wants you to recognize the correct implementation building blocks.

Keep SDK samples small enough to understand

Fundamentals candidates benefit more from a client that clearly shows authentication, request construction, response handling, and error behavior than from a large framework with many abstractions. A minimal example exposes the boundaries between the application and the Foundry service.

Once the basic flow is understood, production concerns such as dependency injection, configuration providers, telemetry, and deployment packaging can be added. The exam is testing whether you recognize the pattern, not whether you can reproduce a large enterprise codebase from memory.

Use asynchronous execution only when the application needs it

Some AI calls can take noticeable time, especially when tools, retrieval, or media processing are involved. Applications with many concurrent requests may benefit from asynchronous patterns, but that is an engineering choice rather than a universal requirement.

The important concept is responsiveness. A client should not freeze an interactive experience unnecessarily, and long-running work should have clear status and timeout behavior. At the fundamentals level, understand why request duration affects the application even if you are not asked to implement advanced concurrency.

Test configuration changes independently from code changes

Model deployment names, endpoints, prompt text, and runtime parameters can change even when the application code does not. Keeping configuration external makes it possible to test those changes without rebuilding every part of the client.

This also helps incident response. If production behavior changes after a model deployment update, teams can compare configuration history separately from code history and identify the likely cause faster.

Design the client so models can be replaced

A tightly coupled client can make a model upgrade unnecessarily risky. If application logic depends on one exact response shape or one deployment name scattered throughout the code, replacement becomes difficult. A small adapter layer can isolate model-specific details from the rest of the application.

This is useful even in a lightweight AI-901 example because it teaches the right mental model: the AI service is a dependency of the application, not the entire application. Good boundaries make experimentation and migration safer.

Walk through a simple client request end to end

A useful mental exercise is to trace one request. The application loads its configuration, authenticates, creates a client for the intended Foundry resource, sends a prompt to the selected deployment, receives the result, validates that the response is usable, and then renders it or passes it to another function. At each stage there is a distinct class of failure: missing configuration, rejected authentication, unavailable deployment, malformed request, service error, or invalid output.

Thinking in this sequence makes troubleshooting systematic. It also mirrors how exam questions describe problems: the symptom may be a failed response, but the correct fix depends on which stage of the client lifecycle broke.