Microsoft AI-103: Multimodal AI Workflows

Multimodal AI works across more than one kind of input or output: text, images, audio, video and documents. The engineering challenge is not simply selecting a model that accepts several modalities. A production workflow still needs to decide how media is ingested, normalized, analyzed, grounded, secured and converted into an output that downstream systems can trust.

AI-103 places multimodal work alongside generative applications, agents and information extraction. That makes it an important part of the Microsoft AI certifications path rather than a separate specialist topic.

Choose the workflow from the question the media must answer

An image can be processed for many different purposes. The application may need a caption, object identification, visual question answering, document extraction, accessibility text or evidence for a broader decision. Those goals imply different prompts, models and evaluation criteria.

Video introduces time. The system may need to analyze selected frames, identify events, summarize a sequence or combine visual evidence with audio. Long media also raises context and cost concerns, so preprocessing and segmentation become architectural choices.

Audio can be transcribed, translated, classified or treated as an input to a conversational agent. The workflow should identify whether the model needs the raw modality or whether a specialized service should first convert it into a more manageable representation.

Multimodal models reduce handoffs but do not remove preprocessing

A multimodal model can interpret an image directly and reason about it with text. This simplifies some applications because the developer does not need separate computer-vision and language pipelines for every case.

However, preprocessing still matters when data quality is inconsistent. Images may need resizing or orientation correction. Audio may contain long silence or noise. Documents may have complex layouts that benefit from OCR and structure extraction before reasoning.

The correct architecture is therefore not “always use the general multimodal model.” Specialized preprocessing is valuable when it increases reliability, reduces cost or produces a structured intermediate result that other services need.

Visual question answering requires evidence discipline

When users ask questions about images, the model should answer from visible evidence rather than from generic knowledge. This matters in inspection, support and compliance scenarios where an incorrect visual inference can trigger a bad action.

Prompts can direct the model to distinguish what is clearly visible from what is uncertain. Applications can also request structured observations before a conclusion, giving downstream logic a chance to validate the evidence.

For multiple images, ordering and labeling matter. The model needs to know which image belongs to which part of the task. A careless media payload can create confusion even when the model itself is capable.

Generation and editing have different control requirements

AI-103 includes image and video generation as well as editing workflows. Generating from a prompt is different from transforming supplied media. Editing may involve masks, reference material, inpainting or localized changes that should preserve unaffected regions.

The application should validate output requirements such as dimensions, format and acceptable content before sending media downstream. If the asset represents a brand, product or regulated communication, additional review may be required.

Human approval is often appropriate when generated media is public-facing. A technically valid image can still be misleading, off-brand or inappropriate for the business context.

Accessibility is a practical multimodal use case

Alt text and extended descriptions can make visual content more accessible, but a generic caption is not automatically useful. The right description depends on the purpose of the image. A chart needs relationships and trends. A UI screenshot may need controls and state. A decorative image may require little or no description.

The model should therefore be given task context. “Describe this image” is less useful than asking for an accessibility description for a specific audience and page purpose.

Evaluation should include whether the description captures the important visual information without inventing details. Accessibility quality is a content requirement, not merely a checkbox attached to model output.

Images and documents can carry prompt-injection risk

Untrusted media can contain text that attempts to influence an agent. A screenshot, scanned document or image may include instructions that are data from the user’s environment, not commands that should override application policy.

Multimodal systems therefore need the same trust boundaries as text-based RAG. Retrieved or extracted content should be treated as evidence. Sensitive tools should be constrained by permissions and deterministic validation so malicious media cannot directly convert into a high-impact action.

This becomes more important when an agent automatically processes email attachments, webpages or uploaded documents and then calls external systems based on what it finds.

Cost and latency rise quickly with rich media

Images, audio and video are heavier than short text prompts. Long media can increase inference time and processing cost. The workflow should avoid sending more content than the task needs.

Sampling, segmentation and preprocessing can help. A system may analyze selected video intervals rather than every frame, or extract document structure once and reuse it across several questions. Caching can improve repeat workloads when the freshness requirements allow it.

Model choice also matters. A high-capability multimodal model may be necessary for complex reasoning but excessive for straightforward classification. Architecture should match capability to the stage.

Evaluation must cover each modality and the transitions between them

A multimodal system can fail even if each component works in isolation. OCR may miss a number that the language model then confidently interprets. A transcription error can change the meaning of a request. A video segmenter may omit the event that matters.

Evaluation should therefore test the end-to-end outcome and important intermediate stages. Labeled examples can verify extraction accuracy. Human review can assess nuanced descriptions. Safety tests can probe whether embedded instructions influence the system improperly.

Operational monitoring should track failures by media type. A rising error rate for one document format may indicate an ingestion problem rather than a model regression.

Multimodal AI is a pipeline design problem

The strongest AI-103 answer usually starts with the requirement and works backward. Identify the media, determine the information or transformation needed, choose specialized preprocessing where useful, select the model, define output structure, enforce security and build evaluation around the actual consequence of an error.

That mindset also connects multimodal development with the broader AI and generative AI certification ecosystem. The value is not in demonstrating that a model can see, hear or generate media. It is in building a workflow where those capabilities are reliable, efficient and safe enough to support a real application.

Document, image, audio and video pipelines need different preprocessing

Multimodal does not mean every media type should be handled identically. Documents often need OCR and layout preservation. Images may need orientation or resolution checks. Audio benefits from segmentation and speaker-aware processing. Video may require frame selection, timestamps and correlation with a transcript.

A shared model can reason across these outputs, but the preparation stage should preserve the features that matter for the downstream task. Flattening every modality into plain text can discard useful evidence.

The application should also carry metadata such as source, time, page or frame. This makes explanations and audits more useful because the final result can point back to the exact piece of media that supported it.

Multimodal agents need a clear evidence hierarchy

When an agent combines a screenshot, a document and an API result, those sources may disagree. The system should define which source is authoritative for which question. A live inventory API may be authoritative for stock count while a product manual is authoritative for operating limits.

Without an evidence hierarchy, the model may resolve conflicts based on wording rather than business truth. Prompts can identify source roles, while deterministic application logic can enforce hard precedence where required.

This becomes especially important in support and operations, where stale screenshots or cached documents can conflict with current system state.

Output review should match the consequence of the media

Generated images used internally for brainstorming need a different review process from media published to customers. Visual analysis that merely helps a human inspect a system needs less automation risk control than analysis that automatically approves or rejects a claim.

Risk-based review keeps the workflow usable. Human approval is concentrated where an error has meaningful consequence, while low-impact transformations remain automated.

This same principle appears throughout AI-103: capability, autonomy and governance must be designed together.

Storage and retention decisions matter for rich media

Raw media can be much larger and more sensitive than derived text. Teams should decide whether original images, audio and video must be retained after analysis or whether a structured representation is sufficient for the business need.

Retention affects cost, privacy and incident response. If original media is kept, access controls and lifecycle policies should match its classification. If it is discarded, the system should retain enough provenance to explain how the resulting output was produced.

These operational choices belong in the workflow design, not in a cleanup phase after deployment.

Human review needs the original evidence

When a multimodal workflow routes a case to a person, the reviewer should see the source media and the model’s interpretation together. Reviewing only the generated text makes it difficult to notice a missed visual detail, transcription error or ambiguous region.

Well-designed review interfaces therefore preserve provenance such as page number, timestamp or image region. This reduces review time and creates better feedback for future evaluation cases because the team can identify exactly where the pipeline failed.