A document classifier usually works during a demonstration. A user uploads a file, a function runs, and the result appears. In production, files arrive in bursts, a downstream service becomes unavailable, and a retry starts while the previous attempt may already have succeeded. Event-driven AI architecture is about preserving the meaning of work under those conditions, not merely connecting services with arrows. This is part of Microsoft AI-200’s backend-development scope alongside the broader Microsoft certifications.
Imagine a claims-processing service that accepts photographs, PDFs and typed incident reports. After upload it must extract text, classify the claim, check supporting records and notify a reviewer. A single synchronous request can do some of that work when traffic is low. It becomes fragile when optical processing or inference latency varies, when a vendor API imposes limits, or when a reviewer requires the original file to be retained independently of the classification result. Separating events from durable work is the foundation of a recoverable design.
An event says something happened; a queue assigns work
Azure Event Grid supports event distribution: a producer announces an event and interested handlers receive it under the service’s delivery model. A new blob or changed record may produce an event indicating that a business object is ready for downstream attention. Event consumers should not assume an event is a complete, authoritative business document. Its payload commonly provides context or a reference, while the handler obtains the current source record through an appropriately authorized access path.
Azure Service Bus provides queues and publish-subscribe messaging features designed around durable messaging scenarios. A work message may represent a requested action that a worker should process and settle. Topics and subscriptions can allow different consumers to receive messages according to routing needs. Knowing that both Event Grid and Service Bus can move notifications does not make their delivery semantics, filtering capabilities or operational ownership identical.
In the claims example, a storage upload event can initiate processing, but a durable job record should identify which file version needs analysis and what completion means. If OCR takes several minutes, a queue-backed worker can process the task independently of the upload API’s HTTP timeout. A workflow might use Service Bus to control work intake and Event Grid for changes that multiple interested systems should observe. Choosing one or both should follow real requirements rather than a rule that every architecture needs every Azure messaging service.
Events should describe committed business changes. Publishing ‘claim approved’ before the database transaction actually commits risks downstream systems acting on an approval that never happened. Conversely, if the database commit succeeds but the publish attempt fails, the organization needs a recovery path. Outbox-style patterns, idempotent publishing and periodic reconciliation can help address this dual-write challenge. The important insight is that reliability cannot be achieved by placing an optimistic send() call immediately after a database update and hoping both operations always succeed.
Contracts for event data should be versioned. A new extractor might add a classification field, but old consumers may still expect the previous shape. Schemas, required identifiers and backward-compatibility rules help services evolve independently. Treat a producer’s event as an integration contract, not as a dump of whichever object its developer had available at the moment.
Azure Functions is execution, not the whole workflow
Azure Functions can react to supported triggers and coordinate discrete units of code. Its serverless execution model can suit bursts of work without requiring every team to manage dedicated virtual machines, depending on selected hosting and runtime configurations. It does not eliminate business design questions. A function triggered twice must not charge the customer twice; a function that times out must not leave a record in an unexplainable intermediate state.
A well-defined function handler reads the message, validates schema and authorization context, performs one bounded piece of work, persists a durable result and then completes processing according to the trigger’s settlement behavior. If the work is too large, divide it into stages with explicit state transitions. A design in which the function performs a ten-minute chain of untracked external calls may be difficult to recover and even harder to explain during an incident.
Azure Container Apps offers another path for event-driven workloads, including supported scaling based on workload signals through KEDA. A specialized AI worker might need a custom container, native processing libraries or runtime resources that make container hosting attractive. That does not make it automatically better than Functions: deployment footprint, cold-start requirements, concurrency management and team expertise matter. Choose the execution environment after defining the task’s duration, throughput and operational constraints.
Concurrency deserves attention before scaling. Suppose a worker issues five expensive model calls per job. Scaling replicas aggressively against queue depth can overwhelm a rate-limited inference endpoint or exhaust a database connection pool. The system needs an end-to-end capacity model, perhaps with bounded concurrency, throttling and backpressure. More worker replicas can reduce a queue only when the downstream services can accept the extra work.
The integration skills in Microsoft AZ-204 overlap with Functions and messaging fundamentals, but the AI-200 perspective includes data retrieval, vector workloads, model-related variability and backend observability. The exam-worthy judgment is not just identifying a trigger. It is explaining why a chosen trigger, queue, function or container handles the workload’s failure and scaling characteristics.
Retries only work when processing is safe to repeat
Distributed systems experience ambiguous outcomes. A client sends a message and loses its connection before receiving acknowledgment; it cannot be certain whether the broker accepted the send. A worker writes a classification result and crashes before completing the message; the broker may deliver the work again. Reliable consumers assume that duplicate processing is possible and design a business-safe response.
Give each job a stable identifier and distinguish attempts from jobs. If a claim needs classification version three, write the output against that claim and processing version. A retry should either confirm the existing result or safely replace an incomplete attempt according to defined rules. A randomly generated result ID on every retry is not enough; it can create an expanding pile of duplicate classifications that operators cannot reconcile.
Service Bus duplicate detection can help with selected send-side repeat-message cases when configured and supported, but it is not a replacement for idempotent consumers. Duplicate messages can still arise outside the deduplication window or from conditions unrelated to duplicate sends. A design should explain what happens when the same legitimate command is delivered twice and whether the second delivery causes another side effect.
Dead-letter queues provide a way to isolate work that could not be processed successfully under configured delivery conditions. They are not an archive for messages nobody wants to read. A support process should identify why messages arrived there, classify permanent versus transient failures, and decide whether replaying a message is safe. An invalid document format may require correcting the input; a temporary dependency outage may justify a delayed retry. Treating both with infinite immediate retries can create an outage from a single bad record.
Time limits and lock renewal are operational details with business consequences. A worker holding a message lock can lose it if processing takes too long or connectivity fails. The same job may then be executed by another worker while the first is still waiting for an external AI response. Idempotent storage updates, bounded timeouts and explicit state transitions are stronger protections than assuming a lock makes every workflow effectively exactly-once.
Keep durable state outside the worker process
An event-driven worker may be restarted, scaled down or replaced during deployment. Important workflow state should not live only in a local variable or temporary container file. Store the authoritative job status and output reference in an appropriate durable system. The worker can then resume based on persisted progress rather than repeating every stage without knowledge of what already happened.
Claims processing often benefits from a state model more expressive than done or failed. Useful states may distinguish received, validating, extracting, awaiting external inference, ready for review, and finalized. Each state should have allowed transitions, timestamps and ownership. When a customer asks why their claim is delayed, the answer should be based on persisted evidence, not on whether a transient worker happens to be running.
Data dependencies also change. If a document is replaced after an upload event but before a worker reads it, which version should be analyzed? A stable object version or content identifier prevents the worker from silently processing a newer file under the old event. If a record is deleted for privacy reasons, retries must not resurrect prohibited data. Idempotency is therefore tied to data lifecycle and governance as well as reliable message delivery.
When a workflow updates external systems, record enough correlation information to determine whether the external effect occurred. A billing API may return a request ID even if the connection closes during the response. A recovery process should reconcile against that ID before reissuing a payment-related operation. The same general problem appears in AI pipelines that send notifications, create tickets or commit generated artifacts. Network retries must not change the business meaning of the original request.
Observe queue health and user outcomes together
Queue depth alone is not a measure of application health. A queue may stay short because workers are discarding messages incorrectly, or grow during a predictable peak while customer deadlines remain satisfied. Track arrival rate, successful processing, retries, dead-letter counts, oldest-message age and end-to-end completion latency. Examine those measurements alongside the capacity and error signals of downstream inference, database and storage services.
Correlation IDs should survive every handoff. An upload request creates a job; an event announces the upload; a queue message assigns extraction; another worker indexes the result. Logs and traces should link these steps without storing confidential document content indiscriminately. Azure Monitor and OpenTelemetry-oriented instrumentation can help reconstruct that path, but the application must consistently propagate identifiers and classify expected versus unexpected failures.
An alert should correspond to an operational decision. For example, a growing oldest-message age may require more worker capacity, while a spike in throttling errors may require reducing concurrency or correcting a retry strategy. A dead-letter trend might indicate an incompatible event schema after deployment. Alerting on every temporary error creates noise, while ignoring the age of stalled work leaves users waiting indefinitely with no explanation.
Operational dashboards should include the business outcome. What percentage of claims reach a reviewer within the expected time? How many documents were processed from the wrong version? Were any duplicate notifications sent? A pipeline can show healthy CPU and queue metrics while violating its real service objective. Defining those measures makes the architecture testable and gives incident responders something more useful than a generic green light.
Practice with the failure that the diagram hides
A strong AI-200 scenario can be approached by drawing the normal path and then deliberately breaking one component. What happens if a document-upload event is delivered twice? What happens if the model service returns a timeout after performing work? What happens if a worker dies after writing results but before settling its message? The answers should identify persistent state, retry ownership, idempotency and diagnostic evidence.
Then change the traffic pattern. Instead of ten uploads an hour, imagine fifty thousand arriving after a partner migration. Which service buffers the work, which component scales, what downstream quotas become limiting and which jobs should be prioritized? A design that survives normal traffic but floods its dependencies under load is not production-ready. The point of event-driven AI architecture is to turn unpredictable work into bounded, observable and recoverable operations.