Virtual-machine availability is not a single feature in Azure. Administrators choose among availability zones, Virtual Machine Scale Sets, availability sets and recovery mechanisms depending on the failure they need to survive and whether the workload needs elastic capacity. The AZ-104 exam expects candidates to distinguish these options rather than assuming that every highly available design uses the same pattern.
The first question should be about failure domain. Are you protecting against a host failure, a rack or datacenter failure, a sudden demand spike, an entire regional outage, or a bad application deployment? Azure provides different controls for each problem. Availability becomes much easier to reason about when those failure modes are separated.
Availability zones protect against datacenter-level failure
Azure availability zones are physically separate locations within supported regions, with independent power, cooling and networking. Deploying redundant workload instances across zones reduces the chance that a datacenter-level failure removes every instance at once.
Zones do not create application resilience automatically. The workload still needs multiple instances, appropriate load distribution and data architecture that can tolerate one zone being unavailable. A single VM placed in one zone is still a single instance. The architecture gains resilience only when redundancy is actually used.
The broader Azure infrastructure certification path builds on this concept because network, storage and application dependencies must also be designed around the availability target. An administrator cannot claim zone resilience if a critical dependency remains tied to one failure point.
Scale sets combine fleet management with elasticity
Virtual Machine Scale Sets let administrators manage a group of VMs as a coordinated resource. Instances can be created from a common model, updated centrally and scaled in or out according to demand or schedule. This is different from manually creating several unrelated VMs.
Scale sets are useful when the workload benefits from interchangeable instances. Web tiers, worker fleets and stateless application components fit naturally because capacity can be added or removed without treating each VM as a special snowflake. Stateful workloads require more planning because instance identity and data placement may matter.
Scale sets can also use availability zones, allowing elasticity and zonal resilience to work together. The design decision is therefore not “zones or scale sets.” A scale set can be the mechanism used to distribute and manage instances across zones.
Flexible orchestration supports modern VM management
Azure supports different scale-set orchestration approaches. Flexible orchestration is designed for scenarios where administrators want scale-set management while retaining many characteristics of individual VMs. It can support heterogeneous configuration patterns more naturally than the older uniform model in some designs.
For AZ-104, focus on the outcome rather than memorizing every feature boundary. Determine whether the workload needs coordinated scaling, individual-instance control, zone distribution or a common image and configuration. Then select the scale-set approach that supports that operating model.
Microsoft also recommends modern scale-set options for many high-availability VM scenarios because they combine centralized management with features that older availability sets do not provide.
Availability sets address a narrower failure model
Availability sets group VMs so Azure can distribute them across fault and update domains within a datacenter-oriented model. They remain relevant for existing workloads and scenarios where zones are not being used, but they do not provide the same datacenter-level isolation as availability zones.
The key comparison is scope of failure. Availability sets reduce correlated host or maintenance impact. Zones separate instances across independent datacenter locations within a region. Site Recovery addresses disaster-recovery scenarios by replicating workloads for recovery elsewhere. These mechanisms should not be collapsed into a single “HA” category.
An existing architecture may also constrain the choice. Migrating a mature VM estate to a different availability model can require redeployment or other changes, so administrators need to evaluate operational impact rather than simply selecting the newest feature.
Autoscale should follow a workload signal
Scale sets can grow or shrink automatically based on metrics or schedules. A useful autoscale rule begins with a signal that actually represents pressure on the workload. CPU can be appropriate for compute-bound services, but queue depth, request rate or another metric may be a better indicator for some applications.
Scale-out and scale-in thresholds should not be so close that the system constantly adds and removes instances. Cooldown periods and sensible boundaries reduce oscillation. Minimum capacity protects the baseline service level, while maximum capacity prevents runaway scaling from creating unexpected cost.
Scaling also depends on startup time. If a new instance takes many minutes to become healthy, a threshold that reacts only after the service is saturated may be too late. Availability and capacity planning therefore remain application-specific even when Azure automates the mechanics.
Images and extensions affect fleet consistency
A scale set is easier to operate when every instance starts from a controlled image and receives predictable configuration. Image versions, VM extensions and initialization scripts should be treated as part of the deployment lifecycle rather than one-time setup conveniences.
Uncontrolled manual changes create configuration drift. When the scale set replaces an instance, that manual change may disappear. If a configuration is necessary for the workload, it should be reproducible through the image, extension, configuration-management process or application deployment.
This is where compute administration intersects infrastructure as code. The Azure Administrator certification coverage is strongest when compute, networking and deployment are understood as one operating system rather than separate exam chapters.
Load balancing and health determine whether redundancy is useful
Multiple healthy VMs do not help if traffic continues to be sent to an unhealthy instance. Load-balancing components and health probes are therefore part of the availability design. The probe needs to test something meaningful enough to distinguish a working application instance from a VM that is merely powered on.
Application Health extensions and load-balancer probes can support automated decisions about instance health, but administrators should understand what each probe observes. A TCP port being open may not prove that the application can complete a user request. A deeper endpoint can provide better evidence but can also create dependencies of its own.
Resilience is an end-to-end property. Compute redundancy, traffic management, data availability and dependency health must all align with the service objective.
Availability does not replace disaster recovery
A zone-resilient application can survive a zone outage inside a region, but it may still be affected by a region-wide incident or a destructive change replicated across all instances. Backup and disaster-recovery mechanisms address different risks.
This distinction is similar to the general difference between high availability, fault tolerance and disaster recovery. The site’s discussion of high availability and fault tolerance uses another cloud platform, but the architectural principle carries across providers: redundancy inside the serving environment and recovery after a larger failure are related but separate design goals.
AZ-104 candidates should therefore resist answers that promise too much from one feature. Scale sets solve fleet and scaling problems. Zones improve isolation. Backup protects recovery points. Site Recovery supports failover and recovery workflows.
Troubleshoot availability by checking the whole path
If a scale set fails to add capacity, inspect autoscale conditions, quota, SKU availability, image access, extension failures and networking. If traffic still reaches an unhealthy instance, inspect load-balancer health and the application probe. If a zone deployment fails, confirm the selected VM size and related resources are available in the chosen zones.
Do not diagnose every symptom as a compute problem. DNS, routes, NSGs, storage dependencies or identity can make a healthy VM appear unavailable to users. The administrator’s job is to identify the layer that actually failed.
This cross-domain troubleshooting is why AZ-104 remains an administrator exam rather than a VM exam. Compute knowledge becomes valuable when it is connected to the rest of the Azure environment.
Choose the availability mechanism from the required outcome
For a scalable stateless tier, think about scale sets and autoscale. For datacenter-level resilience, think about multiple availability zones. For an older multi-VM design without zones, availability sets may still be relevant. For recovery from broader failure, consider backup and Site Recovery rather than expecting high availability to solve disaster recovery.
The right answer is rarely the feature with the most impressive availability description. It is the mechanism whose failure boundary, operational model and recovery behavior match the workload requirement.
Update strategy can create or reduce availability risk
Availability is affected by maintenance and application updates as well as hardware failure. Replacing or updating every VM instance at once can create an outage even when the fleet is spread across zones. Administrators therefore need an update process that preserves enough healthy capacity while new images, extensions or application versions are introduced.
Scale-set designs can support rolling approaches in which only part of the fleet changes at a time. Health checks become important because the platform needs evidence that updated instances are ready before additional capacity is replaced. The application must also tolerate mixed versions for the duration of the rollout if old and new instances overlap.
This is an operational reminder that high availability is not a static architecture diagram. A system can be resilient to infrastructure failure and still be fragile during routine change. AZ-104 candidates should read update and maintenance requirements alongside zone and scaling requirements, because the best availability design is the one that remains available during both unexpected failures and expected administrative work.