Azure Site Recovery Architecture

Azure Site Recovery is often introduced as a disaster recovery service, but the architectural questions begin before replication is enabled. You need to decide which workloads require regional recovery, what recovery point and recovery time objectives are acceptable, where the recovery vault and target resources belong, how network identity changes during failover, and how recovery will be tested without disrupting production.

For AZ-104 and AZ-305 candidates, the most important distinction is between availability and disaster recovery. Availability protects a service from localized failure while it continues to run. Site Recovery is about restoring service in another location after a larger failure. Designing one does not automatically provide the other.

Begin with RTO and RPO, not with replication settings

Recovery time objective describes how quickly a service must be restored after disruption. Recovery point objective describes how much recent data the business can afford to lose. These are business requirements expressed as technical constraints. They determine whether asynchronous replication is acceptable, how much automation is required, and how frequently recovery processes must be validated.

A design that meets a very short RTO may need pre-created network components, capacity planning, automated failover steps, and a well-rehearsed runbook. A looser RTO may permit more infrastructure to be created during recovery. The RTO and RPO distinction should be clear before you evaluate any disaster recovery product.

Place recovery components around the failure domain

A regional disaster recovery design should not depend entirely on the region it is meant to survive. Microsoft guidance for Azure-to-Azure Site Recovery recommends deploying the Recovery Services vault in the target region for replication. The architecture also needs target-region networking, identity access, DNS behavior, and any dependent services that must be available during failover.

Do not think of the recovered VM as an isolated object. Applications depend on load balancers, private endpoints, databases, domain services, secrets, routes, firewalls, monitoring, and external integrations. The recovery design must identify which dependencies are replicated, which are rebuilt, and which are globally available already.

Design the target network before the outage

Site Recovery can create or connect recovered machines to target virtual networks, but network design determines whether users and dependent systems can actually reach the recovered application. Address ranges, subnet names, network security groups, route tables, DNS zones, private connectivity, and hybrid routing should be planned in advance.

If the target network overlaps with on-premises address space or another connected network, failover can succeed technically while the application remains unreachable. Recovery testing should therefore include end-to-end traffic, not just the state of replicated virtual machines.

Protect the replication path and cache design

Azure-to-Azure replication uses supporting resources, including cache storage, and high-change workloads can have different replication demands from ordinary servers. Microsoft recommends zone-redundant storage for the cache storage account in production recovery designs where supported. High-churn support is available for workloads with greater data-change rates.

The architectural lesson is to size and place replication components for the data change profile. A database server with heavy writes cannot be treated like a mostly static web server. RPO objectives are only credible when the replication design can keep pace with change under normal and degraded conditions.

Capacity in the recovery region is part of the plan

A failover requires compute capacity in the target region. During a widespread regional event, many customers may be trying to start resources in the same alternate location. Microsoft recommends considering on-demand capacity reservations for production workloads where guaranteed target capacity is important.

This is a classic architecture tradeoff: reserved recovery capacity costs more than assuming capacity will be available when needed, but it reduces a risk that only appears during the event the design is supposed to handle. Criticality and RTO should determine whether that assurance is justified.

Use recovery plans to sequence dependencies

Multi-tier applications rarely recover correctly if every virtual machine starts at the same time. Identity or database services may need to become available before application tiers, and network or script-based tasks may need to run between groups. Recovery plans help organize that sequence and provide a repeatable failover workflow.

Sequence should reflect application dependencies, not server importance alone. A small service that provides authentication or configuration may be a prerequisite for a larger application tier. Documenting those relationships is one of the main benefits of disaster recovery design even before a real outage occurs.

Test failover without confusing testing with readiness

Site Recovery supports test failover so teams can validate recovery without committing to a production failover. Microsoft recommends regular disaster recovery drills for production workloads. A test should prove more than whether a VM starts. Validate application health, authentication, data consistency, routing, DNS, monitoring, and the operational handoff to the team that would run the service in the target region.

Record the results and update the runbook. A disaster recovery plan that has never been exercised is an assumption. Repeated drills turn it into an operational capability.

Plan failback as carefully as failover

Recovery does not end when users reach the target region. The original region may return, and the business will eventually decide whether to remain on the recovered environment or fail back. Data may have changed significantly while the target was active, so reverse replication and planned failback require care.

Architecture documents should describe who authorizes failback, what conditions must be met, how data synchronization is verified, and how user traffic is switched again. Treating failback as an afterthought can create a second outage after the first crisis has passed.

Monitor replication health before an incident

Replication warnings, agent health, update status, storage issues, and lag should be monitored continuously. Microsoft recommends keeping Site Recovery components current and configuring alerts so replication problems are discovered before they become recovery failures.

This is part of the broader Azure infrastructure certifications mindset: a recovery feature that is configured once and ignored will drift. Operational ownership is part of architecture because someone must respond when protection degrades.

Exam focus: map the failure to the recovery boundary

When an exam scenario mentions a zone failure, region failure, data loss tolerance, or recovery deadline, identify the failure domain first. Then determine whether the requirement is availability, backup, disaster recovery, or a combination. Site Recovery is appropriate when workload replication and orchestrated failover address the stated failure.

Within the wider cloud architecture certifications, this is a transferable skill: design from business continuity objectives, include every dependency needed for service restoration, and test the recovery process often enough that the organization can trust it.

Build recovery documentation around an application rather than around individual virtual machines. List the user entry point, DNS records, certificates, identity dependencies, databases, message systems, storage, network routes, firewalls, monitoring, and external integrations. Then mark which components are replicated by Site Recovery, which are protected by another service, and which must be recreated. This exercise reveals hidden dependencies before a real incident does.

Test failover should use realistic traffic and realistic people. A technical team may prove that a server boots, while the service desk, application owner, or business user later discovers that authentication, a batch integration, or a partner connection does not work. Include representative consumers in drills and measure actual recovery time from the start of the exercise to usable service, not merely to VM startup.

Recovery plans should also include decision points. Who has authority to declare a disaster? Who approves DNS cutover? At what point is failback considered safe? What evidence proves that the target environment is healthy enough to become primary? These questions are operational, but they affect architecture because automation, permissions, and monitoring must support the people making those decisions.

Cost should be reviewed explicitly. Disaster recovery can range from minimal replicated storage with infrastructure created during failover to warm or hot secondary environments with reserved capacity. The correct design is not the one with the most redundancy; it is the one that meets the business continuity objective at a cost the organization accepts. Architecture is strongest when the tradeoff is documented before the outage rather than debated during it.

It is also worth separating disaster recovery testing from backup restore testing. Site Recovery validates workload replication and failover, while backup protects recoverable copies of data. A ransomware event or logical corruption may require an earlier clean recovery point rather than the most recent replicated state. Critical systems often need both capabilities because they address different failure modes.

For architecture review, ask one final question: what could make the recovery plan fail even if replication is healthy? Common answers include missing permissions, expired certificates, unavailable DNS, insufficient target capacity, incompatible network routes, or an undocumented external dependency. Designing and testing around those dependencies is what turns Site Recovery from a configured feature into a credible business continuity solution.

Recovery priorities may also differ by application tier. A customer-facing transaction path might need rapid recovery, while reporting or batch systems can wait. Grouping workloads by business criticality helps avoid spending the same amount on every server and gives operations a clear order during a large incident.