Auditing Information Systems Operations and Resilience

Operational controls are easiest to appreciate after one fails. A server can be patched on schedule yet unavailable because a dependent identity service is down; a backup can complete every night yet be unusable when restoration is attempted; a well-documented incident procedure may collapse because no one knows who has authority to take a system offline. The ISACA CISA certification includes Information Systems Operations and Business Resilience because reliable service delivery depends on coordinated people, procedures, technology and evidence. Auditors in this domain need to follow a business service through normal operations, disruption and recovery without mistaking a dashboard indicator for proof that the service is controlled.

Map critical services and their dependencies

Begin with what the organization must deliver: payments settled, orders fulfilled, clinicians able to view records, employees paid. Map each critical service to applications, databases, identity providers, networks, suppliers, physical locations and support staff. This matters because an apparently healthy application can fail when a shared DNS service or cloud authentication dependency becomes unavailable. A dependency map should distinguish what is actively monitored from what is merely assumed to work and identify where redundant systems still share a single underlying point of failure.

Review service catalogs, asset inventories, configuration data and architecture diagrams, then verify a sample against reality. A change in infrastructure often outpaces documentation. The auditor should not demand one perfect configuration-management database; instead, assess whether information is accurate enough to support incident diagnosis, change impact analysis and recovery. In an audit of a core customer service, an undisclosed dependency on a single external API may be more significant than several outdated labels in an inventory.

Service ownership also matters. Who approves planned downtime, understands the service’s acceptable loss of data and can authorize emergency repair? If an operations team owns servers while an application team owns customer impact, incident leadership may become fragmented. Control design should show how those groups make decisions together. Review incident or change records for evidence that the documented ownership works when priorities conflict.

Evaluate monitoring, event management and escalation

Monitoring should identify conditions that threaten service objectives, not merely produce large quantities of alerts. CPU usage can be useful, but transaction failure, authentication error rates, queue growth and end-to-end latency may be closer to business impact. An auditor considers whether thresholds are based on meaningful operational needs, whether alerts reach a responsible responder and whether teams can separate transient noise from an incident requiring action. An alerting system that floods on-call engineers every night may be less reliable than a smaller set of prioritized signals.

Examine event records across the full process: detection, triage, escalation, resolution and review. Sampling only closed critical incidents can miss alert failures or tickets wrongly classified as low priority. Compare independent sources where practical, such as monitoring event histories, help-desk records and customer complaints. Determine whether severe incidents were escalated within agreed thresholds, whether containment steps were recorded and whether status communications were accurate. A well-organized incident timeline can expose gaps between an automated alarm and the first effective human response.

Log management has both security and operational purposes. Audit whether logging captures significant administrative actions, configuration changes, authentication and relevant application events. Check retention, clock synchronization, tamper resistance and who may delete records. Logs that exist only on a failed server may be unavailable when most needed. However, increasing retention without privacy controls can create unnecessary exposure. A sound design balances investigative needs, cost, regulatory requirements and access restrictions.

Review incident, problem and service restoration controls

Incident management aims to restore service and limit impact; problem management looks for underlying causes that make incidents recur. The same outage can generate both kinds of records, but their success criteria differ. An incident may be resolved by rerouting traffic, while the underlying fault in certificate renewal remains. Review whether the organization distinguishes temporary workarounds from permanent corrections and whether recurring incidents are analyzed rather than treated as unrelated tickets.

A useful case is a payroll application that becomes inaccessible during month-end processing. Operators restart the application and close the incident. If the problem recurs every month, the audit question is whether monitoring, capacity planning and root-cause analysis were sufficient to detect a repeating control weakness. Review escalation to business owners, verification that service was actually restored, and follow-up action after the emergency. A ticket marked resolved when the service was merely restarted may overstate recovery effectiveness.

Assess major-incident authority and communications. Someone must decide when to activate a bridge, notify customers, contact a supplier or begin continuity procedures. Procedures should permit urgent action but still preserve records of decisions. Emergency access can be appropriate for restoring service; after the incident, its use needs reconciliation and independent review. An emergency should not erase segregation-of-duties expectations indefinitely.

Test change, configuration and release management

Changes to production systems should have an identifiable request, impact assessment, appropriate approval, implementation record and verification. The level of control should be proportional to risk. A routine low-impact configuration adjustment may follow an approved standard change, while a database schema migration affecting financial reconciliation needs deeper testing and a recovery approach. Auditors should verify that teams apply these categories consistently rather than labeling most changes as standard or emergency to avoid scrutiny.

A complete test population includes authorized releases and changes performed outside the formal system. Reconcile ticket records with CI/CD deployments, privileged activity logs or configuration histories where feasible. A sampled approved ticket does not prove that every production change was authorized. Watch for retrospective approvals that make an uncontrolled change appear compliant. Examine whether unsuccessful changes resulted in the documented rollback or forward-fix decision and whether the actual live state was reconciled with the configuration record.

Automated pipelines can strengthen control by enforcing peer review, tests and identity restrictions, but they are not inherently trustworthy. Who may modify the workflow definition? Are production credentials protected? Can a developer bypass required checks by editing a branch protection setting? The audit needs to evaluate both the automation and the controls over automation itself. A green build status is limited evidence if tests are routinely skipped or a privileged service account can deploy arbitrary artifacts.

Examine backup design and real restoration capability

Backups are valuable only in relation to recovery objectives and tested restorability. Recovery point objective describes acceptable data loss measured in time; recovery time objective describes the required service restoration time after disruption. These should be derived from business impact analysis, not selected solely because a vendor offers a particular backup interval. A system that must lose no more than fifteen minutes of transactions may need more than nightly backups, and its ability to restore depends on databases, encryption keys, network access and application dependencies.

Review backup coverage, schedules, failures, retention, off-site or isolated copies, access protection and immutability where appropriate. Ransomware can compromise backup infrastructure if backup identities and management paths share excessive privilege with production. Auditors should understand whether an attacker with production administrator permissions could delete or encrypt recovery copies. Encryption protects confidentiality, but lost keys can make perfectly intact backups unusable. A recovery design must protect key access and key availability under disaster conditions.

A restore test should recreate a meaningful business capability, not just prove a compressed file can be opened. Review whether representative data was restored, whether integrity checks passed, whether applications started in the right order and whether users could perform essential transactions. A full regional failover exercise may be expensive, but realistic incremental testing should still demonstrate progress toward the stated objective. Record actual timings, dependencies, data reconciliation and unresolved problems. A completed exercise that exceeded the recovery objective is evidence of a gap, not automatically a success.

Distinguish disaster recovery from business continuity

Disaster recovery focuses on restoring technology capabilities, while business continuity covers how essential operations continue through a disruptive event. A business may temporarily accept manual procedures even when its main platform is unavailable. Audit should examine the relationship among business impact analysis, continuity strategy, technology recovery plans, supplier commitments and crisis communications. It is possible to have excellent database recovery and no viable way to process orders while the warehouse is inaccessible.

Review whether plans account for staff unavailability, geographic outages, utility failures, cyberattacks, compromised identity services and concentration risk in vendors. A primary and secondary system hosted in separate data centers may still rely on one identity tenant or shared connectivity path. Test scenarios that challenge those common dependencies. The plans should specify who declares a disaster, who authorizes recovery spending, how priorities are determined and when normal operations can resume safely.

Exercise evidence is crucial. Tabletop discussions test decision making and communication, while technical tests validate execution. Neither one alone proves complete resilience. A good auditor compares the selected exercises with the organization’s most consequential failure scenarios and examines whether lessons were assigned, funded and retested. The objective is not a thick binder of procedures but an organization able to make coordinated decisions when ordinary assumptions no longer hold.

Review cloud operations and outsourced service management

Cloud operations introduce shared responsibility, dynamic assets and often automated deployment. The organization may delegate physical maintenance to a provider while retaining responsibility for configuration, workload identity, application monitoring and data recovery. Review how teams track changes to cloud services, verify service limits, control administrator access and understand the provider’s recovery commitments. A platform’s stated availability target is not proof that a customer’s multi-component workload will meet its own recovery objective.

Supplier controls should cover incident notification, service monitoring, access management, exit planning and assurance evidence. An auditor may inspect current service reports or independent assurance statements, but must review their scope and exceptions. If a provider excludes a critical subcontractor, the resulting assurance gap needs attention. Contracts should make important responsibilities clear, particularly which party initiates a restore, preserves logs or communicates a data incident. The strongest evidence shows that obligations have been exercised in practice, not simply stated in procurement documents.

Pay attention to cost anomalies when they signal operational defects. Runaway autoscaling, unexpected data transfer or orphaned resources can reflect a weak change process or compromised identity. Financial monitoring alone is not a security control, but it can complement technical signals. Evaluate whether operations teams investigate unusual changes rather than dismissing them as routine cloud variability.

Audit maintenance, capacity and operational records

Reliable systems require patching, lifecycle planning, capacity management and job scheduling. Review whether critical components have supported versions, defined patch responsibilities, testing and exception processes. A missed patch on an isolated development system may create different exposure from an unpatched publicly reachable authentication service. Auditors should evaluate prioritization criteria instead of assuming that every update has the same urgency. Maintenance windows must also reflect business dependencies so that patching one service does not inadvertently disable several downstream applications.

Capacity planning should consider forecasts, demand spikes, storage growth and service limits. A retail platform that fails on a predictable sales holiday may show weak forecasting even if its infrastructure performed normally on ordinary days. Review utilization history, stress-test results and decisions triggered by warning thresholds. Scheduled jobs such as financial reconciliations need monitoring for completion, exceptions and reruns; a job marked successful despite processing incomplete input is a control failure that may not be detected by infrastructure monitoring.

Operational records should make it possible to reconstruct a consequential event. Logs, runbooks, on-call handoffs, change approvals and supplier communications need consistent timestamps and appropriate retention. Excessive documentation can also hinder response if employees cannot find current procedures during an emergency. A realistic test asks an alternate engineer to follow the runbook without help from its author. If the process depends on undocumented tribal knowledge, the organization has a continuity weakness.

Reason through CISA operations and resilience scenarios

CISA scenarios reward attention to the business objective behind a control. When asked whether a backup arrangement is sufficient, examine recovery requirements and restoration evidence rather than counting copies. When a major outage is resolved, consider whether underlying causes, customer impact and emergency access were reviewed. When a monitoring dashboard shows no failures but customers report transactions missing, question the completeness and integrity of monitoring data. The best first action is often to understand the actual condition and relevant objective before recommending a new tool.

A well-supported audit conclusion separates an immediate exception from a systemic failure. One missed service job may be an isolated event; repeated silent failures without reconciliation indicate a process weakness. A plan that has never been exercised may be well drafted but has not demonstrated effectiveness. Audit recommendations should describe the control objective, the exposure that matters and evidence that would establish remediation. Operations and resilience auditing is ultimately about whether the organization can deliver essential outcomes predictably, detect failure promptly and recover within the limits its stakeholders actually require.