AWS SCS-C03: Detect, Contain, and Recover from Cloud Incidents

An AWS account begins launching compute instances in an unfamiliar Region. GuardDuty produces a finding, the finance dashboard shows an unusual cost spike, and a CloudTrail query identifies an assumed role. Is this an approved scaling test, stolen credentials or a malicious workload created by an attacker? A cloud incident responder needs more than one alert: they must establish the identity path, preserve the relevant records, limit damage and restore trusted operation. Those skills are central to AWS Certified Security – Specialty SCS-C03.

The active SCS-C03 blueprint separates Detection from Incident Response, assigning 16% and 14% of scored content respectively. Detection covers monitoring, logging, alerting and troubleshooting those systems; response covers preparation, investigation, containment, eradication and recovery. That distinction is useful in practice. A detection architecture may tell the team something suspicious is happening, but only a prepared response process can turn a finding into a controlled resolution.

Build observability before a suspicious event occurs

A useful cloud security record includes who changed infrastructure, which resources were affected, which actions succeeded and what happened inside workloads. AWS CloudTrail records supported account API activity; AWS Config can help evaluate resource configuration and changes; VPC Flow Logs capture selected network-flow metadata; application and operating-system logs show behavior that cloud control-plane services may not see. None is a complete substitute for the others.

CloudTrail has distinct categories of management, data, network activity and Insights events. Trails and event data stores do not log every optional category by default. An investigator interested in S3 object reads may find that management events show a bucket-policy change but not the individual GetObject requests unless appropriate data-event collection was configured. The automatically available Event history is a useful recent management-event view, not a replacement for durable, appropriately scoped organization-wide security logging.

Centralized collection should account for every relevant account and Region, expected event volumes, retention requirements and the permissions needed to prevent an attacker from modifying evidence. Send logs to appropriately protected destinations with limited write and delete privileges. Encrypt records when needed, monitor ingestion failures and test whether a new account is covered automatically. A dashboard with attractive graphs is not a reliable detection system if one production account silently stopped reporting last week.

Choose findings that drive an investigation

Amazon GuardDuty analyzes supported telemetry for potentially malicious activity and produces findings for investigators. AWS Security Hub can help aggregate and prioritize security findings and posture information, while Amazon Detective can assist with exploring relationships during investigation. These services have different roles; turning on all of them does not guarantee that every intrusion will be detected. Coverage depends on supported resources, enabled protections and whether the activity generates observable signals.

Imagine an API call from an unfamiliar network that creates a new access key. The event might be an approved automation change, but it is significant enough to validate against change records, principal history and subsequent use. Look for related behavior: role assumptions, permissions changes, new resource creation, attempts to disable logging or access to sensitive storage. A single anomalous IP address is weak evidence; a coherent sequence across accounts and resources is stronger.

Detection confidence and business impact should be reported separately. A verified malicious command on an isolated test instance may be high confidence but limited in consequence. A weak indicator involving the account that protects production backups may require urgent review because potential impact is high. Use context to prioritize cases, not the color displayed in the finding. When multiple tools report the same underlying event, deduplicate it without losing distinct evidence.

Correlate identity, resources and time carefully

Cloud APIs are often made using short-lived role sessions rather than permanent user identities. To understand an event, trace the assumed role, session name, source principal, trust relationship and any relevant federation context. A compromised CI workflow may use a legitimate deployment role, making the API caller look like approved automation. Investigators need to compare token issuance, source environment and deployment records before concluding that a known role name means a known human initiated the activity.

CloudTrail event order should not be treated as a perfectly ordered program trace. Services emit events under different conditions, collectors ingest at different times and some workload logs use their own clocks. Normalize event times but retain original fields. Join records through account IDs, resource ARNs, request identifiers and role sessions where possible. An address alone may represent NAT or a shared service rather than one unique instance.

The attack lifecycle can help generate hypotheses about privilege escalation, persistence and exfiltration, but those hypotheses require evidence. A role modified just before a new EC2 instance appears may be related, or they may be independent administrative changes. Investigators should document confidence, competing explanations and missing logs. A rigorous timeline distinguishes confirmed actions from the narrative built around them.

Contain the smallest effective blast radius

Containment in AWS may involve revoking or restricting credentials, updating IAM policies, limiting network connectivity, quarantining a workload or applying an organization-level guardrail. The safest option depends on what the attacker controls and what business function the resource provides. Abruptly disabling a shared role can stop many legitimate applications. Merely blocking one observed attacker address may do nothing if temporary credentials remain valid and can be used from other networks.

For an instance suspected of running malware, consider whether forensic evidence must be captured before termination. An isolation security group can reduce network exposure while investigators collect snapshots or other approved artifacts, but some response functions may require controlled outbound access. The sequence should be determined by a tested runbook. For a stolen identity, session validity and resource-policy access need explicit attention; changing a policy may not have the same effect as invalidating credentials or ending dependent workloads.

Use automation where the outcome is known and reversible. EventBridge rules can route security findings into response workflows involving Lambda, Systems Manager or other orchestrators. A high-confidence event might justify an automatically applied quarantine in a test environment. For a production database or shared networking account, initial automation may only open an incident, enrich context and request authorized approval. Automated remediation needs protection against loops, incorrect triggering and simultaneous actions from different tools.

Preserve evidence without recreating the attacker’s access

Cloud forensics may involve saving CloudTrail and application logs, recording instance metadata, taking EBS snapshots and collecting relevant configuration changes. The team should document which evidence belongs to which account, Region, resource and point in time. Capture hashes where they are meaningful, restrict investigators’ permissions and store working copies in a protected account or repository. A snapshot is not a complete memory image, and a copy of a disk does not include all the cloud control-plane activity that created or modified the instance.

For IAM investigations, preserve the previous and current trust policy, identity policy changes, session history and affected resources. In network investigations, capture route and security-group configuration as it existed during the event if that can be reconstructed. Logs and configuration records can be lost through retention limits or overwritten changes, so evidence collection procedures should run early. Do not grant the compromised account additional access merely to make a collection script easier to execute.

Response reports must avoid overstating outcomes. A GuardDuty finding may support the conclusion that a known malicious pattern was observed, but it does not by itself establish that data was exfiltrated. VPC Flow Logs may show a connection without identifying its contents. An intrusion-detection signal differs from proof of successful exploitation. For every major statement about attacker behavior, record the supporting artifacts and their limits.

Find the control failure before restoring service

Eradication should address how the attacker entered and why the environment allowed the activity to spread. If a build pipeline accepted tokens from an overly broad federation trust policy, rotating one access key does not fix the role’s original weakness. If a public workload exposed credentials through an application error, patch the application and redesign secret handling. If backup permissions enabled deletion across accounts, separate and protect recovery access. Remediation should be specific to the observed failure path.

Recovery plans must also account for infrastructure as code. Recreating a compromised server from an approved image can restore a known baseline, but the template may still contain the vulnerable security group or overprivileged role that made the attack possible. Review the deployment source, not just the running instance. Validate secrets, role trust, networking, monitoring and application behavior before returning the workload to production service.

Successful recovery needs observable acceptance criteria. A security team might require no unexpected role assumptions, restored centralized logs, verified blocked attacker paths and proof that approved users can still complete core transactions. Monitor the environment for recurrence and document residual uncertainty. A “green” deployment pipeline is insufficient if the attacker still has an alternate path through a third-party integration.

Practice a cross-account incident as an integrated problem

Suppose a development account’s CI role begins calling EC2 RunInstances in a Region the organization seldom uses. First verify the finding and identify the exact assumed role session. Check CloudTrail management events, pipeline records and related policy changes. Look for subsequent network connections, new roles and access to sensitive data. Confirm whether optional data-event logs exist for the resources you need to inspect; do not infer that absent data-event records prove no objects were accessed.

Next select containment that addresses the compromised trust or credentials while preserving production deployment continuity. Gather evidence, evaluate the blast radius across connected accounts, then rebuild or remove unauthorized resources. Correct the trust-policy weakness, restrict role permissions and test the blocked path using a controlled identity. Only after confirming logging, monitoring and intended functionality should the incident be closed. The principle of least privilege should be visible in both the repaired role and the response team’s access.

SCS-C03 examines these ideas as part of an architecture rather than isolated product trivia. Design evidence collection before the incident, learn what each service detects, build an accurate timeline, contain the actual trust failure and validate safe recovery. That is the difference between receiving an AWS security notification and operating an effective cloud security response capability.