deniz.in

Markets

Weather

Loading weather

· via dev.to (home feed)

AWS DevOps Agent traced injected AWS faults in minutes but chat answers diverged from its findings

A dev.to workshop report shows AWS DevOps Agent root-causing three fault-injected AWS environments in 10–25 minutes, with IaC-aware fix plans — though chat answers sometimes contradicted its own findings.

AWS DevOps Agent traced injected AWS faults in minutes but chat answers diverged from its findings

A hands-on account published on dev.to puts AWS DevOps Agent, Amazon's take on an automated on-call engineer, through three deliberately broken AWS environments — and finds it capable of real root-cause work, with caveats operators should know about.

According to the write-up by Dmitriy Trunov, the labs came from an incident-investigation workshop organized with AWS by the AWS User Group 3City community, which the author ran on 3 October 2026: one setup lab, then three environments with pre-injected faults. AWS positions the DevOps Agent as a "frontier agent" that both resolves and prevents incidents, drawing on CloudWatch metrics, CloudTrail data, VPC Flow Logs, RDS Performance Insights and infrastructure-as-code stacks to deliver root-cause analysis plus a remediation plan.

Three injected faults, three root causes

In the first scenario, the author pushed a t3.micro instance in an Auto Scaling group to 100% CPU via an SSM Run Command. Roughly one minute 48 seconds in, the agent's first finding appeared: the alarm fired correctly but had no alarm actions, and the group had no scaling policy, so it could never scale out. The agent ruled out CPU credit throttling, showed ALB traffic near zero so the spike was not user load, named the exact SSM command behind the burn, and said plainly that blocked CloudTrail stopped it from attributing an IAM principal. Root cause arrived in about ten minutes.

The second lab simulated an attack from the author's laptop: a port scan, SSH login attempts and an HTTP flood. The agent rated the incident high severity, used VPC Flow Logs to identify the source IP and scanned ports, and confirmed 2,764 accepted packets on port 80 as the genuine attack. It also found a security group open to 0.0.0.0/0 on ports 22, 80, 443 and 8080 since stack creation — and, unprompted, determined that three of four earlier firings of the same alarm were benign S3 return traffic, meaning the alarm itself needed tuning. This was the slowest case, at roughly 20–25 minutes.

The third scenario exhausted an RDS MySQL instance with about 140 connections against a 150-connection limit while a Cartesian join pinned the CPU. Using Performance Insights, the agent split the load into 108 idle sessions and 38 average active sessions on the expensive query, traced it all to a single host and database user, and spotted a broken CloudWatch log export along the way. Root cause took around 16 minutes.

Where it stumbled

That RDS investigation is also where the agent contradicted itself. Follow-up questions in chat produced two disagreements with its own mitigation plan — over whether 150 sits above or below MySQL's default max_connections, and over whether wait_timeout would help. The author's conclusion: the structured findings were more careful than the conversational answers, so chat replies should be verified against the report.

Other limits: the agent only sees what its IAM role and service control policies allow (CloudTrail, GuardDuty and RDS logs were blocked in the workshop account, and those gaps surfaced in the output); investigations take 8–16 minutes; and a spurious "Unable to describe support level" error is expected behavior per the workshop.

What worked

None of the conclusions simply restated the alert; each identified the missing link — an unwired alarm, a world-open security group, a host with no connection limit. Every mitigation plan warned that an API-only fix would drift from the CloudFormation stack and included a code-change spec for the template, structured as five phases (prepare, pre-validate, apply, post-validate, rollback) with copy-paste CLI commands. Each report listed its investigation gaps, the safe-deployment policy refused to auto-run destructive security group changes and flagged them for a human, and remediation can be exported as a spec for a coding agent.

Setup takes about five minutes, per the post: an Agent Space defines what the agent can see and do, with an auto-created read role and an optional write-actions role requiring explicit approval for every change. Capabilities extend to secondary AWS accounts, Azure, GitHub, GitLab and Azure DevOps pipelines, telemetry from Datadog, Dynatrace, New Relic, Splunk and Grafana, Slack, ServiceNow and PagerDuty, any MCP server, and remote agents over the A2A protocol.

The author's verdict: run it as a first responder in the on-call path that hands a human a pre-built timeline — not as an autonomous fixer, not yet.

Why it matters

The first hour of a 2 AM incident is mostly reconstructing a timeline across five console tabs. If an agent can compress that to minutes while honoring approval gates, IaC drift and guardrails around destructive changes, the economics of on-call shift noticeably. The chat contradictions are the caution: the conversational layer remains the least reliable part of the system, so structured findings should be treated as the source of truth. This is one workshop run in a temporary account rather than a benchmark, but it is a concrete signal of where agentic cloud operations are heading.

  • #aws
  • #ai-agents
  • #devops
  • #incident-response
  • #cloud-ops

Related posts