Amazon
Capstone Software Engineer
Problem
Engineering teams monitoring AWS environments often need to reason across multiple services, accounts, regions, metrics, dashboards, and logs before they can identify what happened and where to start troubleshooting.
Built / Contributed
Developed an observability platform that automated anomaly detection across multiple AWS services, accounts, and regions. Built Java backend services and a serverless metric ingestion pipeline using EventBridge, Lambda, and DynamoDB, collected metrics through cross-account IAM role assumption, integrated SageMaker for live time-series anomaly detection, and designed a triage interface with direct CloudWatch links to specific timestamps.
Key Metrics & Highlights
- Reduced manual monitoring for engineering teams by 2+ hours daily.
- Collected metrics across AWS accounts through cross-account IAM role assumption.
- Cut log investigation scope from 10,000+ logs to 50 with targeted CloudWatch links.