Description
Summary:
Lead Site Reliability Engineer to drive reliable operations across infrastructure, owning monitoring, observability, and logging with Dynatrace and Splunk, and partnering on incident response.
Highlights:
1. Drive reliable operations across a broad infrastructure landscape
2. Own monitoring, observability, and logging with Dynatrace and Splunk
3. Partner on incident response and perform Root Cause Analysis
We are seeking a **Lead Site Reliability Engineer** to embed with the team and drive reliable operations across a broad infrastructure landscape. You will own monitoring, observability, and logging with Dynatrace and Splunk, and help evolve dashboards, alerting, and analysis practices. You will also join the on\-call rotation and partner on incident response. Apply now.
**Responsibilities**
* Monitor and sustain the health, performance, and reliability of the client's applications and services
* Operate and continuously enhance Dynatrace and Splunk, including dashboards, alerting, anomaly detection, and log analysis
* Analyze alerts, determine likely root causes, and deliver actionable recommendations to engineering and incident management teams
* Support incident response activities and participate in the on\-call rotation
* Perform Root Cause Analysis (RCA) and drive measurable post\-incident improvements
* Define and monitor SLOs, SLAs, and error budgets
* Detect and remediate observability and monitoring gaps across services
* Maintain operational runbooks and supporting documentation
**Requirements**
* Proven background with 5\+ years in Site Reliability Engineering, Production Operations, DevOps, or a closely related role
* Deep expertise in Dynatrace and Splunk, including APM, alerting, dashboards, RUM, synthetic monitoring, service flow analysis, SPL queries, and log analysis
* Hands\-on experience with production incident management and on\-call support, covering alert triage, incident response, RCA, and post\-incident reviews
* Strong troubleshooting skills and RCA capability across distributed applications and services
* Solid understanding of application architecture, service dependencies, integrations, performance analysis, dependency mapping, and bottleneck identification
* Practical experience with AWS services, including CloudWatch, ECS, EC2, ALB, Route53, RDS, and VPC
* Working knowledge of CI/CD pipelines and release validation processes
* English proficiency at B2 (Upper\-Intermediate) level or higher
**Nice to have**
* Familiarity with AI\-assisted observability capabilities
* Experience optimizing monitoring and alerting strategies
* Exposure to Infrastructure as Code (Terraform or equivalent)
* Travel/Airline industry experience
* Background supporting modernization and cloud transformation initiatives
**We offer**
* International projects with top brands
* Work with global teams of highly skilled, diverse peers
* Healthcare benefits
* Employee financial programs
* Paid time off and sick leave
* Upskilling, reskilling and certification courses
* Unlimited access to the LinkedIn Learning library and 22,000\+ courses
* Global career opportunities
* Volunteer and community involvement opportunities
* EPAM Employee Groups
* Award\-winning culture recognized by Glassdoor, Newsweek and LinkedIn