Faster chat, better deals — Get the App

Lead Site Reliability Engineer

Indeed

Company

Job typeFull-time
Workplace typeOnsite
Experience levelNo experience limit
Education levelNo degree limit

Description

Summary: Lead Site Reliability Engineer to drive reliable operations across infrastructure, owning monitoring, observability, and logging with Dynatrace and Splunk, and partnering on incident response. Highlights: 1. Drive reliable operations across a broad infrastructure landscape 2. Own monitoring, observability, and logging with Dynatrace and Splunk 3. Partner on incident response and perform Root Cause Analysis We are seeking a **Lead Site Reliability Engineer** to embed with the team and drive reliable operations across a broad infrastructure landscape. You will own monitoring, observability, and logging with Dynatrace and Splunk, and help evolve dashboards, alerting, and analysis practices. You will also join the on\-call rotation and partner on incident response. Apply now. **Responsibilities** * Monitor and sustain the health, performance, and reliability of the client's applications and services * Operate and continuously enhance Dynatrace and Splunk, including dashboards, alerting, anomaly detection, and log analysis * Analyze alerts, determine likely root causes, and deliver actionable recommendations to engineering and incident management teams * Support incident response activities and participate in the on\-call rotation * Perform Root Cause Analysis (RCA) and drive measurable post\-incident improvements * Define and monitor SLOs, SLAs, and error budgets * Detect and remediate observability and monitoring gaps across services * Maintain operational runbooks and supporting documentation **Requirements** * Proven background with 5\+ years in Site Reliability Engineering, Production Operations, DevOps, or a closely related role * Deep expertise in Dynatrace and Splunk, including APM, alerting, dashboards, RUM, synthetic monitoring, service flow analysis, SPL queries, and log analysis * Hands\-on experience with production incident management and on\-call support, covering alert triage, incident response, RCA, and post\-incident reviews * Strong troubleshooting skills and RCA capability across distributed applications and services * Solid understanding of application architecture, service dependencies, integrations, performance analysis, dependency mapping, and bottleneck identification * Practical experience with AWS services, including CloudWatch, ECS, EC2, ALB, Route53, RDS, and VPC * Working knowledge of CI/CD pipelines and release validation processes * English proficiency at B2 (Upper\-Intermediate) level or higher **Nice to have** * Familiarity with AI\-assisted observability capabilities * Experience optimizing monitoring and alerting strategies * Exposure to Infrastructure as Code (Terraform or equivalent) * Travel/Airline industry experience * Background supporting modernization and cloud transformation initiatives **We offer** * International projects with top brands * Work with global teams of highly skilled, diverse peers * Healthcare benefits * Employee financial programs * Paid time off and sick leave * Upskilling, reskilling and certification courses * Unlimited access to the LinkedIn Learning library and 22,000\+ courses * Global career opportunities * Volunteer and community involvement opportunities * EPAM Employee Groups * Award\-winning culture recognized by Glassdoor, Newsweek and LinkedIn

Source: indeed

Posted by

Sofía González

Indeed · HR

Location

Sofía González

Indeed · HR

Similar jobs

Lead Site Reliability Engineer job by Indeed in 2026 | ok.com