Faster chat, better deals — Get the App

Senior Site Reliability Engineer

Indeed

Company

Job typeFull-time
Workplace typeOnsite
Experience levelNo experience limit
Education levelNo degree limit

Description

Summary: Seeking a Senior Site Reliability Engineer to monitor and maintain infrastructure, with expertise in Dynatrace and Splunk, and participation in on-call support. Highlights: 1. Monitor and maintain health, performance, and reliability of applications 2. Expertise in Dynatrace and Splunk for observability and monitoring 3. Engage in incident response and root cause analysis We are looking for a **Senior Site Reliability Engineer** to embed with their team and support operations across their infrastructure landscape. The primary monitoring, observability and logging tools used across all client's services are Dynatrace and Splunk. The role includes on\-call support — details to be discussed. **Responsibilities** * Monitor and maintain the health, performance, and reliability of the client's applications and services * Operate and continuously improve Dynatrace and Splunk, including dashboards, alerting, anomaly detection, and log analysis * Investigate alerts, identify probable root causes, and provide actionable recommendations to engineering and incident management teams * Participate in incident response and on\-call support activities * Conduct Root Cause Analysis (RCA) and drive post\-incident improvements * Define and track SLOs, SLAs, and error budgets * Identify and close observability and monitoring gaps across services * Maintain operational runbooks and support documentation **Requirements** * 3\+ years of experience in Site Reliability Engineering, Production Operations, DevOps, or a similar role * Expertise in Dynatrace and Splunk, including APM, alerting, dashboards, RUM, synthetic monitoring, service flow analysis, SPL queries, and log analysis * Experience with production incident management and on\-call support, including alert triage, incident response, RCA, and post\-incident reviews * Skills in troubleshooting and RCA across distributed applications and services * Understanding of application architecture, service dependencies, integrations, performance analysis, dependency mapping, and bottleneck identification * Experience with AWS services, including CloudWatch, ECS, EC2, ALB, Route53, RDS, and VPC * Familiarity with CI/CD pipelines and release validation processes * English proficiency at B2 level or higher **Nice to have** * Knowledge of AI\-assisted observability capabilities * Experience tuning monitoring and alerting strategies * Familiarity with Infrastructure as Code (Terraform or equivalent) * Travel/Airline industry experience * Background in modernization and cloud transformation initiatives **We offer** * International projects with top brands * Work with global teams of highly skilled, diverse peers * Healthcare benefits * Employee financial programs * Paid time off and sick leave * Upskilling, reskilling and certification courses * Unlimited access to the LinkedIn Learning library and 22,000\+ courses * Global career opportunities * Volunteer and community involvement opportunities * EPAM Employee Groups * Award\-winning culture recognized by Glassdoor, Newsweek and LinkedIn

Source: indeed

Posted by

Sofía González

Indeed · HR

Location

Sofía González

Indeed · HR

Similar jobs

Senior Site Reliability Engineer job by Indeed in 2026 | ok.com