Senior Site Reliability Engineer (SRE)

Indeed

Company

Job typeFull-time
Workplace typeOnsite
Experience levelNo experience limit
Education levelNo degree limit

Description

Summary: We are seeking a skilled and proactive Site Reliability Engineer (SRE) to apply software engineering principles to automate operations, scale infrastructure, and ensure systems are highly available and performant. Highlights: 1. Design, build, and maintain cloud infrastructure using modern IaC practices 2. Build and optimize CI/CD pipelines to automate software deployments 3. Work with global teams of highly skilled, diverse peers We are looking for a skilled and proactive **Site Reliability Engineer (SRE)** to join our engineering team. In this role, you will bridge the gap between software development and systems operations, applying software engineering principles to automate operations, scale infrastructure, and ensure systems are highly available, resilient, and performant. Your mission is to build, run, and protect the production environments that power our applications, minimizing downtime and helping us deploy software rapidly and safely. **Responsibilities** * Design, build, and maintain cloud infrastructure using modern IaC practices such as Terraform and CloudFormation * Build and optimize CI/CD pipelines to automate software deployments, configuration management, and repetitive operational tasks * Design and implement robust logging, monitoring, and alerting systems using tools such as Prometheus, Grafana, and Datadog * Establish clear Service Level Objectives (SLOs) and Service Level Indicators (SLIs) * Respond to production incidents and lead troubleshooting efforts to restore services * Conduct blameless post\-mortems to identify root causes and prevent recurrence * Partner with software developers to optimize system performance and plan capacity * Ensure services can scale to handle growth and traffic spikes **Requirements** * 3\+ years of experience in systems administration, DevOps, or systems\-focused software development * Proficiency in at least one scripting or programming language such as Python, Bash, Go, or Rust * Experience with public cloud providers such as AWS, Azure, or GCP, and containerization tools such as Docker and Kubernetes * Understanding of Linux/Unix administration and networking fundamentals such as TCP/IP, DNS, and HTTP/SSL/TLS * Passion for automation, eliminating toil, and building resilient systems that fail gracefully * Advanced proficiency in English (C1\+) **We offer** * International projects with top brands * Work with global teams of highly skilled, diverse peers * Healthcare benefits * Employee financial programs * Paid time off and sick leave * Upskilling, reskilling and certification courses * Unlimited access to the LinkedIn Learning library and 22,000\+ courses * Global career opportunities * Volunteer and community involvement opportunities * EPAM Employee Groups * Award\-winning culture recognized by Glassdoor, Newsweek and LinkedIn

Source: indeed

Posted by

Sofía González

Indeed · HR

Location

Sofía González

Indeed · HR

Similar jobs

Senior Site Reliability Engineer (SRE) by Indeed in 2026 | ok.com