Description
Summary:
Seeking a Senior DevOps Engineer for platform reliability, advanced troubleshooting, and automation, serving as a key escalation point for complex infrastructure and cloud issues.
Highlights:
1. Lead Root Cause Analysis for major or recurring issues
2. Troubleshoot CI/CD pipeline failures and optimize toolchain
3. Plan and execute non-standard, high-risk architectural changes
We are seeking a **Senior DevOps Engineer** to join our team supporting core enterprise platforms through analytical, diagnostic, and project\-based services. This role focuses on platform reliability, advanced troubleshooting, and automation, serving as a key escalation point for complex infrastructure and cloud service issues.
The position requires participation in an on\-call rotation, including weekday coverage (Mon\-Fri 08:00\-16:00 EST) and weekend rotations shared with a regional counterpart to handle escalations and complex deployment issues.
**Responsibilities**
* Serve as the escalation point for advanced incident management, performing deep\-dive troubleshooting of infrastructure, containers, and cloud services to resolve complex or prolonged incidents
* Lead Root Cause Analysis for major or recurring issues, document findings, and implement permanent corrective actions to reduce recurrence
* Troubleshoot CI/CD pipeline failures, maintain infrastructure\-as\-code, and develop and optimize the CI/CD toolchain
* Plan and execute non\-standard, high\-risk, or architectural changes such as cluster migrations and major upgrades
* Analyze system performance, adjust resources, and tweak configurations for stability and cost optimization
* Create and update Standard Operating Procedures (SOPs) and Runbooks
* Conduct training sessions to share knowledge across the team
* Participate in on\-call rotation, operating in stand\-by mode to acknowledge issues within agreed timeframes and actively troubleshooting to resolution when incidents arise
**Requirements**
* 5\+ years of relevant experience in a DevOps or Platform Engineering role
* Proficiency in Windows and Linux operating system administration
* Expertise in scripting with PowerShell and Bash
* Background in web hosting and load balancing using IIS and HAProxy
* Skills in configuration management and automation with Ansible
* Knowledge of Terraform and Helm for infrastructure\-as\-code
* Familiarity with RabbitMQ messaging and Active Directory
* Understanding of security TLS certificates management
* Competency in data platforms, including MS SQL, DWH, and PowerBI
* Capability to support IoT Edge environments
* Flexibility to participate in on\-call rotation, including weekend coverage
* Proficiency in English at an Upper\-Intermediate level (B2\) or higher
**We offer**
* International projects with top brands
* Work with global teams of highly skilled, diverse peers
* Healthcare benefits
* Employee financial programs
* Paid time off and sick leave
* Upskilling, reskilling and certification courses
* Unlimited access to the LinkedIn Learning library and 22,000\+ courses
* Global career opportunities
* Volunteer and community involvement opportunities
* EPAM Employee Groups
* Award\-winning culture recognized by Glassdoor, Newsweek and LinkedIn