Software Expert SRE

Company
Description
Job Summary: As a Software Expert at Mercado Libre, you will design and scale innovative and secure systems, leading structural initiatives and building a robust SRE culture powered by AI. Key Highlights: 1. Design and scale innovative and secure systems using cutting-edge technology. 2. Lead initiatives to improve resilience, uptime, and scalability. 3. Mentor teams and team leads, fostering a strong SRE culture. As a **Software Expert** at Mercado Libre, you will design and scale innovative and secure systems that solve real-world, high-impact problems. You will work in a dynamic environment leveraging cutting-edge technology, applying sound engineering practices, proprietary AI models, and continuous learning — all aimed at democratizing e-commerce and financial services across Latin America. Imagine tackling challenging, dynamic, and innovative projects while being **responsible for:** * Leading structural initiatives to improve resilience, uptime, and scalability, and governing resilient design patterns and cross-cutting best practices. * Designing secure, resilient, business-aligned architectures, and designing and integrating AI-based solutions to optimize system reliability — including proactive anomaly detection, incident prediction, and intelligent automation of operational responses. * Modeling and prioritizing uptime based on service criticality, and assessing technical maturity to design domain-specific evolution plans. * Driving advanced observability and custom instrumentation practices, and promoting adoption of AI tools for tasks and automations. * Identifying technical bottlenecks through deep profiling and troubleshooting. * Leading and elevating technical rigor in war rooms and postmortems, and mentoring teams and team leads to foster a strong SRE culture. **What are we looking for?** * Minimum 5 years of experience in software development, with expertise in distributed systems, microservices, scalable APIs, and cloud architectures. * Experience managing uptime, high availability, and scalability — with both preventive and reactive incident response approaches — and deep knowledge of uptime management, incident response, and SRE reliability engineering practices. * Proven ability to lead continuous improvement processes, including operational failure and problem management, and mastery of executive summary writing, KPI management, and technical objective definition. * Advanced proficiency with observability and troubleshooting tools — such as Datadog, New Relic, Kibana — and experience designing and evolving advanced observability systems, including custom instrumentation strategies, end\-to\-end traceability, correlation of technical and business metrics, effective threshold definition for alerting, and focus on efficient diagnosis. * Experience with event sourcing, and experience designing and integrating AI/ML solutions applied to operations — e.g., anomaly detection, capacity forecasting, automated incident classification, or LLM-assisted runbook generation — ability to design systems that capture and store key business events with full data traceability as input for AI models, and familiarity with AI ecosystem frameworks and tools oriented toward reliability and infrastructure use cases — e.g., Python, scikit\-learn, and MLOps platforms. **Ready to leave your mark on Latin American technology?** Apply now and join our mission! Hybrid work mode. Site: Polo Dot, Saavedra, Buenos Aires.
Posted by

Sofía González
Indeed · HR




