Description
Summary:
Improvado is seeking a DevOps/SysOps engineer to manage and evolve cloud infrastructure, ensuring reliability and scalability for an AI-powered marketing intelligence platform.
Highlights:
1. Manage and evolve cloud infrastructure on AWS/Azure (Kubernetes)
2. Design and architect scalable, reliable infrastructure
3. Drive infrastructure security and optimize cloud costs
Improvado is an AI\-powered marketing intelligence platform trusted by enterprise brands like ASUS, Activision, Docker, and H\&R Block. We raised $34M Series A and are scaling fast — which means our infrastructure needs to be rock\-solid. We're looking for a DevOps/SysOps engineer who takes ownership and keeps things running.
**What you'll do**
* Manage and evolve our cloud infrastructure on AWS/Azure (Kubernetes) \- ensure cluster reliability: capacity planning, autoscaling, incident response, and post\-mortems
* Design and architect scalable, reliable infrastructure to support AI\-driven analytics and data processing at scale
* Build and maintain Helm charts, Terraform, and multi\-environment setups
* Own monitoring and alerting across the stack (Prometheus/Mimir, Grafana, CloudWatch)
* Administer and optimize storages: PostgreSQL, ClickHouse, Redis
* Support message brokers: RabbitMQ, AWS SQS, Temporal
* Drive infrastructure security: secrets management, IAM policies, vulnerability scanning, and encryption
* Own network design, configuration — VPCs, subnets, firewalls, load balancers, VPNs
* Monitor and optimize cloud costs — identify waste, right\-size resources, and report on spend efficiency
* Participate in on\-call rotation
**What we're looking for**
* 5\+ years in a DevOps/SRE role
* Solid hands\-on experience with most of our stack (80%): AWS/Azure/GCP, Kubernetes (cluster design, capacity planning, and reliability at scale), Helm charts, CI/CD pipelines (GitHub CI), Terraform, PostgreSQL, ClickHouse, Redis, RabbitMQ, AWS SQS, Temporal, and monitoring/alerting stacks (Prometheus/Mimir, Grafana, CloudWatch)
* AI\-assisted development in practice — actively uses Claude Code / agentic coding to ship, knows where AI\-generated code needs validation
* Strong Linux fundamentals
* Bash and/or Python scripting
* Comfortable with ambiguity and high\-velocity, fast\-shifting priorities — we move like a startup, not a committee
* Detail\-oriented and pedantic, with sharp attention to detail in high\-volume, high\-stakes systems — accountable and genuinely invested in the quality of your work
* EST timezone availability
* Fluency in English
**Nice to have**
* Golang or Python development background
* HashiStack: Vault, Packer
* ClickHouse administration
* Experience supporting data\-intensive pipelines
**What we offer**
* Remote\-first environment
* 20 days PTO \+ US holidays
* Optional relocation support to Latin America
* Modern AI\-native tech stack
* Stock options
* Professional development reimbursement
* A genuinely fun, transparent startup culture
dOPxGCqsbY