Senior DevOps/Site Reliability Engineer
Role in brief
Limit Break, a company focused on digital markets and communities, is hiring a Senior DevOps/Site Reliability Engineer. This role focuses on improving performance and scalability within their AWS-based EKS environment, involving CI/CD, Kubernetes, and security. Candidates with strong backgrounds in SRE, DevOps, or Systems engineering and expertise in AWS, Kubernetes, and automation tools should apply.
About the role
This Senior DevOps/Site Reliability Engineer will be responsible for enhancing the performance and scalability of Limit Break's multi-cluster EKS environment on AWS. The role involves identifying and executing improvements, measuring system health, and troubleshooting production issues. A key aspect is using code to address operational challenges within the company's infrastructure and platform.
The position requires close collaboration with the broader engineering team to ensure testing environments accurately reflect production. A core responsibility is to define and improve SLOs, SLIs, monitoring, alerting, and incident response practices. The engineer will also continuously refine the observability stack, which includes tools like Grafana, Thanos, and Loki, to support global scalability.
Success in this role means proactively addressing system bottlenecks, ensuring robust and scalable infrastructure, and maintaining high operational standards. The engineer will contribute to a secure and efficient environment, working with various tools for monitoring, automation, and security to protect the system in real-time and remediate vulnerabilities.
The salary for this position ranges from $160,000 to $218,000 USD annually.
Skills that matter here
- Kubernetes: The role requires operating and optimizing multiple EKS clusters in a production environment.
- AWS: Extensive experience with AWS services, including Aurora, RDS, and networking, is essential for managing the cloud infrastructure.
- Terraform: This tool is used for infrastructure as code, requiring strong experience in its application.
- Ansible: The role involves using Ansible for automation and configuration management.
- CI/CD: Experience with continuous integration and continuous delivery pipelines is necessary for deploying services and automating workflows.
- Grafana: This tool is part of the observability stack that the engineer will continuously improve for monitoring and alerting.
Who this role suits
- A person with at least five years of experience in SRE, DevOps, or Systems engineering.
- Someone who is comfortable participating in an on-call rotation and troubleshooting production issues.
- An individual who can clearly explain their reasoning and thought processes to colleagues.
- A collaborative team member who works effectively with product engineers and product owners.
From the employer
- Identify, propose and execute improvements to performance and scalability bottlenecks across our multi-cluster EKS environment on AWS.
- Measure systems health, scalability and performance metrics and identify areas of improvement.
- Deploy services and troubleshoot production issues day-to-day, using code to solve broad operational challenges within the Limit Break Infrastructure and Platform.
- Work with the wider engineering team to identify how we can provide the most production-like environment for running both manual and automated testing.
- Define SLOs, SLIs, monitoring, alerting and incident response practices — and continuously improve our observability stack (Grafana, Thanos, Loki) to be ready for worldwide scale.
- 5+ years experience in SRE, DevOps or Systems engineering.
- Strong background in Kubernetes, including operating multiple EKS clusters in production.
- Extensive experience in Terraform and Ansible.
- CI/CD and automation experience with tools such as GitHub Actions, Jenkins, or GitLab CI.
- Solid background in AWS, including experience with Aurora, RDS (MySQL/SQL), and networking.
- Ability to participate in an on-call rotation.
- Effective communication skills to clearly explain your reasoning and thought process.
- Excellent collaboration skills to work closely with product engineers and product owners.
- Implementation of in-house monitoring and observability infrastructure (e.g. Grafana, Thanos, Loki, or equivalents).
- Implementation of ElasticSearch stack or equivalent solutions for capturing logs from all environments.
- Experience with CloudFlare, CDN technologies, and edge/perimeter networking.
- Exposure to cloud security and perimeter tooling such as Wiz (or equivalent CSPM/vulnerability detection), AWS GuardDuty, CloudFlare Zero Trust, and secrets management platforms.
- Experience addressing vulnerabilities — comfortable finding issues, digging deep to root cause, and driving remediation.
- Implement various tools to monitor and protect the environment in real-time.
Questions about this role
What is the remote work policy for this role?
This is a remote position with no specific location restrictions mentioned beyond being remote.
What level of seniority is expected for this position?
This is a senior-level role, requiring at least 5 years of experience in SRE, DevOps, or Systems engineering.
What are some key technologies used in this role?
Key technologies include Kubernetes, AWS, Terraform, Ansible, CI/CD tools like GitHub Actions, and observability tools such as Grafana, Thanos, and Loki.