Site Reliability Engineer
Role in brief
Offchainlabs is seeking a Site Reliability Engineer to manage and scale their blockchain infrastructure. This role involves operating Kubernetes clusters, designing CI/CD pipelines, and ensuring system reliability and security. Candidates with experience in distributed systems, cloud platforms, and a proactive approach to problem-solving will find this role engaging.
About the role
This Site Reliability Engineer position focuses on the operational aspects of Offchainlabs' blockchain technology. The role requires managing production Kubernetes clusters and building infrastructure using tools like Terraform. Responsibilities also include deploying and maintaining Kubernetes environments, troubleshooting applications, and designing CI/CD workflows for both infrastructure and application deployments.
A key part of the role involves designing and operating observability systems using various tools for metrics, logs, and dashboards. The engineer will diagnose complex networking and storage issues in distributed systems and implement secure-by-default infrastructure. Automation of operational workflows through scripting in Python, Go, or Bash is also expected.
Success in this role means contributing to a reliable and scalable blockchain platform. This involves active participation in on-call rotations, responding to incidents, and driving post-mortem analyses to improve system reliability. The engineer will also contribute to architecture reviews and threat models, ensuring security is integrated into system design.
The salary range for this position is between $98,000 and $162,000 USD.
Skills that matter here
- Kubernetes: This role involves operating production Kubernetes clusters and deploying/maintaining Kubernetes environments.
- Terraform: The engineer will build scalable, declarative infrastructure using Terraform or similar tools.
- CI/CD: This position requires designing CI/CD workflows using tools like ArgoCD, GitHub Actions, or CodeBuild.
- Python: The role involves automating operational workflows using scripting or programming in Python.
- AWS: The engineer should be comfortable operating within cloud platforms like AWS, understanding underlying components.
- On-call Rotation: Participation in an on-call rotation is required, including incident response and post-mortem analysis.
Who this role suits
- Someone who is curious about how systems work at a deep level, not satisfied with surface-level fixes.
- An individual who enjoys solving infrastructure problems in unconventional ways and thinking beyond standard patterns.
- A person who takes ownership of their work, collaborates openly, and contributes to a culture of clarity and continuous improvement.
- A candidate comfortable with Linux and shell scripting, and productive in programming languages like Python or Go.
From the employer
What You Will Do
- Operate production Kubernetes clusters and build scalable, declarative infrastructure using Terraform or similar tools.
- Deploy and maintain Kubernetes environments, manage system components, and troubleshoot applications running on the platform.
- Design CI/CD workflows with ArgoCD, GitHub Actions, CodeBuild, or similar tools, covering both infra and app deployments.
- Design and operate observability systems using time-series metrics, logs, and dashboards with tools like Prometheus, Loki, Mimir, Grafana, and CloudWatch.
- Diagnose tough networking and storage issues across complex, distributed systems.
- Implement secure-by-default infrastructure and contribute to architecture reviews and threat models.
- Automate operational workflows using scripting or programming in Python, Go, or Bash.
What You've Done
- Eager to dive into blockchain technology, even if it’s new territory.
- Enjoy solving infrastructure problems in unconventional ways and thinking beyond standard patterns.
- Use tools like k9s or ArgoCD for speed and abstraction, but comfortable dropping into YAML, logs, or low-level debugging when things go sideways.
- Experienced with GitOps-style systems and treating both infrastructure and application delivery as code.
- Have scaled deployment automation using patterns like ArgoCD ApplicationSets or similar tooling.
- Curious about how things work under the hood and not satisfied with surface-level fixes.
- Comfortable in Linux, fluent in shell scripting, and productive in languages like Python or Go.
- Comfortable operating within a cloud platform (e.g., AWS, GCP, Azure), with a strong understanding of the underlying components making it easy to adapt to or migrate across providers.
- Participated in an on-call rotation, responding to incidents, troubleshooting under pressure, and driving postmortems to improve system reliability over time.
- Design systems with security in mind, applying principles like least privilege and threat modeling.
- Bring a strong technical foundation, excellent problem-solving skills, and a genuine commitment to high-quality work.
- Take ownership, collaborate openly, and contribute to a culture of clarity, curiosity, and continuous improvement.
Perks
- Remote-first global workforce + NY office.
- Annual company offsite + team onsites.
- Professional reimbursement program (facilitates industry conference attendance, certifications, and more).
- Medical, dental & vision coverage (US + some other countries).
- 401k retirement plan + company match (US only).
- Wellness stipend.
- Home office set up / ergonomic equipment program.
Questions about this role
What is the remote work policy for this role?
This is a remote-first position, and the company supports a global workforce.
What level of experience is expected for this position?
The company seeks individuals eager to engage with blockchain technology, experienced with GitOps, cloud platforms, and comfortable with Linux and scripting.
What are the primary technical skills required for this role?
Key technical skills include experience with Kubernetes, Terraform, CI/CD tools, cloud platforms like AWS, GCP, or Azure, and programming in Python or Go.