Senior Site Reliability Engineer

Remote $82k–$112k senior English B2 4 months ago full-time quality 8.6/10

Role in brief

Bloxstaking is hiring a Senior Site Reliability Engineer to build and maintain infrastructure, focusing on automation and reliability. This role involves using modern cloud-native tools and AI to enhance developer productivity and operational efficiency. Ideal for engineers with strong Kubernetes and IaC experience, who are also proficient in Go or Python and skilled in deploying AI-powered developer tools.

KubernetesArgoCDTerraformCrossplaneGrafanaGoPythonLinuxAI tools

About the role

This Senior Site Reliability Engineer role centers on designing and implementing robust infrastructure and tools that enable product teams to develop and deploy applications quickly and securely. The position requires a focus on automation and reliability, contributing to the strategic direction of infrastructure and operational practices to support organizational growth. Success in this role means ensuring seamless operations through proactive incident resolution and continuous improvement.

A key aspect of this position involves collaborating closely with product teams on critical initiatives such as production deployments, release management, and incident handling. The engineer will provide technical expertise to modernize the platform and infrastructure. This includes fostering a culture of continuous learning and adaptation within the team, ensuring that operational processes are always evolving and improving.

A significant part of the role is dedicated to building and deploying AI-powered tooling, including autonomous coding agents, LLM-assisted CI/CD, and automated incident triage. The goal is to create sandboxed environments where agents can write, test, and verify code independently, thereby increasing the overall productivity of the engineering organization. This requires not just using AI tools, but actively developing and implementing them.

The estimated annual salary for this position ranges from $82,000 to $112,000.

Skills that matter here

  • Kubernetes: The role requires expertise in managing and maintaining Kubernetes clusters, understanding its core concepts for infrastructure design.
  • ArgoCD: This tool is used for GitOps, ensuring continuous deployment and synchronization of application states.
  • Terraform: Terraform is essential for Infrastructure as Code (IaC), provisioning and managing cloud resources.
  • Go: Proficiency in Go is required for development and scripting tasks related to infrastructure and tooling.
  • Python: Python is also a required language for scripting and development, similar to Go.
  • AI tools: The role demands the ability to build and deploy AI-powered developer tooling and autonomous agents to enhance engineering productivity.

Who this role suits

  • A person who thrives on building and automating infrastructure, always looking for ways to improve system reliability and team efficiency.
  • Someone with a proactive approach to problem-solving, capable of leading incident resolution and fostering a blameless learning environment.
  • An engineer who is passionate about leveraging AI to create innovative developer tools and increase organizational productivity.
  • A collaborative individual who enjoys working closely with product teams and providing technical guidance for platform modernization.

From the employer

Responsibilities

  • Design and implement infrastructure and tools that empower our product teams to rapidly and securely iterate, emphasizing reliability and automation.
  • Influence the strategic direction of our infrastructure and operational practices, ensuring that we are well-positioned to scale and support our growing organization.
  • Take a proactive role in the resolution of production issues, ensuring that we are well-prepared to handle incidents and that we learn from them in a blameless manner.
  • Work closely with product teams on crucial initiatives such as production deployments, release management, and incident handling, aiming for seamless operations.
  • Offer technical expertise and input to support the continual adoption and modernization of our platform and infrastructure.
  • Build and deploy AI-powered tooling (autonomous coding agents, LLM-assisted CI/CD, automated incident triage) that makes the engineering org more productive. Think: sandboxed environments where agents can write, test, and verify code without human babysitting.
  • Foster a culture of continuous learning and improvement, encouraging constructive review and adaptation processes.

Your Experience & Qualifications

  • Kubernetes expertise, with a strong understanding of its core concepts and the ability to manage and maintain clusters.
  • Expertise within modern cloud native tools, e.g. ArgoCD for GitOps, Terraform/Crossplane for IaC, and the Grafana LGTM stack (Loki, Grafana, Tempo, Mimir) for observability.
  • 3-5 years of experience in using Infrastructure as Code and tools for cloud provisioning - Must
  • 3-5 years of practice in development and scripting in languages like Go, Python, or similar - Must
  • Proficient in both written and spoken English, with exceptional communication abilities.
  • Expertise when it comes to Linux environments, containerization, and cloud technologies.
  • Comprehensive knowledge of production management concepts for distributed systems.
  • A history of 3-5 years in operational roles, overseeing production settings.
  • AI fluency. You use AI coding tools daily and have opinions about what works. More importantly, you can build and deploy LLM-powered developer tooling and autonomous agents, not just consume them. We want someone who thinks about how to make an entire engineering team more productive with AI.
  • Networking knowledge: bonus points for service mesh experience, platform engineering and cross-cloud networking.
  • Familiarity with the Ethereum ecosystem, staking, and blockchain technologies - Advantage.

Conditions

  • Salary: $82k - $112k estimated.

Questions about this role

What is the remote work policy for this role?

This is a fully remote position, allowing candidates to work from any location.

What level of seniority is expected for this position?

This is a senior-level role, requiring 3-5 years of experience in operational roles and Infrastructure as Code.

What are the key technical skills required for this role?

Key technical skills include expertise in Kubernetes, modern cloud-native tools like ArgoCD and Terraform, proficiency in Go or Python, and the ability to build and deploy AI-powered tooling.

Similar jobs

Before you apply

  • Legitimate employers never ask you to pay anything to apply or get hired.
  • Never share seed phrases or private keys. No real job needs them.
  • Do not install software ("test tasks", "trading tools", "video call clients") sent during hiring.
  • Check that the application page's domain really belongs to Bloxstaking.