Senior DevOps / Infrastructure Engineer

Remote $180k–$250k senior 23 days ago full-time quality 8.6/10

Role in brief

Category Labs is building a high-performance blockchain and seeks a Senior DevOps Engineer to manage its node infrastructure. This role involves operating and enhancing the Monad node fleet, automating release pipelines, and developing AI-driven operational tools. Candidates with strong Linux, infrastructure-as-code, and observability experience, who are comfortable with AI-assisted engineering, should apply.

LinuxsystemdnetworkingshellAnsibleTerraformKubernetesGitOpsPrometheusGrafanaLokiPython

About the role

This Senior DevOps Engineer role at Category Labs focuses on operating and enhancing the infrastructure supporting Monad, a high-performance blockchain. Responsibilities include managing the Monad node fleet, encompassing health monitoring, synchronization, upgrades, and recovery across various node types on both mainnet and testnet. The position requires implementing safe, staged rollouts and responding to incidents.

A core aspect of the role is owning the infrastructure-as-code, utilizing tools like Ansible for fleet configuration, and Terraform with Atlantis for cloud and DNS management. The engineer will also be responsible for Kubernetes and Flux (GitOps) for platform services. Building and maintaining observability and alerting systems using Prometheus, Grafana, and Loki is key to proactively identifying and resolving issues.

Success in this role involves automating the release pipeline for node upgrades, canary rollouts, and snapshot/restore processes, while implementing safeguards to minimize impact. The engineer will also design and build agentic operations, developing AI agents and tools to investigate, diagnose, and execute routine operations with human oversight. Codifying operational knowledge into reusable tools and automation is also a critical part of the job.

The salary for this position ranges from $180,000 to $250,000 USD annually.

Skills that matter here

  • Linux: The role requires strong fundamental knowledge to debug live systems and manage the Monad node fleet.
  • Ansible: This tool is used for fleet configuration as part of owning the infrastructure-as-code.
  • Terraform: This is essential for managing cloud resources and DNS as part of the infrastructure-as-code responsibilities.
  • Kubernetes: Experience with Kubernetes and GitOps is a plus for managing platform services.
  • Prometheus: This tool is used for building and operating observability and alerting systems.
  • Python: Programming and scripting experience is required for automation and tool development.

Who this role suits

  • A candidate with over five years of experience operating production systems at scale in DevOps, SRE, or Infrastructure Engineering roles.
  • Someone comfortable debugging live systems via SSH and skilled in designing automation with robust guardrails.
  • An individual who uses coding agents and LLM tooling in their daily workflow and can discern where AI assistance is beneficial or risky.
  • A methodical responder to incidents, capable of bringing calm and structured approaches to problem-solving.

From the employer

What You'll Do

  • Operate the Monad node fleet: health, sync, upgrades, and recovery across validators, full nodes, archive/historical, and indexer nodes on mainnet and testnet, including safe, staged rollouts and incident response.
  • Own our infrastructure-as-code: Ansible for fleet configuration, Terraform + Atlantis for cloud and DNS, and Kubernetes/Flux (GitOps) for platform services.
  • Build and operate observability and alerting (Prometheus, Grafana, Loki); create dashboards and alerts that catch problems before they page while minimizing false positives.
  • Automate the release pipeline: node upgrades, canary rollouts, snapshot/restore, and the guardrails that bound blast radius (e.g., protecting validators from automated changes).
  • Design and build agentic operations: develop AI agents, tooling (e.g., MCP servers), and runbooks-as-code that let agents safely investigate, diagnose, and execute routine operations, with deterministic guardrails and human oversight.
  • Codify operational knowledge into tools and automation that the whole team, and its agents, can reuse.
  • Harden nodes and services, manage secrets, and continuously drive down manual toil.

Who You Are

  • You have 5+ years in DevOps, SRE, or Infrastructure Engineering, operating production systems at scale.
  • You have strong Linux, systemd, networking, and shell fundamentals, and you're comfortable debugging live systems over SSH.
  • You have deep, hands-on infrastructure-as-code experience with Ansible and Terraform.
  • You have experience with observability stacks (Prometheus, Grafana, Loki, or equivalents).
  • You have hands-on fluency with AI-assisted engineering: you use coding agents and LLM tooling in your daily workflow and have judgment on where it helps and where it's risky.
  • You have experience designing automation with safe guardrails, and you bring calm, methodical incident response.
  • You have programming and scripting experience (e.g., Python, bash).
  • Experience with Kubernetes and GitOps (Flux or Argo) is a plus.
  • Experience building AI agent tooling, MCP servers, or agent orchestration frameworks is a plus.
  • Experience serving inference, either locally or as a service is a plus.
  • Previous experience with blockchain clients or node operations is a plus.
  • A Bachelor of Science in Computer Science, Engineering, or a related field is a plus.

Why Work with Us

  • Challenging problems: You’ll work on extremely challenging problems with massive impact. See our [Blogs](https://www.category.xyz/blogs) and [Publications & Talks](https://www.category.xyz/papers-talks) for a flavor of the problems we are solving in the real world.
  • Huge opportunity: The Ethereum Virtual Machine (EVM) standard is ubiquitous, but existing EVM-compatible chains are very slow. Monad’s core innovations offer developers the best of both worlds (portability and performance) and are a game-changer for mass user adoption in crypto.
  • The right team: You’ll be part of a small, exceptional team (engineers and researchers make up 90% of the team).
  • Open by default: Our core software is public on GitHub. You’ll build in the open, and your work ships where the whole ecosystem can see it.
  • Culture: We’re a lean team working together to achieve very ambitious goals. We are united in our culture of collaboration, low ego, and high-quality output. As an early member of our team, you’ll help to shape our culture.
  • Compensation: You’ll receive a competitive salary and equity package.
  • Resources and growth: We’re well-capitalized, with [backing](https://x.com/monad_xyz/status/1777687376136982767) from leading venture funds like Paradigm, Electric Capital, Greenoaks, Dragonfly, and Coinbase Ventures. We keep a lean team, and this is a rare opportunity to join. You’ll learn a lot and grow as our company scales.

Questions about this role

What is the remote work policy for this role?

This is a fully remote position, with no specified geographic restrictions.

What level of seniority is expected for this position?

This is a senior-level role, requiring significant experience in operating production systems.

What are the key technical skills required for this role?

Key skills include strong Linux, systemd, networking, shell fundamentals, hands-on experience with Ansible and Terraform, and familiarity with observability stacks like Prometheus, Grafana, and Loki. Programming in Python or bash is also required.

Similar jobs

Before you apply

  • Legitimate employers never ask you to pay anything to apply or get hired.
  • Never share seed phrases or private keys. No real job needs them.
  • Do not install software ("test tasks", "trading tools", "video call clients") sent during hiring.
  • Check that the application page's domain really belongs to Category Labs.