Cloud Infrastructure Engineer

Remote $135k–$240k senior 4 months ago full-time quality 9.2/10

Role in brief

Alchemy is seeking a Senior Cloud Infrastructure Engineer to build and maintain scalable, multi-region cloud infrastructure for their web3 developer platform. This role involves architecting solutions with Kubernetes and Terraform, driving AI enablement in infrastructure, and developing internal developer tools. Candidates with strong reliability engineering experience and expertise in cloud-native technologies should apply.

KubernetesTerraformAWSGCPPrometheusGrafanaOpenTelemetryHelmGitOpsArgoCDIstioClaude Code

About the role

This role focuses on designing and operating highly scalable, self-healing infrastructure. The work involves leveraging Kubernetes, Terraform, and other cloud-native tools across multi-region deployments. A key aspect of the role is driving AI enablement within engineering by optimizing workflows and tooling for agentic development, using tools like Claude Code, Cursor, and Codex. This includes building AI-powered automation for tasks such as Kubernetes upgrades and cost optimization.

The engineer will also be responsible for building and maintaining internal developer platform capabilities, ensuring self-service deployments, robust observability, and overall reliability. This involves developing observability frameworks with Prometheus and Grafana for metrics, dashboards, and alerting. Additionally, the role includes leading incident management, defining and enforcing service level indicators and objectives, and managing error budgets across services.

Success in this position means architecting and managing complex multi-cloud, multi-region network architectures, including VPC design, DNS, and cross-cloud connectivity. Collaboration with security teams is crucial to embed compliance into infrastructure practices. The role also requires providing technical leadership and mentorship to enhance the team's operational capabilities, ensuring a focus on systemic improvement and effective incident response.

The salary for this role ranges from $135,000 to $240,000 annually, in addition to equity.

Skills that matter here

  • Kubernetes: This role requires architecting and operating scalable, self-healing infrastructure leveraging Kubernetes across multi-region deployments.
  • Terraform: The engineer will use Terraform for infrastructure as code, with an automation-first mindset for managing cloud resources.
  • AWS: Deep experience with cloud infrastructure, including AWS, is necessary for designing and managing multi-cloud environments.
  • GCP: Deep experience with cloud infrastructure, including GCP, is necessary for designing and managing multi-cloud environments.
  • Prometheus: The role involves developing observability frameworks using Prometheus for metrics, dashboards, and alerting.
  • ArgoCD: Proficiency with GitOps workflows, such as ArgoCD, is required for managing deployments and infrastructure changes.

Who this role suits

  • A person with at least five years of experience in reliability-focused infrastructure roles, such as SRE or Platform Engineer.
  • Someone who is calm and effective during incidents, focusing on systemic improvements rather than blame.
  • An individual who can communicate clearly and effectively across SRE, security, and product engineering teams.
  • A professional who has driven company-wide reliability efforts, including the implementation of SLO frameworks and error budget policies.

From the employer

  • Architect and operate scalable, self-healing infrastructure leveraging Kubernetes, Terraform, and cloud-native tools across multi-region deployments.
  • Drive AI enablement across engineering — ensuring repos, tooling, and workflows are optimized for agentic development with tools like Claude Code, Cursor, and Codex.
  • Build AI-powered infrastructure tooling and automation (e.g., automated K8s upgrades, IaC plan analysis, cost optimization advisors, MCP servers, n8n workflows).
  • Build and maintain internal developer platform (IDP) capabilities for self-service deployments, observability, and reliability.
  • Develop observability frameworks using Prometheus and Grafana for metrics, dashboards, and alerting.
  • Lead incident management with blameless post-mortems; define and enforce SLIs, SLOs, and error budgets across services.
  • Design and manage multi-cloud, multi-region network architecture — VPC design, IPAM, DNS (Cloudflare), cross-cloud connectivity, security groups, and edge-proxy/istio gateway configuration.
  • Collaborate with security teams to embed compliance into infrastructure, including IaC scanning and runtime protection.
  • Provide technical leadership and mentorship to elevate the team’s operational capabilities.
  • 5+ years as an Infrastructure Engineer focused on reliability (SRE, Production Engineer, Platform Engineer).
  • Experience driving company-wide reliability efforts, including SLO frameworks and error budget policies.
  • Strong proficiency with observability stacks: OpenTelemetry, Prometheus/Grafana.
  • Deep experience with cloud infrastructure (AWS/GCP), Kubernetes, and multi-region architectures.
  • Skilled with Terraform, Helm, and GitOps workflows (e.g., ArgoCD) with an automation-first mindset.
  • Experience leveraging agentic development tools (Claude Code, Cursor, Codex) and workflow automation (n8n) to accelerate IaC and build internal tooling is a strong plus.
  • Solid networking fundamentals — VPC design, DNS, IPAM, security groups, cross-cloud connectivity, and service mesh (e.g., Istio) experience is a plus.
  • Calm and effective incident responder with a focus on systemic improvement.
  • Strong cross-functional communicator across SRE, security, and product engineering.
  • Blockchain infrastructure, distributed systems, or high-throughput RPC experience — not required but a plus.
  • Medical, Dental, & Vision
  • Gym Reimbursement
  • Home Office Build-out Budget
  • In-Office Group Meals
  • Wellbeing & Mental Health Perks
  • Learning & Development Stipend
  • Company Sponsored Conferences & Events
  • HSA and FSA Plans
  • Fertility Benefits
  • Competitive compensation, including base salary as well as equity
  • Comprehensive medical, dental, and vision coverage
  • 401k and unlimited flexible time off

Questions about this role

What is the remote work policy for this position?

This is a fully remote position.

What level of seniority is expected for this role?

This is a senior-level position.

What are the key technical skills required for this role?

Key technical skills include strong proficiency with Kubernetes, Terraform, AWS/GCP, Prometheus/Grafana, and GitOps workflows like ArgoCD.

Similar jobs

Before you apply

  • Legitimate employers never ask you to pay anything to apply or get hired.
  • Never share seed phrases or private keys. No real job needs them.
  • Do not install software ("test tasks", "trading tools", "video call clients") sent during hiring.
  • Check that the application page's domain really belongs to Alchemy.