Role in brief
Alchemy seeks a senior Cloud Infrastructure Engineer to design and operate scalable, multi-region blockchain infrastructure. This role involves using Kubernetes, Terraform, and cloud-native tools, with a focus on AI enablement and internal developer platforms. Candidates with strong reliability engineering experience, proficiency in observability stacks, and expertise in AWS/GCP and Kubernetes should apply.
About the role
This role focuses on architecting and operating scalable, self-healing infrastructure for blockchain applications, utilizing Kubernetes, Terraform, and other cloud-native tools across multiple regions. A key aspect involves driving AI enablement within engineering, optimizing repositories, tooling, and workflows for agentic development, and building AI-powered infrastructure tools and automation. The position contributes to a platform that supports a significant portion of web3 teams and major global brands.
The engineer will build and maintain internal developer platform capabilities to support self-service deployments, observability, and system reliability. This includes developing observability frameworks using Prometheus and Grafana, and designing multi-cloud, multi-region network architectures. Collaboration with security teams is essential to integrate compliance into infrastructure design.
Success in this position involves providing technical leadership, mentoring team members, and effectively managing incidents with blameless post-mortems. The ideal candidate will have a background in driving company-wide reliability efforts and possess strong communication skills to work cross-functionally within a team that draws expertise from leading technology companies and universities.
The annual salary for this role ranges from $135,000 to $240,000.
Skills that matter here
- Kubernetes: This role requires designing and operating scalable infrastructure using Kubernetes across multi-region deployments.
- Terraform: The engineer will use Terraform to provision and manage infrastructure as code.
- AWS: Experience with cloud infrastructure, specifically AWS, is essential for managing multi-cloud environments.
- GCP: Proficiency in GCP is required for managing multi-cloud infrastructure and deployments.
- Prometheus: The role involves developing and utilizing observability frameworks with Prometheus.
- Grafana: This position requires building and maintaining observability dashboards and alerts using Grafana.
Who this role suits
- A person who has spent at least five years focused on infrastructure reliability.
- Someone who can calmly and effectively respond to incidents and lead post-mortems.
- An individual who can communicate complex technical information clearly across different teams.
- A candidate with experience driving company-wide initiatives related to system reliability.
From the employer
- Architect and operate scalable, self-healing infrastructure leveraging Kubernetes, Terraform, and cloud-native tools across multi-region deployments.
- Drive AI enablement across engineering — ensuring repos, tooling, and workflows are optimized for agentic development.
- Build AI-powered infrastructure tooling and automation.
- Build and maintain internal developer platform (IDP) capabilities for self-service deployments, observability, and reliability.
- Develop observability frameworks using Prometheus and Grafana.
- Lead incident management with blameless post-mortems.
- Design and manage multi-cloud, multi-region network architecture.
- Collaborate with security teams to embed compliance into infrastructure.
- Provide technical leadership and mentorship.
- 5+ years as an Infrastructure Engineer focused on reliability.
- Experience driving company-wide reliability efforts.
- Strong proficiency with observability stacks: OpenTelemetry, Prometheus/Grafana.
- Deep experience with cloud infrastructure (AWS/GCP), Kubernetes, and multi-region architectures.
- Skilled with Terraform, Helm, and GitOps workflows.
- Experience with agentic development tools and workflow automation.
- Solid networking fundamentals.
- Calm and effective incident responder.
- Strong cross-functional communicator.
- Blockchain infrastructure experience is a plus.
- Medical, Dental, & Vision
- Gym Reimbursement
- Home Office Build-out Budget
- In-Office Group Meals
- Wellbeing & Mental Health Perks
- Learning & Development Stipend
- Company Sponsored Conferences & Events
- HSA and FSA Plans
- Fertility Benefits
- Competitive compensation including base salary and equity
- Comprehensive medical, dental, and vision coverage
- 401k and unlimited flexible time off
Questions about this role
What is the remote work policy for this role?
This is a fully remote position.
What level of seniority is expected for this position?
This is a senior-level role.
What are the key technical skills required?
Key technical skills include Kubernetes, Terraform, AWS, GCP, Prometheus, Grafana, OpenTelemetry, Helm, and GitOps workflows.