Role in brief
Gauntlet is building financial systems for onchain finance, managing over $1.5B in client assets. This Infrastructure Engineer role focuses on building and maintaining the core cloud infrastructure, CI/CD pipelines, and observability systems. It suits an engineer with strong cloud and Kubernetes experience who values security, automation, and operational excellence.
About the role
This role involves supporting application teams by handling infrastructure requests, owning CI/CD processes, and managing deployments. A key focus will be on automating deployments to ensure security and efficiency, moving towards a system where direct engineer access to production is minimized. The engineer will also be responsible for building and maintaining infrastructure as code using Terraform across GCP environments.
The position requires running Kubernetes effectively, managing service deployments with Helm, and ensuring async workloads are healthy on Dagster. A significant initial project will be unifying observability across the platform, consolidating alerting, and routing incident notifications to the correct teams. The role also contributes to advancing system resilience, aiming for cloud-agnostic services that can quickly recover from failures.
Strengthening security is a core responsibility, involving the application of IAM, secrets management, and least privilege principles, while also contributing to SOC 2 readiness. The role also explores automating routine tasks using AI agents to free up engineering time, and leveraging AI for more complex problem-solving within the infrastructure domain.
The salary range for this position is between $82,000 and $138,000 USD.
Skills that matter here
- Python: Valued for software engineering fundamentals, scripting, and shell work.
- GCP: Required for hands-on experience with cloud infrastructure and core cloud services.
- Kubernetes: Essential for operating large-scale production systems and managing service deployments.
- Terraform: Used for authoring and updating infrastructure as code modules across GCP.
- GitHub Actions: Involves maintaining and extending CI/CD workflows and migrating to dedicated CD tools.
- IAM: Applied for strengthening security, access control, and least privilege principles.
Who this role suits
- An engineer with a strong background in software engineering fundamentals and cloud infrastructure.
- Someone who prioritizes security, automation, and operational excellence in system design.
- A clear communicator who can articulate incident details, design decisions, and operational procedures.
- A problem-solver adept at debugging production issues using a variety of tools and source code.
From the employer
What you'll do;
- Support the application teams: turn around infra requests (permissions, roles, service setup, project peering) so product engineers stay focused on shipping.
- Own CI/CD and deployments: maintain and extend our GitHub Actions workflows and help migrate toward a dedicated CD tool with proper permissioning — the goal is fully automated, locked-down deploys via service accounts, no direct engineer access to production.
- Build and maintain infrastructure as code: author and update Terraform modules for new and existing services across GCP environments.
- Run Kubernetes the right way: manage service deployments via Helm (we're on Helm 4) keep async workloads healthy on Dagster.
- Unify observability (likely first project): consolidate today's per-team alerting into a single view — system-to-system dashboards plus incident alerting that routes upstream service/vendor failures to the right impacted teams and on-call rotations.
- Advance resilience: help move us toward a fully region- and cloud-agnostic posture so services can pick up and move if something fails.
- Strengthen security & access: apply IAM, secrets management, least privilege, and auditability; contribute to SOC 2 readiness.
- Automate with AI: build agent skills / agents.md so routine tasks (provisioning access, simple changes) can be handled by an agent instead of human engineering hours, and use AI to reason through bigger problems.
What you bring;
- Strong software-engineering fundamentals in at least one production language (Python, Go, TypeScript, or Rust); Python especially valued, plus comfort scripting and working in the shell.
- Hands-on experience with cloud infrastructure and core cloud services, especially GCP (AWS/Azure transferable).
- Experience operating large-scale Kubernetes production systems.
- Experience with Infrastructure as Code, especially Terraform.
- Familiarity with CI/CD systems, especially GitHub Actions or Octopus Deploy.
- Ability to debug production issues using logs, metrics, traces, shell tools, and source code.
- Security and access-control fundamentals: IAM, secrets management, least privilege, and auditability.
- Clear written communication around incidents, design decisions, and operational procedures.
Bonus points
- Supporting SOC 2 controls - evidence collection, access reviews, change management, or audit readiness.
- Observability with Datadog, Prometheus, Grafana, OpenTelemetry, Honeycomb, or similar.
- Improving developer experience through internal tooling, templates, scripts, or platform APIs.
- Incident response experience, including postmortems and follow-up remediation.
- Experience with Dagster, Helm 3+, high-scale CD tooling (Bazel, Octopus), or AI/agent-assisted ops.
- Basic web3 / DeFi literacy (transactions, wallets) and genuine curiosity about onchain — the role doesn't touch chain directly, but the business is onchain.
Questions about this role
What is the remote work policy for this role?
This is a fully remote position.
What technical skills are most important for this role?
Key skills include strong software engineering in Python, Go, TypeScript, or Rust, hands-on experience with GCP and Kubernetes, and proficiency with Terraform and CI/CD systems like GitHub Actions.
How do I apply for this position?
The job posting does not specify an application method, but typically applications are submitted via the company's website.