Role in brief
Alpaca, a financial services company offering brokerage infrastructure, seeks a DevOps Team Lead for its Core Foundation pod. This role involves managing a globally distributed team responsible for critical infrastructure, including networking, data, and cloud. Ideal for an experienced engineering manager with a strong technical background in DevOps/SRE, who can lead and mentor teams while driving large-scale infrastructure upgrades and ensuring platform resilience.
About the role
As the DevOps Team Lead for the Core Foundation pod, you will oversee the critical infrastructure initiatives that power Alpaca's financial services platform. This includes managing heavy compute, core networking, stateful data, observability, and both cloud and physical infrastructure. Your primary focus will be on leading a globally distributed team of engineers, guiding them through complex projects, and ensuring the stability and performance of the platform.
A key aspect of this role is leadership, encompassing both people and technology. You will be responsible for mentoring your team, fostering a positive team culture across different time zones, and owning the execution of critical roadmap items. This involves strategic planning for large-scale infrastructure upgrades, ensuring high availability, and enhancing platform resilience, all within a dynamic and diverse environment.
Success in this position means effectively managing change, implementing robust support frameworks, and refining agile planning methodologies. You will also manage and optimize the global on-call rotation, ensuring team well-being while maintaining high availability for Alpaca's services. The role requires a strategic mindset to navigate shifting priorities and a commitment to operational excellence.
The salary for this position ranges from $82,000 to $138,000 annually, in addition to stock options.
Skills that matter here
- Kubernetes: This role requires a solid technical background in Kubernetes, specifically GKE, as part of modern DevOps/SRE ecosystems.
- Terraform: Expertise in Terraform is essential for managing infrastructure as code within the team's responsibilities.
- PostgreSQL: Experience with relational databases like PostgreSQL is necessary for managing stateful data infrastructure.
- Prometheus: Knowledge of Prometheus is required as part of the observability stack for monitoring critical infrastructure.
- Grafana: Familiarity with Grafana is expected for visualizing data from the observability stack.
- Thanos: Experience with Thanos is needed to manage and scale the observability infrastructure.
Who this role suits
- You have a proven background as an Engineering Manager, DevOps Lead, or Site Reliability Engineering Lead.
- You excel at coaching, mentoring, and fostering team culture across multiple time zones.
- You possess deep expertise in engineering support frameworks, roadmap planning, and team prioritization methodologies.
- You are adept at owning Change Management and Incident Management lifecycles, including running sustainable global on-call rotations.
From the employer
Your Role:
As the DevOps Team Lead for our Core Foundation pod, you will lead the "engine room" of Alpaca’s most critical infrastructure initiatives. You will manage a talented, globally distributed team of engineers responsible for Heavy Compute, Core Networking, Stateful Data, Observability, and Cloud/Physical Infrastructure.
Things You Get To Do:
- People & Tech Leadership: Lead, mentor, and foster a healthy, high-performing globally distributed engineering team.
- Prioritization & Planning: Own the execution and delivery of highly critical, complex yearly roadmap items centered around large-scale foundational infrastructure upgrades, high availability, and platform resilience.
- Change Management Ownership: Own and drive the change management processes across engineering and product domains.
- Support Frameworks & Methodologies: Design, implement, and refine robust support workflows, agile planning methodologies, and deployment/rollout strategies to ensure operational excellence.
- On-Call & Incident Management: Manage and optimize the global on-call rotation to ensure team well-being while maintaining high availability.
Who you are (must-haves):
- Proven experience as an Engineering Manager, DevOps Lead, or Site Reliability Engineering Lead, with a track record of successfully managing globally distributed teams.
- Exceptional people management skills, with a deep focus on coaching, mentoring, and fostering team culture across multiple time zones.
- Deep expertise in engineering support frameworks, roadmap planning, and team prioritization methodologies.
- Proven experience owning Change Management lifecycles.
- Extensive experience managing Incident Management lifecycles and running sustainable, global on-call rotations.
- Incredibly strong communication and organizational skills.
- A solid technical background in modern DevOps/SRE ecosystems, including Kubernetes (GKE), Infrastructure as Code (Terraform), Relational Databases (PostgreSQL), and Observability stacks (Prometheus, Grafana, Thanos).
- A strategic mindset capable of navigating shifting priorities.
How We Take Care of You:
- Competitive Salary & Stock Options
- Health Benefits
- New Hire Home-Office Setup: One-time USD $500
- Monthly Stipend: USD $150 per month via a Brex Card.
Questions about this role
What is the remote work policy for this role?
This is a fully remote position.
What is the seniority level for this position?
This is a lead-level role, requiring proven experience in managing engineering teams.
What technical skills are important for this role?
Key technical skills include Kubernetes (GKE), Terraform, PostgreSQL, Prometheus, Grafana, and Thanos.