Vice President Site Reliability Engineering (Data Centers)
Role in brief
Galaxydigitalservices, a leader in digital assets and data center infrastructure, seeks a Vice President of Site Reliability Engineering. This remote role involves leading a specialized SRE team to develop and maintain automation toolsets, enforce Infrastructure as Code standards, and manage monitoring for automation platforms. This position is ideal for an experienced SRE or DevOps professional with a strong background in infrastructure automation and team leadership.
About the role
This Vice President role focuses on leading a Site Reliability Engineering team to drive automation and infrastructure initiatives. The core work involves designing, deploying, and maintaining automation toolsets, establishing Infrastructure as Code governance, and optimizing configuration management with tools like Ansible and Packer. The role also requires managing the monitoring and health of these automation platforms, implementing SLIs/SLOs to ensure high availability.
A key aspect of this position is leading the automated lifecycle management of physical and virtual assets, from initial deployment to patching and scaling. This includes developing custom scripts and internal providers using languages like Python, Go, PowerShell, and Bash to enhance system insights and tooling. Collaboration with the broader Datacenter team is essential to facilitate overall team needs and workflows.
Success in this role means fostering an "automate-first" culture, providing technical guidance and mentorship to SREs, and continuously improving processes. The Vice President will analyze system behavior and resource utilization in virtual environments to optimize automated deployments, ensuring performance and efficiency across the infrastructure ecosystem.
The salary for this position is listed between $120,000 and $200,000 USD.
Skills that matter here
- Terraform: This role requires strong proficiency in Terraform for establishing and enforcing Infrastructure as Code standards across the infrastructure ecosystem.
- Ansible: The Vice President will lead the strategy for automated configuration and state management, optimizing Ansible playbooks for Windows, Linux, and ESXi platforms.
- Packer: Experience with Packer is needed for building standardized, hardened images for both Windows and Linux in hybrid environments.
- Python: High-level scripting skills in Python are necessary for developing custom tooling and internal providers to gain better insights into systems.
- Site Reliability Engineering: This leadership role is centered on Site Reliability Engineering principles, focusing on automation, infrastructure stability, and continuous improvement.
- Observability Tools: The role involves managing monitoring and health of automation platforms, requiring experience with observability tools to implement SLIs/SLOs.
Who this role suits
- A leader with 6-10 years of experience in infrastructure, SRE, or DevOps, specifically focused on large-scale infrastructure automation.
- An individual who can establish and enforce technical standards, particularly for Infrastructure as Code and configuration management.
- Someone adept at mentoring SREs and fostering a culture of automation and continuous improvement.
- A hands-on professional comfortable with both Windows Server and Linux environments, including troubleshooting and performance tuning.
From the employer
Responsibilities
- Automation Platform Leadership: Oversee a specialized SRE team focused on the design, deployment, and maintenance of automation toolsets as well as the systems they interact with.
- Infrastructure as Code (IaC) Governance: Establish and enforce standards for IaC to ensure consistent, repeatable, and secure deployments across an entire infrastructure ecosystem. Strong proficiency in Terraform is required.
- Configuration Management: Lead the strategy for automated configuration and state management, ensuring Ansible playbooks and Packer image pipelines are optimized for both Windows, Linux, and ESXi Platforms.
- Monitoring & Observability: Manage the monitoring and health of the automation platforms themselves. Implement SLIs/SLOs to ensure the "tools that build the servers" are highly available and performant.
- Lifecycle Management: Drive the automated lifecycle of both physical and virtual assets, from initial template creation/deployment to automated patching, scaling, and decommissioning.
- Custom Tooling & Scripting: Lead the development of custom scripts and internal providers (Python, Go, PowerShell, Bash) to provide better insights and tooling for our systems.
- Collaboration: Outside of the automation team you will need to be able to collaborate and foster workflows alongside the rest of the Datacenter team and be able to facilitate needs for the team as a whole.
- Capacity & Performance: Analyze system behavior and resource utilization in virtual environments to optimize the performance of automated deployments.
- Mentorship & Growth: Provide technical guidance and career mentorship to SREs, fostering a culture of "automate-first" and continuous improvement.
Requirements
- 6-10 years’ experience in Infrastructure, SRE or DevOps, specifically focused on infrastructure automation at scale.
- Deep proficiency with Terraform (providers, modules, state management) and Ansible (roles, playbooks, Tower/AWX).
- Hands-on experience with Image Creation (i.e. Packer, Ansible, SCCM) to build standardized, hardened images for both Windows and Linux in hybrid environments.
- Strong experience managing and automating virtual platforms such as VMware (vSphere/vCenter) as well as Cloud providers such as Azure and AWS.
- High-level scripting skills in mediums such as Python, Go, PowerShell, and Bash.
- Experience with observability tools (Splunk, ELK, Prometheus, or Grafana) to monitor infrastructure health and automation telemetry.
- Good understanding of Network topology and design as well as experience with platforms such as Juniper Networks or Palo Alto.
- Strong mastery of Git (branching strategies, PR workflows) and CI/CD platforms (Jenkins, GitLab CI, or GitHub Actions).
- Equal comfort managing, troubleshooting, and tuning performance for both Windows Server and Linux.
Conditions
- Competitive salary range of $120K-$200K.
- Remote work environment.
- Opportunities for career mentorship and growth.
Questions about this role
What is the remote work policy for this position?
This is a fully remote position.
What is the salary range for this role?
The competitive salary for this position ranges from $120,000 to $200,000 USD.
What level of experience is required for this role?
Candidates should have 6-10 years of experience in Infrastructure, SRE, or DevOps, with a focus on infrastructure automation at scale.