Senior IaaS / Kubernetes Platform Engineer
Role in brief
CloudLinux is looking for a Senior IaaS / Kubernetes Platform Engineer to design and operate a multi-tenant Kubernetes platform and manage infrastructure across multiple data centers. This role is ideal for an experienced engineer with a strong background in Kubernetes, IaaS, and distributed storage systems like Ceph, who thrives in a remote, collaborative environment.
About the role
This role focuses on Kubernetes platform engineering, specifically designing and operating a multi-tenant platform using Cluster API and bare-metal providers. A significant part of the work involves implementing hard multi-tenancy with tools like vCluster and orchestrating virtual machines within Kubernetes using KubeVirt, including CPU pinning and NUMA awareness. The position also entails implementing GitOps-driven infrastructure with ArgoCD or Flux and managing policy-as-code using Kyverno or OPA Gatekeeper.
The position also involves substantial storage engineering, including optimizing and operating Ceph distributed storage clusters and managing Rook-Ceph operator deployments. The engineer will implement various storage tiers, design I/O isolation for VMs on shared Ceph clusters, and manage Containerized Data Importer for VM image lifecycles. Additionally, networking responsibilities include deploying overlay networks, implementing Cluster Mesh for multi-datacenter connectivity, and configuring Multus CNI and SR-IOV.
A portion of the role is dedicated to reliability and operations, practicing SRE disciplines such as defining SLOs, implementing capacity management, and conducting chaos engineering experiments. The engineer will participate in an on-call rotation for IaaS infrastructure and proactively identify and solve reliability risks and performance bottlenecks. Infrastructure as Code and automation are also key, involving the development of Terraform/OpenTofu modules and Ansible playbooks for server configuration and fleet management.
The competitive salary for this position ranges from $115,000 to $195,500 USD.
Skills that matter here
- Kubernetes: This role is primarily focused on designing, building, and operating a multi-tenant Kubernetes platform, including VM orchestration with KubeVirt.
- IaaS: The position requires proven experience in Infrastructure as a Service platform engineering and managing infrastructure across multiple data centers.
- Ceph: The engineer will operate and optimize Ceph distributed storage clusters and manage Rook-Ceph operator deployments at scale.
- Terraform: The role involves developing and maintaining Terraform/OpenTofu modules for multi-cloud infrastructure provisioning.
- Ansible: Proficiency in Ansible is required for writing playbooks for bare-metal server configuration and fleet management.
- GitOps: The engineer will implement GitOps-driven infrastructure using tools like ArgoCD or Flux for cluster configurations.
Who this role suits
- A person with a strong background in Kubernetes platform engineering and cloud infrastructure.
- Someone who can work independently while also contributing effectively to a remote team.
- An individual who is proactive in identifying and solving reliability and performance issues.
- A candidate experienced in SRE practices, including defining SLOs and participating in on-call rotations.
From the employer
What You Will Do
- Kubernetes Platform Engineering (Primary Focus — 40%)
- Design, build, and operate a multi-tenant Kubernetes platform using Cluster API (CAPI) with bare-metal providers (Metal3/Sidero).
- Implement hard multi-tenancy using vCluster (Loft Labs) or similar technology, providing isolated Kubernetes API servers per tenant.
- Deploy and manage KubeVirt for VM orchestration within Kubernetes, including CPU pinning, NUMA awareness, and HugePages configuration.
- Implement GitOps-driven infrastructure using ArgoCD or Flux as the single source of truth for all cluster configurations.
- Deploy and manage Policy-as-Code using Kyverno or OPA Gatekeeper for admission control, resource quotas, and security policies.
- Build self-service capabilities using Crossplane or similar Kubernetes-native infrastructure provisioning tools.
- Storage Engineering (20%)
- Operate and optimize Ceph distributed storage clusters (currently 1 PiB raw, 149 OSDs, Quincy 17.2.5).
- Manage Rook-Ceph operator deployments at scale on modern Kubernetes (v1.28+).
- Implement storage tiering: Ceph for bulk storage, local NVMe for high-IOPS workloads, LINSTOR/DRBD or TopoLVM for ultra-fast replicated storage.
- Design and implement per-VM / per-tenant I/O isolation on shared Ceph clusters.
- Manage CDI (Containerized Data Importer) for VM image lifecycle in KubeVirt environments.
- Networking (15%)
- Deploy and manage overlay networks for pod networking, micro-segmentation, and WireGuard/IPsec encryption.
- Implement Cluster Mesh for multi-datacenter pod-to-pod connectivity.
- Configure Multus CNI and SR-IOV for multi-NIC VM support in KubeVirt.
- Work with physical network infrastructure: Juniper switches (JunOS), BGP (eBGP/iBGP), EVPN/VXLAN, VLANs.
- Maintain IPSec site-to-site connectivity between datacenters.
- Reliability and Operations (15%)
- Practice SRE discipline: define and maintain SLOs with error budgets, implement proactive capacity management with 6-12 month forecasting.
- Design and execute chaos engineering experiments to validate system resilience.
- Participate in on-call rotation for IaaS infrastructure (OpenNebula, Ceph, networking).
- Write and maintain runbooks, DRP documentation, and postmortem analyses.
- Drive proactive improvement: identify reliability risks, performance bottlenecks, and toil — then propose and implement solutions without waiting for incidents.
- Infrastructure as Code and Automation (10%)
- Develop and maintain Terraform/OpenTofu modules for multi-cloud infrastructure provisioning.
- Write Ansible playbooks for bare-metal server configuration and fleet management.
- Automate infrastructure lifecycle: PXE
Requirements
- Proven experience in Kubernetes platform engineering and IaaS.
- Strong understanding of cloud infrastructure, networking, and storage solutions.
- Experience with GitOps practices and tools (ArgoCD, Flux).
- Familiarity with Ceph and distributed storage management.
- Proficiency in Terraform and Ansible for automation and infrastructure as code.
- Ability to work independently and collaboratively in a remote team environment.
What We Offer
- Competitive salary ranging from $115,000 to $195,500 USD.
- Fully remote work environment.
- Supportive team culture focused on collaboration and success.
- Opportunities for professional growth and development.
Questions about this role
What is the remote work policy for this role?
This is a fully remote position with a global remote-first company.
What is the seniority level for this position?
This is a senior-level position.
What are the primary technologies used in this role?
The primary technologies include Kubernetes, IaaS, Ceph, Terraform, Ansible, GitOps, and various networking solutions.