Role in brief
P2P.org, a major institutional staking provider, is seeking a Senior SRE to lead the launch and ensure the stability of new blockchain networks. This remote role involves designing deployment architectures, implementing production readiness standards, and collaborating with protocol teams. It's ideal for experienced SREs with strong Kubernetes, Terraform, and cloud skills, who are adept at building and operating high-availability infrastructure.
About the role
This Senior SRE role at P2P.org focuses on the full lifecycle of new blockchain network launches, from testnet to mainnet. The work involves designing and implementing deployment architectures for critical components like validators, full nodes, and RPCs. A key responsibility is ensuring these new networks meet stringent production readiness standards, covering monitoring, alerting, backups, failover, and security, to maintain high reliability and performance.
The role requires close collaboration with various protocol teams to understand specific network requirements, potential risks, and failure modes. A significant part of the job is to create repeatable launch patterns and runbooks, which will help reduce the time it takes to bring new networks to market. This involves building and operating infrastructure across both cloud and bare-metal environments, with an emphasis on automation and standardization using tools like Terraform and Helm.
Success in this position means implementing highly available and fault-tolerant setups for validator infrastructure, ensuring all services are fully observable with actionable, low-noise alerts. The SRE will also participate in on-call rotations and incident response, applying security best practices to all deployments, including secrets management, access control, and network isolation. The goal is to continuously improve infrastructure resilience and operational efficiency for P2P.org's expanding product line.
The compensation for this full-time contractor role ranges from $138,000 to $220,000 annually, with the option to be paid in crypto.
Skills that matter here
- Kubernetes: This role requires hands-on experience with Kubernetes for deploying and managing containerized applications in production environments.
- Terraform: Terraform is used to automate and standardize infrastructure deployments across various environments.
- Linux: Strong experience with Linux systems is necessary for operating and troubleshooting the underlying infrastructure.
- GCP: Experience with at least one major cloud provider like GCP is required for building and operating cloud-based infrastructure.
- Prometheus: Prometheus is part of the observability tooling used to monitor services and define actionable alerts.
- Go: Solid scripting or programming skills in languages like Go are needed for automation and tooling development.
Who this role suits
- Someone with a proven track record of operating production systems at scale and ensuring their stability.
- An individual who excels at designing and implementing robust, high-availability, and fault-tolerant infrastructure.
- A person who is proactive in creating repeatable processes and runbooks to streamline complex operations.
- A problem-solver who can debug and resolve issues under pressure, with strong communication skills for cross-team collaboration.
From the employer
- Lead the end-to-end launch of new blockchain networks—from testnet to mainnet
- Design and implement deployment architectures for validators, full nodes, RPCs, and supporting services
- Ensure all new networks meet production readiness standards—monitoring, alerting, backups, failover, and security
- Collaborate with protocol teams to understand network-specific requirements, risks, and failure modes
- Create repeatable launch patterns and runbooks to reduce time-to-market for new networks
- Build and operate infrastructure across cloud and bare-metal environments
- Improve automation and standardisation of deployments using Terraform, Helm, and internal tooling
- Implement high-availability and fault-tolerant setups for validator infrastructure
- Ensure all services are fully observable—metrics, logs, and traces
- Define and implement alerts that are actionable and low-noise
- Participate in on-call rotations and incident response
- Apply security best practices to all deployments—secrets management, access control, and network isolation
- 5+ years of experience in SRE, DevOps, or infrastructure engineering
- Strong experience operating production systems at scale
- Hands-on experience with Kubernetes, Terraform, Linux systems, and networking fundamentals
- Experience with at least one cloud provider (GCP preferred, AWS, Azure, OCI)
- Experience with observability tooling (Prometheus, Grafana, Loki, or similar)
- Familiarity with CI/CD systems and GitOps workflows (e.g., ArgoCD)
- Solid scripting or programming skills (Go, Python, or similar)
- Experience working in distributed systems or high-availability environments
- Strong debugging and problem-solving skills under pressure
- Good communication skills and ability to work across teams (English B2 minimum)
- Nice to have: Experience with blockchain infrastructure, bare-metal environments, distributed tracing, or advanced observability setups
- Fully remote
- Full-time contractor (Indefinite-term Consultancy Agreement)
- Competitive salary level in $ (we can also pay in crypto)
- Paid vacation and sick leave
- Well-being program
- Mental Health care program
- Compensation for education, including foreign language & professional growth courses
- Equipment & co-working reimbursement program
- Overseas conferences, community immersion
Questions about this role
What is the remote work policy for this role?
This is a fully remote position, allowing candidates to work from any location.
What is the seniority level for this position?
This is a senior-level role, requiring at least 5 years of experience in SRE, DevOps, or infrastructure engineering.
What skills are essential for this role?
Essential skills include Kubernetes, Terraform, Linux systems, networking fundamentals, experience with a major cloud provider (GCP preferred), observability tooling, CI/CD, GitOps, and scripting in Go or Python.