Senior DevOps Engineer
Role in brief
BitFit Labs is building a blockchain custody platform and seeks a Senior DevOps Engineer. This role involves managing hybrid cloud and on-premise infrastructure, ensuring high availability and security for their MPC-TSS solution. Candidates with extensive AWS, Kubernetes, and CI/CD experience, particularly in distributed systems, should apply to help secure digital assets for institutions.
About the role
As a Senior DevOps Engineer at BitFit Labs, you will be responsible for the infrastructure supporting an enterprise-grade blockchain custody platform. This includes designing and maintaining secure AWS environments using Infrastructure as Code, and managing on-premise Trusted Execution Environment (TEE) infrastructure for secure Multi-Party Computation (MPC) key operations. The role involves deploying and maintaining critical components such as MPC TSS nodes, backend API servers, and blockchain full nodes across various chains like Bitcoin, Solana, and EVM-compatible networks.
A key aspect of this position is ensuring the reliability and performance of the platform through robust CI/CD pipelines and comprehensive monitoring. You will design and maintain automated deployment strategies using GitHub Actions or GitLab CI, implementing blue-green and canary deployments for zero-downtime releases. Additionally, you will build and maintain monitoring, logging, alerting, and tracing systems using tools like Prometheus, Grafana, and the ELK Stack, ensuring proactive incident response and adherence to strict SLA commitments.
Success in this role means maintaining a highly secure and available infrastructure for digital asset custody. This involves enforcing strict IAM, network segmentation, and zero-trust principles, integrating Hardware Security Modules (HSM) for key management, and conducting regular security assessments. You will also design and implement cross-AZ and cross-environment disaster recovery plans, performing regular drills to ensure 99.9%+ uptime for critical wallet services, thereby safeguarding institutional digital assets.
The salary for this position ranges from $150,000 to $180,000 USD annually.
Skills that matter here
- AWS: This role requires expert-level knowledge of AWS to design, build, and maintain secure and highly available cloud environments for the blockchain custody platform.
- Terraform: You will use Terraform for Infrastructure as Code to manage and automate the provisioning of AWS environments.
- Docker: Deep hands-on experience with Docker is required for containerizing applications and managing them in production environments.
- Kubernetes: You will leverage Kubernetes to orchestrate and manage containerized applications, ensuring scalability and reliability in production.
- GitHub Actions: This role involves designing and maintaining end-to-end CI/CD pipelines using GitHub Actions for automated deployments.
- Prometheus: You will build and maintain monitoring systems using Prometheus to track the performance and health of the platform.
Who this role suits
- A candidate with at least five years of experience in DevOps, SRE, or cloud infrastructure, specifically with large-scale distributed systems.
- Someone who possesses strong scripting or programming abilities in Go, Python, or Shell, and solid Linux administration skills.
- An individual who understands networking concepts like firewalls, SSL/TLS, and load balancing.
- A professional adept at creating clear technical documentation and communicating effectively across teams.
From the employer
Key Responsibilities
Infrastructure & Environment Management
- Design, build, and maintain secure, highly available AWS environments using Infrastructure as Code (Terraform / CDK)
- Set up and manage on-premise TEE infrastructure for secure MPC key computation
- Deploy and maintain all critical components: MPC TSS nodes (BTC, SOL, EVM, TRON), backend API servers, frontend web application, and blockchain full nodes
Blockchain Infrastructure
- Configure and monitor full node synchronization for Ethereum and other supported chains
- Manage RPC endpoints and load balancing across multiple nodes
- Ensure high availability of blockchain connectivity with automated failover and health checks
CI/CD & Deployment Automation
- Design and maintain end-to-end CI/CD pipelines using GitHub Actions or GitLab CI
- Implement blue-green and canary deployment strategies for zero-downtime releases
- Provide self-service build, test, and deployment tooling for the development team
Monitoring & Observability
- Build and maintain monitoring, logging, alerting, and tracing systems (Prometheus, Grafana, ELK Stack, Jaeger)
- Monitor all layers: AWS infrastructure, on-prem TEE nodes, blockchain nodes, MPC computation, and application performance
- Maintain SLA commitments with proactive alerting and incident response
Security & Compliance
- Enforce strict IAM, network segmentation, and zero-trust principles across both cloud and on-prem environments
- Integrate HSM/KMS for secure key management and MPC TSS operations
- Conduct regular security scans, vulnerability assessments, and penetration testing
- Manage SSL/TLS, Nginx reverse proxies, and firewall rules
High Availability & Disaster Recovery
- Design and implement cross-AZ and cross-environment disaster recovery plans
- Execute regular DR drills and maintain runbooks
- Maintain 99.9%+ uptime for all critical wallet services
Requirements
- 5+ years of DevOps, SRE, or cloud infrastructure engineering experience with large-scale distributed systems
- Expert-level knowledge of AWS and its managed services
- Deep hands-on experience with Docker and Kubernetes in production environments
- Proficiency with GitHub Actions or GitLab CI/CD
- Strong scripting or programming ability in at least one of: Go, Python, or Shell
- Solid Linux administration experience (Ubuntu / RedHat)
- Strong understanding of networking: firewalls, SSL/TLS, load balancing, DNS
- Hands-on experience with Prometheus, Grafana, or equivalent monitoring solutions
- Excellent technical documentation and cross-team communication skills
Nice to Have
- Direct experience operating Ethereum, Bitcoin, or other blockchain full nodes
- Familiarity with Web3 technologies, smart contracts, or DeFi protocols
- Understanding of MPC, cryptographic key management, or HSM integration
- Experience with service mesh technologies (Istio, Linkerd)
- Chaos engineering or fault injection testing experience
- Experience with HashiCorp Vault, AWS Secrets Manager, or similar secrets management tools
- Familiarity with Redis, TimescaleDB, or InfluxDB
- Security certifications (CISSP, AWS Security Specialty, or equivalent)
- Experience with multi-cloud or hybrid cloud architectures
Questions about this role
What is the remote work policy for this position?
This is a fully remote position, allowing candidates to work from any location.
What is the seniority level for this role?
This is a senior-level position, requiring significant experience in DevOps, SRE, or cloud infrastructure engineering.
What are some of the key technical skills required for this role?
Key technical skills include expert-level knowledge of AWS, deep experience with Docker and Kubernetes, proficiency with GitHub Actions or GitLab CI/CD, and strong scripting in Go, Python, or Shell.