Role in brief
Copperco, a digital asset infrastructure firm, seeks a Principal Site Reliability Engineer. This role involves defining and implementing SRE principles, automating systems for reliability, and mentoring teams. It is suited for an experienced engineer with a background in distributed systems, observability, and incident management, ideally with financial services or blockchain exposure, who can drive organizational change.
About the role
This Principal Site Reliability Engineer will be responsible for shaping how Copperco approaches reliability, observability, and operational excellence. The role involves defining key metrics like SLIs, SLOs, and error budgets, and then building the systems and processes to measure these. A core part of the work will be to champion architectural improvements that enhance both system reliability and deployment speed, ensuring that SRE principles are adopted across the organization.
The role also includes providing expert consultation on system architecture, developing reusable platforms, and planning capacity needs. The Principal SRE will conduct production readiness reviews to ensure successful service launches and operations, and will be involved in the entire lifecycle of microservices, from initial design through deployment, ongoing operation, and continuous refinement. Technical leadership is key, as the role requires partnering with engineering and product leaders to integrate reliability into the product development process.
A significant aspect of this position is leading through influence and mentorship. The Principal SRE will conduct blameless postmortems to drive systemic improvements in incident management and will mentor engineers across the organization on SRE practices. This involves empowering teams to take ownership of their service reliability, fostering a culture where every team contributes to the overall stability and performance of Copperco's digital asset infrastructure.
The annual salary for this Principal Site Reliability Engineer position is between $200,000 and $300,000.
Skills that matter here
- distributed systems: The role requires experience in designing, analyzing, and troubleshooting these systems to ensure reliability and performance.
- micro-services architectures: Expertise in these architectures is needed for improving the lifecycle from inception to operation.
- observability: This is a key area of expertise for defining reliability standards and monitoring system health.
- incident management: The role involves improving processes and conducting postmortems to enhance system resilience.
- AWS: Experience with production workloads in AWS is desirable for this position.
- financial services or similarly regulated environments: Experience in these sectors is beneficial due to the nature of Copperco's business.
Who this role suits
- A systematic problem-solver who can drive organizational change.
- An experienced engineer with a background in distributed systems and microservices.
- Someone who thrives on defining and implementing reliability standards.
- A leader capable of mentoring other engineers and influencing technical direction.
From the employer
Key Responsibilities:
- Shape SRE; Define how we think about reliability, observability, and operational excellence.
- Drive the adoption of SRE principles across the organization while building the systems and processes that make those principles measurable – think SLIs, SLOs and error budgets.
- Scale Through Automation; Champion architectural improvements that enhance both system reliability and deployment velocity.
- Provide consultation on system architecture, building reusable platforms and frameworks, planning capacity needs, and conducting production readiness reviews to ensure services launch and operate successfully.
- Drive Technical Excellence; Engage in and improve the lifecycle of microservices, from inception through deployment, operation, observability, and continuous refinement.
- Lead Through Influence; Partner with engineering and product leadership to embed reliability into our product development lifecycle.
- Conduct blameless postmortems and drive systemic improvements in incident management.
- Mentor engineers across the organisation on SRE practices, helping teams take ownership of their service reliability.
Skills and Experience:
- Essential Experience in designing, analysing, and troubleshooting distributed systems or micro-services architectures.
- Established expertise in observability and incident management.
- Proven experience in driving organizational Change.
- Excellent communication skills, with a systematic problem-solving approach.
- Desirable Experience working with production workloads in AWS.
- Experience working in financial services or similarly regulated environments.
- Interest in blockchain based technologies and/or ‘decentralised finance.
- Master’s degree in Computer Science or Engineering.
Why Copper?
- At Copper, we keep innovation, openness, and curiosity at the centre of everything we do.
- Here, bold ideas get the spotlight, learning is constant, and diversity shapes our team from the ground up.
- Jump into a fast-moving, dynamic team that loves a challenge and knows how to have fun along the way.
- Collaboration is just as important as results—you’ll be surrounded by smart, driven colleagues in London and across our APAC, Switzerland, UAE, and US offices.
- Hybrid working model – we believe in the value of bringing people together and at the same time we embrace the adaptability of flexibly working.
- Diversity and inclusion matter to us – they’re woven into Copper life.
- From employee-led groups like Women at Copper to a committee focused on community and wellbeing, you’ll have a network that supports you from day one.
- Everyone’s voice matters.
- If you’re looking to ramp up your career, or keen to do something new in your field, with us, you’ll keep moving forward.
Questions about this role
What is the remote work policy for this role?
This is a fully remote, full-time position. Copperco also mentions a hybrid working model in general, but this specific role is listed as remote.
What level of experience is expected for this position?
This is a Principal-level role, indicating a need for established expertise and proven experience in driving organizational change and technical leadership.
What is the compensation for this role?
The salary for this position ranges from $200,000 to $300,000 per year.