Role in brief
Block is seeking a Senior Site Reliability Engineer to improve platform stability and incident response. This role involves building reliability tools, leading incident management, and leveraging AI for system improvements. Ideal for experienced engineers focused on distributed systems and proactive reliability.
About the role
This Senior Site Reliability Engineer role at Block focuses on enhancing the reliability of their platform and critical infrastructure. The position involves both proactive development of reliability tools and reactive incident management. A core part of the work includes standardizing reliability practices across various platforms and organizations within Block, aiming for fewer customer-facing incidents and faster resolution times.
The successful candidate will be responsible for leading incident response for critical issues, including on-call duties and structured escalation. A key aspect of this role is the integration of AI-driven systems to improve signal detection, reduce operational noise, and accelerate root cause analysis. This shift aims to move Block from reactive incident handling to a more preventative, system-wide approach to reliability.
Success in this role means driving platform-wide reliability improvements, developing shared operational tooling, and implementing safe deployment patterns like progressive delivery and automated rollbacks. The team's goal is to increase product velocity and reduce burnout by building scalable and resilient distributed platforms that support safe product development across the company.
The listed salary of $170,100 represents the base pay for Zone C locations in the U.S., with actual compensation varying based on location, skills, and experience.
Skills that matter here
- Kotlin: Used for building and extending platforms that improve system reliability.
- Kubernetes: Involves managing and improving the reliability of containerized applications and infrastructure.
- Amazon Web Services: Requires experience running production systems on AWS for high availability.
- DataDog: Used for monitoring and observability, including building and tuning alerts for system health.
- Terraform: Applied in managing infrastructure as code to ensure consistent and reliable deployments.
- Event driven architectures: Relevant for building and maintaining distributed platforms that enable scalable product development.
Who this role suits
- A technically driven individual who enjoys deep diving into complex systems to find and fix root causes.
- Someone who naturally considers how AI can accelerate problem-solving and reduce manual effort in operations.
- An engineer with strong leadership in incident management, capable of structured triage and mitigation under pressure.
- A curious and autonomous professional with a strong sense of accountability for system reliability.
From the employer
- Build and extend platforms to improve system reliability
- Work on team goals that encompass reliability for the entire company
- Standardize reliability tools across multiple platforms and organizations
- Triage, coordinate, and lead stabilization of sev 0–1 incidents
- Serve as primary oncall, maintaining structured escalation paths and exercising leadership escalation
- Drive platform-wide reliability improvements, shared operational tooling, and deploy-safety patterns
- Use AI-driven systems to improve signal detection, reduce noise, and accelerate root cause analysis
- Design and implement safe deployment patterns (progressive delivery, automated rollback, guardrails)
- Drive to root cause systems with many moving parts and take the necessary steps to fix them
- Demonstrated technical initiative and leadership on previous projects, especially those with a backend/platform focus
- Familiarity with AI-driven tooling for observability, incident analysis, or automation
- A mindset that naturally reaches for AI to accelerate problem-solving and reduce toil
- Experience running production oncall for high-availability systems
- Strong incident management skills — structured triage, mitigation under pressure, blameless postmortems
- Fluency with CI/CD pipelines, progressive rollout strategies, and rollback automation
- Monitoring & observability expertise — building/tuning alerts for uptime, error rates, latency regression, and resource exhaustion
- Ability to create and maintain evidence-based maturity assessments using trailing 90-day data windows
- Comfort with vendor/dependency management — maintaining validated escalation contacts reachable within ≤ 5 minutes
- Boundless curiosity, autonomy, and a strong sense of accountability
- A strong desire to perform and grow as an engineer
- 5+ years of software development experience
- This program shifts Block from reactive incident handling to repeatable, system-wide reliability gains — fewer customer-visible incidents, faster response, higher product velocity, and lower burnout across the organization.
- Block takes a market-based approach to pay, and pay may vary depending on your location. U.S. locations are categorized into one of four zones based on a cost of labor index for that geographic area. The successful candidate’s starting pay will be determined based on job-related skills, experience, qualifications, work location, and market conditions. These ranges may be modified in the future.
- Zone A: USD $189,000 - USD $283,600
- Zone B: USD $179,600 - USD $269,400
- Zone C: USD $170,100 - USD $255,100
- Zone D: USD $160,700 - USD $241,100
Questions about this role
What is the remote work policy for this role?
This is a fully remote position.
What level of seniority is expected for this position?
This is a senior-level role, requiring demonstrated technical initiative and leadership.
What are the core technical skills required?
Key skills include Kotlin, Modern Java, HTTP, JSON, gRPC, MySQL, Vitess, DynamoDB, event-driven architectures, DataDog, LaunchDarkly, Terraform, Kubernetes, Istio, Envoy, and Amazon Web Services.