Role in brief
Hibachi is seeking a Production Engineer to manage alert responses, investigate incidents, and enhance observability tools. This role involves maintaining live production systems and improving monitoring infrastructure. Candidates with experience in incident response and strong diagnostic skills across logs, databases, and code are encouraged to apply.
About the role
This role focuses on maintaining the stability and performance of live production systems. The Production Engineer will be the first line of defense against system issues, responding to alerts and triaging problems as they arise. A key part of the work involves detailed investigation, tracing issues through various system components like logs, databases, and code to pinpoint root causes and ensure quick resolution.
Beyond immediate incident response, the engineer will contribute to the long-term health of the system by improving observability. This includes maintaining and enhancing dashboards, building automation tools, and identifying areas where visibility can be strengthened. The goal is to proactively reduce incidents and improve the team's ability to understand system behavior.
Success in this position means consistently providing clear context during incidents, writing thorough summaries and runbooks for future reference, and continuously improving monitoring and tooling. The engineer will play a crucial role in ensuring system reliability and contributing to a robust operational environment.
The listed salary for this position is between $80,000 and $120,000 USD.
Skills that matter here
- Datadog: This tool will be used for monitoring, creating dashboards, and responding to alerts in live production systems.
- Grafana: The role requires familiarity with Grafana for building and maintaining dashboards to visualize system metrics.
- CloudWatch: Experience with CloudWatch is necessary for monitoring and gathering data from cloud-based services.
Who this role suits
- A person who thrives on solving complex problems under pressure.
- Someone who enjoys deep-diving into system diagnostics across different layers.
- An individual comfortable with on-call duties and taking ownership of production issues.
- A clear and concise communicator, especially when documenting incidents and solutions.
From the employer
- Respond to production alerts, triage, investigate, and provide context.
- Trace issues across logs, databases, and codebase to identify root causes.
- Write incident summaries and runbooks.
- Improve and maintain dashboards.
- Build tooling and automation.
- Identify and close visibility gaps.
- Experience with live production systems.
- Ability to read and trace codebases.
- Comfort writing database queries.
- Familiarity with monitoring tools like Datadog, Grafana, CloudWatch.
- Clear written communication.
- Comfort with on-call responsibilities.
Questions about this role
What is the remote work policy for this position?
This is a fully remote position.
What kind of experience is required for this role?
Candidates should have experience with live production systems, be able to read and trace codebases, and be comfortable writing database queries.
What is the salary range for this position?
The salary for this role ranges from $80,000 to $120,000 USD.