Role in brief
Hibachi is seeking a Data Engineer to build and maintain data pipelines using PySpark, AWS Glue, and Airflow. This role involves designing data models, implementing Change Data Capture, and developing dashboards for insights. Candidates with experience in AWS data services and strong SQL skills, who enjoy working with both batch and streaming data, should consider applying.
About the role
This Data Engineer position focuses on the full lifecycle of data, from ingestion to visualization. The role requires architecting, building, and maintaining both batch and streaming data pipelines. A key responsibility is implementing Change Data Capture (CDC) using tools like AWS DMS to ensure incremental data updates are efficiently handled. The work involves creating modular, reusable, and scalable data models to support various analytical needs.
The role involves managing ETL/ELT processes, ensuring data is efficiently ingested, cleansed, and aggregated for analysis. A significant part of the job includes developing dashboards using tools like QuickSight to present actionable insights derived from the processed data. This requires a strong understanding of how to transform raw data into meaningful business intelligence.
Success in this role means ensuring data pipelines are robust and efficient, accurately capturing and processing data, and providing clear, insightful visualizations. The ideal candidate will be adept at using PySpark, AWS Glue, and Airflow to deliver reliable data solutions that support the company's analytical capabilities.
The salary for this position ranges from $140,000 to $200,000 annually.
Skills that matter here
- PySpark: This role uses PySpark for building and maintaining both batch and streaming data pipelines.
- AWS Glue: Proficiency in AWS Glue is required for developing and managing data pipelines within the AWS ecosystem.
- Airflow: Airflow is used for orchestrating and scheduling data pipeline workflows.
- AWS DMS: This skill is essential for implementing Change Data Capture to manage incremental updates in data.
- SQL: Advanced SQL knowledge is necessary for data manipulation, querying, and model design.
- QuickSight: This tool is used to develop dashboards that surface actionable insights from processed data.
Who this role suits
- A candidate with a Bachelor’s or Master’s Degree in Computer Science, Engineering, or a related field.
- Someone with at least two years of experience specifically with PySpark in data pipeline development.
- An individual who is proficient in AWS data services and understands ETL/ELT processes.
- A person who can effectively communicate technical concepts and collaborate on data solutions.
From the employer
- Architect, build, and maintain batch and streaming data pipelines using PySpark, AWS Glue, and Airflow.
- Implement Change Data Capture (CDC) with AWS DMS to capture incremental updates.
- Design modular, reusable, and scalable data models.
- Manage ETL/ELT pipelines ensuring efficient data ingestion, cleansing, and aggregation.
- Develop QuickSight dashboards to surface actionable insights.
- Bachelor’s or Master’s Degree in Computer Science, Engineering, or a related field.
- 2+ years of experience with PySpark for batch and streaming pipelines.
- Proficiency in AWS Glue, Apache Airflow, and Iceberg.
- Experience with AWS DMS or other CDC tools.
- Advanced SQL knowledge.
- Experience with BI platforms like QuickSight, Tableau, or Power BI.
- Understanding of testing frameworks like Pytest.
- Excellent communication skills.
Questions about this role
What is the remote work policy for this position?
This is a fully remote position.
What level of experience is required for this role?
Candidates should have a minimum of 2 years of experience with PySpark for batch and streaming pipelines.
What are the primary technical skills needed for this role?
Key technical skills include PySpark, AWS Glue, Airflow, AWS DMS, advanced SQL, and experience with BI platforms like QuickSight, Tableau, or Power BI.