Machine Learning Data Analyst

Remote $68k–$112k middle 3 months ago full-time quality 8.6/10

Role in brief

Incode is looking for a Machine Learning Data Analyst to ensure the quality and reliability of data used in their identity verification ML models. This role involves investigating data issues, building automated pipelines, and defining evaluation metrics. Candidates with strong SQL and Python skills, experience with data orchestration, and a background in mathematics or engineering should consider applying.

Data QualityPythonSQLAutomated Data PipelinesData infrastructureIdentity VerificationStatistical foundationCryptoML ModelsPrefectMLOpsDocument Intelligence

About the role

This role focuses on maintaining the integrity of data that powers machine learning models for identity verification solutions. The Machine Learning Data Analyst will independently investigate data quality issues, forming hypotheses, querying various data sources, and providing clear, evidence-backed answers. This involves taking ownership of problems when model metrics decline or data appears incorrect, ensuring that the underlying data is accurate and reliable for model performance.

A key responsibility is the design, construction, and maintenance of automated data pipelines. These pipelines handle data collection, labeling, validation, and metric computation, all essential for ML model training and evaluation. The analyst will also establish and monitor data and labeling quality standards, conducting accuracy audits and root-cause analyses when issues arise. Success in this position means developing robust systems for tracking model performance and building scalable dashboards to support data-driven decisions across teams.

The role requires developing and operating reliable workflow orchestration using tools like Airflow or Prefect to manage end-to-end pipelines. This includes scheduling, observing, and troubleshooting data flows. Collaboration is crucial; the analyst will partner with ML engineers, other analysts, and product stakeholders to prioritize work, resolve blockers, and continuously improve internal tools for data analysis and evaluation. The ideal candidate will write clean, maintainable SQL and Python code to efficiently investigate problems, query databases, and process data from multiple sources.

The salary for this role ranges from $68,000 to $112,000 annually.

Skills that matter here

  • Data Quality: This role involves establishing and monitoring consistency checks and accuracy audits to ensure the reliability of data impacting ML model outcomes.
  • Python: You will use Python for efficiently investigating data problems, querying databases, and processing data across multiple sources.
  • SQL: Strong SQL skills are essential for data investigation, querying across various data sources, and performing root cause analysis.
  • Automated Data Pipelines: A core responsibility is to design, build, and maintain pipelines for data collection, labeling, validation, and metric computation that support ML training and evaluation.
  • AWS Redshift: Hands-on experience with AWS Redshift or similar columnar/cloud databases is required for data investigation and management.
  • Prefect: You will develop and operate reliable orchestration using tools like Prefect to schedule, observe, and troubleshoot end-to-end pipelines.

Who this role suits

  • A person who thrives on independently investigating complex data issues, forming hypotheses, and delivering evidence-backed solutions.
  • Someone with a proactive mindset, eager to identify problems and implement robust data quality and pipeline solutions.
  • An individual who enjoys defining and automating evaluation metrics that accurately reflect real-world product usage and business objectives.
  • A collaborative professional who can effectively partner with ML engineers, analysts, and product stakeholders to prioritize work and improve tooling.

From the employer

What You'll Own & Drive

  • Root Cause Investigation — Independently investigate data quality issues end-to-end. When a model metric drops or data looks wrong, you own the investigation: forming hypotheses, querying across data sources, and delivering a clear, evidence-backed answer.
  • Automated Data Pipelines — Design, build, and maintain pipelines for collection, labeling, validation, and metric computation that support ML training and evaluation.
  • Data & Labeling Quality Standards — Establish and monitor consistency checks, accuracy audits, and root-cause analysis when issues impact model outcomes.
  • Model Evaluation Metrics — Define, implement, and automate evaluation metrics and reporting that reflect real-world product use cases and business goals.
  • Performance Tracking Systems — Build scalable dashboards and monitoring to enable fast, data-driven decisions across teams.
  • Workflow Orchestration — Develop and operate reliable orchestration (Airflow, Prefect, or similar) to schedule, observe, and troubleshoot end-to-end pipelines.
  • Clean, Maintainable Code — Write SQL and Python to efficiently investigate problems — querying databases, calling internal APIs, and processing data across multiple sources.
  • Cross-Functional Partnership — Partner closely with ML engineers, analysts, and product stakeholders to prioritize work by impact, unblock execution, and continuously improve internal tooling for analysis and evaluation.

Your Background

  • 3+ years of experience as a Data Analyst or in a similar data infrastructure role.
  • Strong SQL and Python skills for data investigation and root cause analysis.
  • Hands-on experience with AWS Redshift or a similar columnar/cloud database (BigQuery, Snowflake, etc.).
  • Solid statistical foundation — you can reason about rates, distributions, significance, and sampling bias.
  • Hands-on experience with workflow orchestration tools (Airflow, Prefect, Dagster, etc.).
  • Proven experience in data quality management, data preparation, or ML data pipelines.
  • A proactive mindset toward identifying problems.
  • Strong collaboration and problem-solving skills.
  • Background in mathematics, physics, or engineering.

Why Incode?

  • Mission with Meaning — Build systems that enable ethical, seamless identity verification for millions.
  • Rocket-Ship Growth — Join a company scaling globally with AI at its core.
  • Elite Team & Technology — Collaborate with top engineers and data scientists redefining document intelligence.
  • Ownership & Autonomy — Operate with end-to-end responsibility for impactful data pipelines.
  • Global Impact — Your work will power real-world AI experiences trusted by major enterprises.

Benefits & Perks:

  • Flexible Working Hours & Workplace
  • Open Vacation Policy
  • Equal Opportunities: Incode is an equal opportunity employer, committed to creating a diverse and inclusive work environment.

Questions about this role

What is the remote work policy for this position?

This is a fully remote position.

What level of seniority is this role?

This position is for a middle-level professional.

What are the key technical skills required?

Candidates should have strong SQL and Python skills, experience with AWS Redshift or similar databases, a solid statistical foundation, and hands-on experience with workflow orchestration tools like Airflow or Prefect.

Similar jobs

Before you apply

  • Legitimate employers never ask you to pay anything to apply or get hired.
  • Never share seed phrases or private keys. No real job needs them.
  • Do not install software ("test tasks", "trading tools", "video call clients") sent during hiring.
  • Check that the application page's domain really belongs to Incode.