Data Engineer

Leonberg, GermanyFullTimePosted Jul 22, 2026

The Role 

As a Data Engineer within the Machine Learning team in Application Software, you’ll contribute to critical initiatives that push the frontier of model-based autonomous driving—both in terms of core driving performance and feature-level intelligence such as personalization, comfort, and collaboration.

You’ll design and deliver scalable data pipelines that transform vast amounts of data from diverse internal and external sources into structured, reliable, and model-ready datasets. Your work will span data ingestion, data quality assurance, transformation, curation, evaluation and ML support. You’ll collaborate deeply with Wayve’s Data Corpus teams and ML engineers to build systems that are performant, adaptable, and ready for production.

 

Key Responsibilities:

  • Build and improve scalable data pipelines that support model development, evaluation, and production ML workflows for autonomous driving.

  • Ingest, transform, and curate large-scale real-world, synthetic, and partner-provided datasets into structured, reliable, and model-ready formats aligned with standardised taxonomies and coordinate systems.

  • Develop data quality checks, validation processes, and monitoring to ensure both raw data from our vehicle platforms and processed datasets are high-quality, complete, consistent, traceable, and fit for ML use cases.

  • Curate and mine real-world and synthetic data to drive scenario diversity, coverage, and feature-specific development.

  • Improve pipeline performance, reliability, and usability, helping reduce bottlenecks and increase iteration velocity across ML development.

  • Collaborate closely with Machine Learning engineers, Data Corpus, AI Platform, and external partners to ensure data pipelines integrate effectively with production-scale learning systems.

 

About You 

In order to set you up for success as a Data Engineer at Wayve, we’re looking for the following skills and experience.  

Essential 

  • Proven experience building and operating scalable data pipelines or distributed data processing systems in production environments.

  • Strong software engineering skills in Python, with a solid foundation in maintainable, reliable, and well-tested software development practices.

  • Proficient in SQL and PySpark, with experience using warehouse/OLAP concepts, window functions, and Spark for distributed data processing.

  • Experience with modern data pipeline architectures, including workflow orchestration and DAG-based systems such as Airflow, Flyte, Ray, or similar.

  • Solid understanding of robotics and automated driving data concepts, including sensor characteristics, timestamping and clock synchronisation, coordinate transformations, calibration, and ego-motion signals such as GNSS/IMU and vehicle odometry.

  • Understanding of machine learning development workflows, including training data generation, evaluation datasets, scenario mining, and model iteration.

  • Excellent communication and collaborative skills, capable of working effectively with interdisciplinary teams.

 

Desirable 

  • Prior work in calibration, perception, imitation learning, or trajectory prediction.

  • Experience with third-party dataset ingestion and transformation.

  • Familiarity with automated driving data and scenario taxonomies, including ODD definitions, manoeuvre and behavior labels, scene classification, event tagging, semantic understanding.

 

This is a full-time role based in our office in Leonberg. At Wayve we want the best of all worlds so we operate a hybrid working policy that combines time together in our offices and workshops to fuel innovation, culture, relationships and learning, and time spent working from home. We operate core working hours so you can determine the schedule that works best for you and your team.  

#LI-KM1

Want jobs like this matched to you?

Swoopd scores fresh postings against your résumé so you only see the matches that matter.

Get started free