About the Role
We are building the data foundation that powers the full machine learning lifecycle for autonomous robotics. Our robots generate large-scale, multimodal datasets across real-world deployments, and turning that raw experience into reliable, discoverable, high-quality ML data is a core part of improving our autonomy systems.
As a Data Platform Engineer, you will help design and build the platform that manages data from ingestion through processing, validation, labeling, dataset generation, training, and evaluation.
This is not a traditional analytics data engineering role. You will work closely with ML engineers, researchers, labeling teams, robotics engineers, and infrastructure engineers to build scalable systems for robotics and ML data.
What You'll Do
- Design and build scalable data architecture for large-scale multimodal robotics and ML datasets.
- Build abstractions and services for ingestion, processing, datasets, metadata, lineage, and data quality.
- Develop reliable pipelines for transforming raw robot data into versioned, ML-ready data products.
- Define data models, schemas, contracts, and lifecycle states across data-processing workflows.
- Build systems for tracking provenance and lineage across raw data, derived artifacts, labels, datasets, and downstream ML workloads.
- Develop automated validation and data-quality frameworks that detect incomplete, corrupted, or unusable data early.
- Design for incremental processing, reprocessing, backfills, and versioned transformations.
- Improve observability and failure diagnosis across complex data workflows.
- Partner with ML and robotics teams to understand domain-specific data requirements and turn recurring patterns into reusable platform capabilities.
- Work closely with infrastructure/platform teams on storage, compute, orchestration, reliability, and scalability.
What We're Looking For
- Strong experience building production data platforms or large-scale data-processing systems.
- Strong software engineering skills, preferably Python and/or C++/Java/Go.
- Experience with distributed data processing and workflow orchestration.
- Experience with data lakes/lakehouses, object storage, metadata systems, schemas, and data versioning.
- Strong understanding of data quality, lineage, reproducibility, and reliable pipeline design.
- Experience with technologies such as S3, Airflow/Dagster, Spark/Ray, Kubernetes, Parquet, or similar systems.
- Ability to work across ambiguous organizational and technical boundaries.
- Strong systems-design and engineering judgment.
Nice to Have
- Experience with ML datasets or ML infrastructure.
- Robotics, autonomous vehicles, sensor, video, image, LiDAR, or other multimodal data.
- Experience building internal developer/platform products.
- Experience operating pipelines at TB/PB scale.
FieldAI Irvine, California, USA Office
Irvine, California, United States, 92602
Similar Jobs
What you need to know about the Los Angeles Tech Scene
Key Facts About Los Angeles Tech
- Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
- Key Industries: Artificial intelligence, adtech, media, software, game development
- Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
- Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering


