Veeda AI Logo

Veeda AI

Senior Machine Learning Engineer — Data Infrastructure

Posted Yesterday
Remote or Hybrid
Hiring Remotely in California, USA
Senior level
Remote or Hybrid
Hiring Remotely in California, USA
Senior level
Design, build, and operate large-scale crawling, ingestion, and ETL/ELT pipelines for petabyte-scale image and video datasets. Optimize distributed processing (Ray/Spark/Dask), ensure dataset provenance/versioning, reduce cost/latency, and collaborate with CV engineers and ML scientists. Provide technical leadership and mentor junior engineers.
The summary above was generated by AI
About Us

Veeda AI is building the next generation of multimodal foundation world models for Physical AI. We're a small, fast-moving team of engineers and researchers from leading AI labs, tackling some of the most challenging problems at the intersection of AI, robotics, and embodied intelligence. If you're excited about pushing the boundaries of what's possible with Physical AI, you'll have the opportunity to make an outsized impact from day one.

Responsibilities
  • Crawling & Ingestion at Scale: Architect and own the scraping and crawling infrastructure that acquires petabytes of image and video data — designing systems that are resilient, compliant, and able to keep pace with a fast-moving content landscape.

  • Image-Data Pipeline Design: Build and maintain production data pipelines that clean, transform, and process raw visual assets—handling format diversity, resolution variance, metadata extraction, and noise at scale.

  • ETL/ELT for Training & Evaluation: Design ETL/ELT workflows that prepare image and video datasets for ML model training and evaluation, including annotation ingestion, quality scoring, and reproducible dataset versioning.

  • Performance & Cost Optimization: Continuously profile and optimize pipelines for throughput, latency, and cloud storage/compute costs—ensuring research velocity without unnecessary overhead.

  • Distributed Processing Ownership: Select, deploy, and operate distributed data-processing frameworks (e.g., Ray, Spark, Dask) to handle compute-intensive tasks like large-batch image transformation, embedding generation, and feature extraction.

  • Cross-Functional Partnership: Work directly with Computer Vision Engineers, ML Scientists, and product teams to translate data requirements into reliable, well-documented pipelines and close the feedback loop between model performance and data quality.

  • Technical Leadership: Lead data-pipeline design reviews, mentor junior and mid-level data engineers, establish best practices for the data team, and raise the overall engineering bar.

Requirements
  • You have 5+ years of professional experience in data engineering or a closely related discipline with a strong foundation in distributed systems, database design, and handling large binary assets (images, video) at scale.

  • You possess strong programming skills in Python and SQL, with proven experience designing end-to-end production systems rather than basic scripts.

  • You have designed and operated web scraping or data ingestion systems running reliably in production at significant global scale.

  • You are hands-on with data-processing frameworks (e.g., PySpark, Pandas, Ray) and comfortable with image-processing libraries (OpenCV, Pillow, scikit-image).

  • You have built ETL/ELT pipelines for ML/CV use cases with a strong emphasis on dataset provenance, lineage, and reproducibility.

  • You have extensive experience with cloud object storage (AWS S3, GCP GCS) and understand the cost, latency, and durability trade-offs when managing petabyte-scale data assets.

  • Production Quality, Agent Velocity: Your daily workflow runs through AI coding harnesses (e.g., AI agents/assistants) without sacrificing software engineering rigor. You review agent code diffs with the same scrutiny as a team member's PR, recognize AI code generation failure modes, and ship rapidly without introducing technical debt or "slop."

Nice to Have
  • Experience with Ray for distributed data processing or ML-adjacent workflows.

  • Hands-on experience with data orchestration platforms (e.g., Airflow, Dagster, Prefect) to manage complex processing DAGs in production.

  • Familiarity with Computer Vision concepts and image annotation formats (COCO, PASCAL VOC, YOLO).

  • Practical knowledge of MLOps concepts, including data versioning (e.g., DVC, LakeFS) and experiment tracking for CV pipelines.

  • Experience with containerization (Docker) and orchestration (Kubernetes) for scaling data processing workloads.

  • Experience navigating multi-region data compliance challenges (GDPR, content licensing, and image rights across jurisdictions).

Similar Jobs

7 Hours Ago
In-Office or Remote
Site of Old Bullion, NV, USA
177K-294K Annually
Senior level
177K-294K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Lead HEOR and evidence strategy for obesity and internal medicine early-pipeline assets. Design and oversee RWD studies, economic models, and patient-reported outcomes; influence trial design; manage evidence generation, publications, vendor relationships, and cross-functional stakeholder engagement to secure reimbursement and patient access.
Top Skills: ChatgptMicrosoft Copilot
7 Hours Ago
In-Office or Remote
Site of Old Bullion, NV, USA
177K-294K Annually
Senior level
177K-294K Annually
Senior level
Artificial Intelligence • Healthtech • Machine Learning • Natural Language Processing • Biotech • Pharmaceutical
Lead global HEOR and evidence-generation strategy for obesity assets. Design and execute RWE, economic models, and PRO strategies; create launch dossiers; manage cross-functional teams, vendors, budgets, and external partnerships to support reimbursement and patient access.
Top Skills: ChatgptMicrosoft Copilot
12 Hours Ago
Remote
Senior level
Senior level
Artificial Intelligence • Big Data • Cloud • Information Technology • Software • Cybersecurity • Data Privacy
Own technical relationship for Swiss enterprise and mid-market accounts: run discovery, demos, and technical validations (POC), architect Rubrik solutions across on‑prem, cloud, SaaS, identity and AI data workflows, qualify opportunities, present to technical and CxO stakeholders, and partner with account teams to close deals.
Top Skills: Ai-Driven Data WorkflowsBackup And Disaster RecoveryCloudCyber ResilienceData ProtectionData ResilienceIdentity SecurityObservabilityRemediationRubrikRubrik Agent CloudRubrik Security CloudSaaS

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account