Troveo AI Logo

Troveo AI

Data Engineer

Posted 22 Days Ago
Remote
Hiring Remotely in USA
100K-140K Annually
Senior level
Remote
Hiring Remotely in USA
100K-140K Annually
Senior level
Design, build, and maintain scalable ELT/ETL and streaming data pipelines into cloud warehouses. Implement transformations and dimensional models (dbt), ensure data quality and observability, write SQL for analysis, support BI/dashboarding, monitor SLAs, troubleshoot incidents, participate in on-call rotations, and collaborate cross-functionally to deliver analytics-ready datasets.
The summary above was generated by AI

About Troveo

Troveo builds the data platform that AI labs and model builders need to train the next generation of models. We have created the world's largest licensed platform of scarce, proprietary data for AI, spanning video, audio, text, and business workflows.

Troveo indexes, enriches, and packages this high-quality data into formats ready for training, fine-tuning, evaluation, and agentic use cases. Backed by top investors, we’re a small, high-impact team solving one of the biggest bottlenecks in AI development.

Role Overview

We are seeking a versatile, hands-on Data Engineer to build and maintain a scalable analytics data warehouse while contributing to data modeling, performing data analysis, and ensuring the reliable delivery of data to downstream teams and systems. This hybrid role combines core data engineering responsibilities with data modeling, analytics, and operational support. You will own the full analytics data lifecycle, from ingestion and transformation to modeling, quality assurance, and timely delivery, while partnering closely with software engineers and business stakeholders.

Key Responsibilities

Data Pipeline Engineering

  • Design, build, and maintain ELT/ETL data pipelines (batch and streaming), optimizing for performance, reliability, scalability, and cost.

  • Design and implement conceptual, logical, and physical data models (including dimensional modeling, star/snowflake schemas).

  • Build and maintain transformation layers using modern tools (e.g., dbt) to create clean, well-documented, analytics-ready datasets.

  • Apply data modeling best practices, versioning, testing, and documentation to ensure consistency and reusability.

Data Analysis & Reporting Support

  • Write optimal SQL queries for data exploration, ad-hoc analysis, and troubleshooting.

  • Support the creation of reports, dashboards, and self-service analytics assets in collaboration with data analysts and business teams.

  • Translate business questions into data requirements and deliver actionable insights or datasets.

Operational Support & Data Deliveries

  • Monitor data pipelines and data delivery processes to ensure SLAs for timeliness, freshness, and accuracy are consistently met.

  • Proactively identify, troubleshoot, and resolve data issues impacting downstream consumers or business operations.

  • Manage incidents related to data availability and quality; participate in on-call rotations as needed.

  • Implement data quality checks, observability, and alerting to maintain high reliability of data deliveries.

  • Automate operational tasks and continuously improve data delivery processes.

Collaboration & Best Practices

  • Work cross-functionally with analysts, data scientists, engineers, and business stakeholders to understand data needs and deliver solutions.

  • Document data pipelines, models, lineage, and processes.

  • Contribute to data governance, security, and best practices across the data platform.

Requirements

  • 7+ years of professional experience in data engineering or a closely related role (analytics engineering experience is highly relevant).

  • Strong proficiency in SQL and Python.

  • Hands-on experience building and maintaining data pipelines and working with cloud data platforms/warehouses. (Snowflake, BigQuery, Redshift, Databricks, etc.).

  • Experience with data orchestration tools (Apache Airflow, Dagster, Prefect, or similar).

  • Solid understanding of data modeling techniques and dimensional modeling.

  • Experience performing data analysis and working with BI/visualization tools (Looker, Tableau, Power BI, or similar).

  • Proven ability to troubleshoot data issues and support operational reliability/SLAs.

  • Strong communication skills and ability to collaborate with both technical and non-technical stakeholders.

Bonus Points

  • Experience with DBT for data transformation and modeling.

  • Knowledge of data observability/monitoring tools.

  • Experience with real-time/streaming data technologies (Kafka, Flink, etc.).

  • Familiarity with CI/CD practices for data pipelines.

  • Experience in data quality frameworks and governance.

  • Bachelor’s degree in Computer Science, Engineering, or a related quantitative field (or equivalent practical experience).

Compensation

  • Base Salary: $100,000 – $140,000 (depending on experience and location)

  • Equity: Competitive equity package in a well-funded AI startup with significant upside

Compensation is location-adjusted for cost of living. We are open to candidates in California, New York, and select other states.

What We Offer

  • Comprehensive Health Benefits: Medical, dental, and vision coverage (100% employer-paid for employees)

  • Flexible PTO & Paid Holidays: Unlimited PTO with encouragement to actually use it

  • Remote First Policy: Work from anywhere in the US (with occasional team offsites)

  • Learning & Growth: Annual learning stipend, access to top conferences, and direct mentorship from experienced founders

  • Equity Ownership: Competitive equity package with clear growth potential as we scale

  • Modern Tech Stack & Tools: Budget for the best equipment and software

  • Strong Culture: High-trust, low-ego environment focused on impact, transparency, and work-life balance. We believe great work happens when people are supported, challenged, and given ownership.

Equal Opportunity Employer

Troveo is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees. We do not discriminate based on race, religion, color, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status.

Similar Jobs

10 Days Ago
Remote or Hybrid
United States
70K-120K Annually
Mid level
70K-120K Annually
Mid level
Cloud • Insurance • Payments • Software • Business Intelligence • App development • Big Data Analytics
Build, model, and maintain scalable BigQuery-based data solutions on GCP. Implement performant data models, storage partitioning/clustering, ETL improvements, data validation, and documentation. Collaborate with architects, data scientists, and engineers to deliver governed, high-quality data for reporting and AI initiatives.
Top Skills: Ansi SqlBigQueryBigquery SqlConfluenceGCPJIRAPythonSQL
3 Days Ago
Easy Apply
Remote
United States
Easy Apply
Senior level
Senior level
Enterprise Web • Mobile • Professional Services • Software
Design, build, and own scalable data pipelines and evaluation systems that power production AI features and internal analytics. Ensure data quality across ingestion, modeling, and reporting, collaborate with ML and analytics teams, deploy and monitor ML systems, and establish standards for data work and evaluation.
Top Skills: AirflowAWSDagsterGCPPostgresPythonSnowflake
11 Days Ago
Remote
United States
225K-225K Annually
Expert/Leader
225K-225K Annually
Expert/Leader
Security • Cybersecurity
Lead design and modernization of scalable, secure data systems and real-time pipelines for xOT cybersecurity. Create data contracts, observability, lineage, and cross-functional solutions supporting cloud and on-prem deployments, Kubernetes, and stream processing to enable threat detection and intelligence.
Top Skills: CloudData Observability FrameworksDockerGoJvm (Java/Scala/Kotlin)KubernetesMessage QueuesNode.jsOn-PremisesPythonRustStream Processing

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account