Wynd Labs Logo

Wynd Labs

Data Engineer

Posted 5 Days Ago
Remote
Hiring Remotely in USA
Entry level
Remote
Hiring Remotely in USA
Entry level
Build and operate large-scale data pipelines, web-scraping systems, and data-collection tools for AI datasets. Maintain analytical databases, optimize queries, troubleshoot pipeline and data-quality issues, manage containerized workloads across Kubernetes and Linux infrastructure, and support CI/CD workflows. The role also involves documenting technical processes, researching improvements, and developing scalable APIs in a fully remote environment.
The summary above was generated by AI

Who We Are:

We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models.

We're the team that helps to power and support Grass, a bandwidth-sharing network that lets us operate a massive distributed crawler, giving us unique access to high-quality public web data at global scale. On top of that, we’ve built pipelines for ingesting, segmenting, and annotating billions of videos, transcripts, and audio files, powering dataset creation for frontier labs.

We’re lean, technical, and move fast. No red tape, no slow decision-making; just a team of builders pushing to expand what’s possible for open web data and AI.

The Role:

We are seeking a Data Engineer to support and improve large-scale data pipelines and infrastructure. You’ll work across data collection, processing, transformation, validation, and delivery, with a focus on scalability, reliability, and performance.
This is a hands-on role where you’ll work with distributed systems, large datasets, web scraping infrastructure, and production data workloads.

Please note: This role requires a work schedule that overlaps sufficiently with EST business hours to collaborate effectively with the team.

Who You Are:

  • Bachelor’s degree or equivalent work experience

  • Python (advanced) — strong grasp of async programming, multiprocessing, and writing production-grade code for long-running data jobs

  • Web scraping at scale — hands-on experience with high-volume scraping (proxies, rate limiting, anti-bot evasion). Experience with platform APIs and large media/metadata datasets (video platforms, social media)

  • Distributed data pipelines — experience designing and operating pipelines across many workers/servers using task queues (Celery, Kafka, RabbitMQ, or similar)

  • Data warehousing — practical experience with columnar/analytical warehouses; Databend, ClickHouse, or BigQuery strongly preferred; comfortable with complex analytical queries, partitioning strategies, cost-aware querying on cloud warehouses

  • Docker & Kubernetes — containerizing workloads, writing Helm charts/manifests, managing deployments, autoscaling scraping/processing workloads

  • Linux & bare-metal ops — comfortable managing services on Linux servers, debugging performance issues (disk I/O, network, memory) without managed-cloud abstractions

  • CI/CD for data workflows (GitHub Actions, ArgoCD)

  • Writing Scalable API

What You'll Be Doing:

  • Maintain, optimize, and troubleshoot database queries and related data systems to support efficient data access, processing, and reliability.

  • Assist in creating, maintaining, and improving data pipelines used to collect, process, transform, validate, and deliver large-scale datasets.

  • Support web scraping and data collection initiatives, including developing, testing, and maintaining scripts or tools used to gather publicly available data in accordance with Company requirements.

  • Monitor and troubleshoot data pipeline issues, identify data quality concerns, and help implement timely fixes to maintain data accuracy and operational continuity.

  • Document engineering work, including database queries, pipeline processes, scraping workflows, technical decisions, issues encountered, and resolutions implemented.

  • Participate in research and development projects to improve the Company’s data products and workflows.

Why Work With Us:

  • Opportunity. We are at the forefront of developing a web-scale crawler and knowledge graph that improves access to public web data and extends the value of AI to the people.

  • Culture. We're a lean team with a high bar. We come to work not to be comfortable, but to find out what we're capable of and to do work that matters. We're not calling for people who keep things moving. We're calling for people who make everyone around them better.
    We prioritize low ego and high output. This is a fully remote team.

  • Compensation. You’ll receive a competitive salary, benefits and equity package.

Similar Jobs

3 Days Ago
Easy Apply
Remote
USA
Easy Apply
191K-225K Annually
Senior level
191K-225K Annually
Senior level
Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Build and operate low-latency market data systems for institutional trading, including feed handlers, normalization pipelines, venue connectivity, and distribution services. Develop high-throughput services, improve reliability and performance through observability and incident response, participate in on-call support, and collaborate with engineering and product teams. The role requires production backend engineering experience, market data infrastructure expertise, Java or C++, messaging frameworks, and exchange connectivity protocols.
Top Skills: AeronC++FixItch/OuchJavaMulticastSbe
3 Days Ago
Remote or Hybrid
United States
160K-260K Annually
Expert/Leader
160K-260K Annually
Expert/Leader
Artificial Intelligence • Cloud • Payments • Software • Business Intelligence • Generative AI • Automation
Define and govern enterprise-scale data architecture across batch, streaming, warehouse, lakehouse, transactional, and AI use cases. Establish standards for data quality, lineage, access, cataloging, governance, observability, and SLAs. Architect AI-enabled workflows, resolve complex architecture issues, influence roadmaps, and mentor engineers through hands-on technical leadership. The role requires 15+ years of software, data engineering, or architecture experience and expertise in large-scale data platforms and modeling.
Top Skills: AIBatch ProcessingBigQueryData CatalogsData WarehousesDbtFeature StoresGCPLakehousesOlapOltpStreaming ArchitecturesVector Stores
13 Days Ago
Easy Apply
Remote or Hybrid
United States
Easy Apply
Senior level
Senior level
Fintech • News + Entertainment • Software • Database • Financial Services
Lead the architecture and development of scalable AWS-based data ingestion, transformation, and orchestration pipelines. Build reliable data infrastructure using Python, SQL, Airflow, Lambda, ECS, SQS, and Terraform. Establish data modeling, quality, lineage, monitoring, testing, and observability practices while partnering with analysts, scientists, and backend engineers. Mentor senior engineers, guide technical decisions, and provide hands-on leadership for complex data platform initiatives.
Top Skills: Amazon EcsAmazon KinesisAmazon RedshiftAmazon S3Amazon SqsApache AirflowApache FlinkAWSAws GlueAws LambdaBeautifulsoupCi/CdDatabricksDockerGreat ExpectationsKafkaMonte CarloMwaaPythonScrapySnowflakeSQLTerraform

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account