Seneca Holdings Logo

Seneca Holdings

Senior Data Engineer

Posted 7 Days Ago
Remote
Hiring Remotely in USA
Senior level
Remote
Hiring Remotely in USA
Senior level
Designs and maintains scalable Databricks data pipelines on AWS, covering ingestion, transformation, streaming, data quality, and analytics-ready dataset delivery. Builds medallion architecture layers using Spark, Python, SQL, Delta Lake, and AWS services. Responsibilities include source-to-target mapping, unit testing, CI/CD deployment, anomaly analysis, and compliance with FISMA High and multi-tenant governance requirements. Collaborates in Agile teams and uses AI automation tools to accelerate pipeline development and validation.
The summary above was generated by AI

Great Waters Federal is part of the Seneca Nation Group (SNG) portfolio of companies. SNG is Seneca Holdings' federal government contracting business that meets mission-critical needs of federal civilian, defense, and intelligence community customers. Our portfolio comprises multiple subsidiaries that participate in the Small Business Administration 8(a) program. To learn more about SNG, visit the website and follow us on LinkedIn.

Our team of talented individuals is what makes us successful. To support our team, we provide a balanced mix of benefits and programs. Your total rewards package includes competitive pay, benefits, and perks, flexible work-life balance, professional development opportunities, and performance and recognition programs. We offer a comprehensive benefits package that includes medical, dental, vision, life, and disability, voluntary benefit programs (critical illness, hospital, and accident), health savings and flexible spending accounts, and retirement 401K plan. One of our fundamental principles is to offer competitive health and welfare benefits to our team members, providing coverage and care for you and your family. Full-time employees working at least 30 hours a week on a regular basis are eligible to participate in our benefits and paid leave programs. We pride ourselves on our collaborative work environment and culture, which embraces our mission of providing financial and non-financial benefits back to the members of the Seneca Nation.


We are seeking an experienced Data Engineer to design, build, and maintain scalable data pipelines within a Databricks E2 environment hosted on AWS and governed by FISMA High and multi-tenant security and compliance standards.

The ideal candidate has strong hands-on experience with Databricks, Apache Spark, and the broader AWS data ecosystem. This individual should be comfortable working across the full data lifecycle—from ingestion and transformation to the delivery of analytics-ready datasets.

Key Responsibilities

  • Analyze and collect data from various sources, including relational databases, APIs, external data providers, and real-time streaming sources.
  • Design and implement efficient, scalable data pipelines to cleanse, transform, and aggregate data for downstream reporting and analytics.
  • Build and maintain Databricks solutions using medallion architecture, including Bronze, Silver, and Gold layers, to create structured, reliable, and progressively refined datasets.
  • Develop source-to-target mapping documentation for all data pipelines and conduct thorough unit testing to validate data accuracy and pipeline logic.
  • Write advanced SQL to create aggregations across complex datasets and analyze data for anomalies, quality issues, and inconsistencies.
  • Implement real-time and near-real-time data ingestion using AWS Database Migration Service (DMS) and other AWS-native services.
  • Develop data-processing logic using Python and/or R, leveraging Spark and related packages such as PySpark and Pandas for large-scale data transformation.
  • Manage source code, versioning, and deployments using GitLab, and support automated build and release processes through CI/CD pipelines.
  • Collaborate with cross-functional teams in an Agile project environment, participating in sprint planning, stand-ups, and iterative delivery.
  • Ensure that data engineering practices align with FISMA High and multi-tenant security, governance, and compliance requirements.
  • Leverage AI automation tools to support data engineering pipeline development, testing, and validation and to accelerate delivery.

Required Qualifications

  • Proven, hands-on experience as a Data Engineer or Databricks Developer building production-grade data pipelines.
  • Strong working knowledge of Databricks on AWS, including E2 architecture, cluster configuration, job orchestration, and workspace management.
  • Hands-on experience with Apache Spark, PySpark, and Pandas for large-scale, distributed data processing.
  • Solid programming skills in Python; working knowledge of R is a plus.
  • Experience with Databricks Auto Loader for scalable, incremental file ingestion.
  • Practical experience with AWS Database Migration Service (DMS) for change data capture and real-time data replication.
  • Strong experience with Delta Lake and Delta tables, including schema evolution, time travel, and optimization techniques such as OPTIMIZE, Z-ORDER, and VACUUM.
  • Experience with Amazon RDS and other relational database sources used in data extraction and integration workflows.
  • Advanced SQL skills, including complex joins, window functions, aggregations, and performance tuning.
  • Solid understanding of medallion architecture and modern data lakehouse design principles.
  • Experience with GitLab and CI/CD pipelines for automated testing, build, and deployment of data engineering code.
  • Experience working in Agile/Scrum project environments.

Preferred Skills

  • Broader knowledge of AWS services beyond DMS and RDS, including S3, Glue, Lambda, Step Functions, CloudWatch, and IAM.
  • Experience operating within FISMA High, FedRAMP, or other regulated, multi-tenant compliance environments.
  • Familiarity with data quality frameworks, anomaly detection techniques, and automated data validation.
  • Exposure to infrastructure-as-code tools such as Terraform or CloudFormation for provisioning data platform resources.
  • Knowledge of streaming technologies such as Kafka, Kinesis, or Spark Structured Streaming.

Equal Opportunity Statement:
Seneca Holdings provides equal employment opportunities to all employees and applicants without regard to race, color, religion, sex/gender, sexual orientation, national origin, age, disability, marital status, genetic information and/or predisposing genetic characteristics, victim of domestic violence status, veteran status, or other protected class status. This policy applies to all terms and conditions of employment, including, but not limited to, hiring, placement, promotion, termination, layoff, recall, transfer, leave of absence, compensation and training. The Company also prohibits retaliation against any employee who exercises his or her rights under applicable anti-discrimination laws. Notwithstanding the foregoing, the Company does give hiring preference to Seneca or Native individuals. Veterans with expertise in these areas are highly encouraged to apply.
 


Similar Jobs

3 Hours Ago
Remote or Hybrid
115K-145K Annually
Senior level
115K-145K Annually
Senior level
AdTech • Cloud • Digital Media • Information Technology • News + Entertainment • App development
Designs, builds, and supports scalable data pipelines, cloud-native processing systems, data models, APIs, and analytics solutions. The role ensures data quality, observability, testing, governance, CI/CD, and production reliability while partnering with engineering, product, analytics, BI, and business stakeholders. Responsibilities include architecture decisions, troubleshooting, documentation, Agile delivery, incident response, and mentoring engineers.
Top Skills: Amazon AthenaAmazon EmrAmazon MwaaAmazon RedshiftAmazon SnsAmazon SqsApache AirflowApache HiveApache IcebergSparkAPIsAws GlueAws LambdaAws S3Aws Step FunctionsCi/CdDatabricksGithub ActionsInfrastructure As CodeJavaLookerMicrostrategyPostgresScalaSinglestoreSnowflakeSQLTableau
3 Hours Ago
In-Office or Remote
139K-219K Annually
Senior level
139K-219K Annually
Senior level
Cloud • Information Technology • Productivity • Security • Software • App development • Automation
Designs scalable data architectures, ELT/ETL pipelines, data models, quality frameworks, and self-service data platforms. Partners with engineering, product, program, leadership, and data science teams to translate business needs into technical solutions. The role includes end-to-end technical ownership, large-scale batch and streaming data processing, mentoring junior engineers, system optimization, and on-call support for platform stability.
Top Skills: Amazon RedshiftApache AirflowBatch ProcessingBitbucketCi/CdData LakesDatabricksDatabricks ApisDynamoDBGitMachine LearningMongoDBPostgresPythonScalaSparkSQLStreaming
3 Days Ago
In-Office or Remote
92K-164K Annually
Senior level
92K-164K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Designs, develops, deploys, and supports scalable cloud data engineering solutions for HEDIS processing and healthcare quality reporting. Builds ETL pipelines, data integrations, SQL transformations, and AI-enabled workflow automation. Collaborates with stakeholders, auditors, and product teams to ensure data quality, regulatory compliance, and operational reliability. Troubleshoots production issues, supports cloud modernization, applies CI/CD and testing practices, and mentors engineers as a senior individual contributor.
Top Skills: Agentic AiAICi/CdDatabricksETLFhirGenaiAzureSQL

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account