Granica Logo

Granica

Senior Software Engineer — Lakehouse Systems

Posted 16 Days Ago
Be an Early Applicant
In-Office
Mountain View, CA
160K-240K Annually
Senior level
In-Office
Mountain View, CA
160K-240K Annually
Senior level
Build foundational lakehouse infrastructure for exabyte-scale AI data environments. Responsibilities include metadata and transaction systems, table maintenance, schema and partition evolution, snapshot isolation, compaction, clustering, file-layout optimization, object-store performance, columnar-format optimization, and query performance across major lakehouse engines. The role also involves debugging distributed systems, implementing compression and data-efficiency algorithms, and contributing to open-source or research efforts.
The summary above was generated by AI
Senior Software Engineer — Lakehouse Systems

Location: Mountain View, CA — On-site

 
About Granica

Granica builds AI infrastructure for enterprises operating massive data environments.

Our platform helps data and engineering teams reduce storage and compute costs, improve performance and reliability, and prepare large datasets for analytics and AI.

 

Granica’s products include:

  • Crunch — continuous optimization for enterprise lakehouse data

  • Myelin — stateful infrastructure for long-running AI agents

  • Large Tabular Models — foundation models designed for enterprise tables

 

Together, we are building the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently.

 

Granica has demonstrated approximately $200K in annualized value per petabyte and verified customer value within weeks.

 
About the Role

Granica is hiring a Senior Software Engineer to build foundational lakehouse systems for AI.

You will work on the core infrastructure behind Crunch, Granica’s continuous optimization product for enterprise lakehouse data. This includes systems for metadata management, transaction semantics, table maintenance, object-store-backed storage layouts, file-level optimization, and lakehouse cost/performance across petabyte- and exabyte-scale environments.

You will own core systems that directly affect customer infrastructure cost, query performance, table reliability, and the operational health of large lakehouse environments.

This is a hands-on engineering role for someone who has deep systems experience and wants to build at the intersection of data lakes, table formats, metadata systems, storage layout, query performance, and AI infrastructure.

You will work on lakehouse systems involving Apache Iceberg, Delta Lake, Apache Hudi, Parquet, ORC, cloud object stores, and query engines such as Spark, Trino, Presto, Flink, Databricks, and Snowflake-adjacent environments.

 
What You’ll Do
  • Build metadata and transaction systems for large-scale tabular datasets

  • Design systems that support time travel, schema evolution, partition evolution, snapshot isolation, and atomic consistency

  • Develop table-maintenance infrastructure for lakehouse formats such as Apache Iceberg, Delta Lake, and Apache Hudi

  • Build systems for manifests, snapshots, transaction logs, metadata pruning, snapshot expiration, table garbage collection, and catalog consistency

  • Optimize file layout, clustering, compaction, file sizing, data skipping, indexing, and read-path performance

  • Improve performance and cost efficiency across object-store-backed lakehouse environments such as S3, GCS, and ADLS

  • Work with columnar formats such as Parquet and ORC, including encoding, compression, layout, pruning, and read-path optimization

  • Build systems that make lakehouse tables faster, cheaper, and more reliable across engines and platforms such as Spark, Flink, Trino, Presto, Databricks, and Snowflake-adjacent environments

  • Debug performance bottlenecks across storage, metadata, table maintenance, query execution, network, and compute layers

  • Develop workload-aware table optimization systems that learn from access patterns and reorganize data automatically

  • Implement algorithms in compression, representation, layout optimization, and data efficiency

  • Contribute to open-source or publish research when appropriate

 
What We’re Looking For
  • Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure

  • Production experience with modern data lake or lakehouse technologies such as Iceberg, Delta Lake, Hudi, Spark, Trino, Presto, Flink, Hive Metastore, Unity Catalog, or similar systems

  • Hands-on experience with columnar formats such as Parquet or ORC

  • Understanding of metadata-driven architectures, table formats, transaction semantics, query planning, and physical data layout

  • Experience with table maintenance, compaction, clustering, file sizing, metadata pruning, snapshot expiration, or garbage collection

  • Familiarity with cloud object storage systems such as S3, GCS, or ADLS and the performance tradeoffs of building lakehouse systems on top of them

  • Strong programming skills in Java, Scala, Go, Rust, C++, or similar systems-oriented languages

  • Curiosity about compression, entropy, information theory, and how data representation affects AI efficiency

  • A pragmatic builder’s mindset: rigorous, hands-on, and comfortable owning complex systems end to end

 
Bonus
  • Experience contributing to Apache Iceberg, Delta Lake, Apache Hudi, Spark, Flink, Trino, Presto, Velox, DuckDB, Polars, Parquet, ORC, or related systems

  • Experience with manifests, snapshots, metadata catalogs, schema evolution, partition evolution, delete handling, transaction logs, or table garbage collection

  • Experience solving the small-file problem, optimizing object-store access patterns, or improving table health at scale

  • Background in storage engines, query engines, indexing, caching, encoding, compression, or adaptive query optimization

  • Research or open-source contributions in distributed systems, databases, storage, compression, indexing, or data processing

  • Interest in how physical data representation affects model training, inference, retrieval, and reasoning efficiency

 
Why Join Granica
  • Build foundational infrastructure for enterprise data and AI

  • Work on deep systems problems across lakehouse metadata, transaction semantics, table maintenance, storage layout, object-store behavior, query performance, and AI efficiency

  • Partner directly with Product, Engineering, and company leadership

  • Help shape Crunch, Granica’s production data optimization platform for enterprise-scale lakehouse environments

  • Work with a small, high-caliber team solving high-value infrastructure problems at massive scale

  • Have direct influence on architecture, product direction, customer outcomes, and company growth

Compensation & Benefits
  • Competitive salary, meaningful equity, and performance bonus for top performers

  • 401(k) with company match, comprehensive health coverage, and unlimited PTO

  • Daily catered meals in our Mountain View office

  • Support for research, publication, and conference participation

At Granica, you'll help build the next generation of enterprise AI—from exabyte-scale data infrastructure, Large Tabular Models (LTMs), and stateful AI agents. Together, we're creating the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently.

 

Similar Jobs at Granica

16 Days Ago
Hybrid
160K-240K Annually
Senior level
160K-240K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
Build and optimize distributed compute infrastructure for large-scale analytical and AI workloads. Responsibilities include improving query execution, scheduling, resource allocation, reliability, workload routing, and compute efficiency across Spark and related systems. The role involves debugging performance bottlenecks, optimizing joins, scans, shuffles, caching, partitioning, and memory usage, and working with lakehouse formats and cloud object storage. Candidates will implement workload optimization algorithms and may contribute to open source or research.
Top Skills: Adaptive Query ExecutionAmazon EmrAmazon S3Apache FlinkApache HiveApache HudiApache IcebergSparkAws GlueAzure Data Lake StorageC++CatalystDatabricksDatafusionDelta LakeDuckdbGoGoogle Cloud StorageJavaOrcParquetPrestoRustScalaSnowflakeSpark SqlTrinoVelox
20 Days Ago
In-Office
140K-180K Annually
Senior level
140K-180K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
Lead finance strategy including capital raising, forecasting, performance modeling, and risk management to support company’s growth and IPO readiness.
Top Skills: Ai InfrastructureFinancial Modeling
22 Days Ago
In-Office
160K-240K Annually
Senior level
160K-240K Annually
Senior level
Artificial Intelligence • Big Data • Cloud • Machine Learning • Software • Business Intelligence • Data Privacy
Develop and prototype novel diffusion and generative learning algorithms for Large Tabular Models. Design efficient training methods, representation learning techniques, benchmarks, and collaborate with academic and engineering teams to translate research into production.
Top Skills: Diffusion ModelsJaxLarge Tabular ModelsProbabilistic ModelingPythonPyTorchRepresentation LearningScalable Ml SystemsScore-Based Generative Modeling

What you need to know about the Los Angeles Tech Scene

Los Angeles is a global leader in entertainment, so it’s no surprise that many of the biggest players in streaming, digital media and game development call the city home. But the city boasts plenty of non-entertainment innovation as well, with tech companies spanning verticals like AI, fintech, e-commerce and biotech. With major universities like Caltech, UCLA, USC and the nearby UC Irvine, the city has a steady supply of top-flight tech and engineering talent — not counting the graduates flocking to Los Angeles from across the world to enjoy its beaches, culture and year-round temperate climate.

Key Facts About Los Angeles Tech

  • Number of Tech Workers: 375,800; 5.5% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: Snap, Netflix, SpaceX, Disney, Google
  • Key Industries: Artificial intelligence, adtech, media, software, game development
  • Funding Landscape: $11.6 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Strong Ventures, Fifth Wall, Upfront Ventures, Mucker Capital, Kittyhawk Ventures
  • Research Centers and Universities: California Institute of Technology, UCLA, University of Southern California, UC Irvine, Pepperdine, California Institute for Immunology and Immunotherapy, Center for Quantum Science and Engineering

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account