Senior Software Engineer — Lakehouse Systems
3 дн. назад
160k–240k USD / yearUSASeniorOnsite
metadata managementtransaction semanticsfile optimizationobject storage
Senior Software Engineer to develop core infrastructure systems for enterprise lakehouse AI products.
Обязанности
- is hiring a Senior Software Engineer to build foundational lakehouse systems for AI.
- You will work on the core infrastructure behind Crunch , ’s continuous optimization product for enterprise lakehouse data. This includes systems for metadata management, transaction semantics, table maintenance, object-store-backed storage layouts, file-level optimization, and lakehouse cost/performance across petabyte- and exabyte-scale environments.
- You will own core systems that directly affect customer infrastructure cost, query performance, table reliability, and the operational health of large lakehouse environments.
- This is a hands-on engineering role for someone who has deep systems experience and wants to build at the intersection of data lakes, table formats, metadata systems, storage layout, query performance, and AI infrastructure.
- You will work on lakehouse systems involving Apache Iceberg, Delta Lake, Apache Hudi, Parquet, ORC, cloud object stores, and query engines such as Spark, Trino, Presto, Flink, Databricks, and Snowflake-adjacent environments.
Будет плюсом
- Experience contributing to Apache Iceberg, Delta Lake, Apache Hudi, Spark, Flink, Trino, Presto, Velox, DuckDB, Polars, Parquet, ORC, or related systems
- Experience with manifests, snapshots, metadata catalogs, schema evolution, partition evolution, delete handling, transaction logs, or table garbage collection
- Experience solving the small-file problem, optimizing object-store access patterns, or improving table health at scale
- Background in storage engines, query engines, indexing, caching, encoding, compression, or adaptive query optimization
- Research or open-source contributions in distributed systems, databases, storage, compression, indexing, or data processing
- Interest in how physical data representation affects model training, inference, retrieval, and reasoning efficiency
Условия
- Build foundational infrastructure for enterprise data and AI
- Work on deep systems problems across lakehouse metadata, transaction semantics, table maintenance, storage layout, object-store behavior, query performance, and AI efficiency
- Partner directly with Product, Engineering, and company leadership
- Help shape Crunch, ’s production data optimization platform for enterprise-scale lakehouse environments
- Work with a small, high-caliber team solving high-value infrastructure problems at massive scale
- Have direct influence on architecture, product direction, customer outcomes, and company growth
- Competitive salary, meaningful equity, and performance bonus for top performers
- 401(k) with company match, comprehensive health coverage, and unlimited PTO
- Daily catered meals in our Mountain View office
- Support for research, publication, and conference participation
- At , you'll help build the next generation of enterprise AI —from exabyte-scale data infrastructure , Large Tabular Models (LTMs) , and stateful AI agents . Together, we're creating the infrastructure that enables enterprises to own their data , own the intelligence built on it , and scale both efficiently .
Другое
- builds AI infrastructure for enterprises operating massive data environments.
- Our platform helps data and engineering teams reduce storage and compute costs, improve performance and reliability, and prepare large datasets for analytics and AI.
- ’s products include:
- Crunch — continuous optimization for enterprise lakehouse data
- Myelin — stateful infrastructure for long-running AI agents
- Large Tabular Models — foundation models designed for enterprise tables
- Together, we are building the infrastructure that enables enterprises to own their data, own the intelligence built on it, and scale both efficiently.
- has demonstrated approximately $200K in annualized value per petabyte and verified customer value within weeks.
- Build metadata and transaction systems for large-scale tabular datasets
- Design systems that support time travel, schema evolution, partition evolution, snapshot isolation, and atomic consistency
- Develop table-maintenance infrastructure for lakehouse formats such as Apache Iceberg, Delta Lake, and Apache Hudi
- Build systems for manifests, snapshots, transaction logs, metadata pruning, snapshot expiration, table garbage collection, and catalog consistency
- Optimize file layout, clustering, compaction, file sizing, data skipping, indexing, and read-path performance
- Improve performance and cost efficiency across object-store-backed lakehouse environments such as S3, GCS, and ADLS
- Work with columnar formats such as Parquet and ORC, including encoding, compression, layout, pruning, and read-path optimization
- Build systems that make lakehouse tables faster, cheaper, and more reliable across engines and platforms such as Spark, Flink, Trino, Presto, Databricks, and Snowflake-adjacent environments
- Debug performance bottlenecks across storage, metadata, table maintenance, query execution, network, and compute layers
- Develop workload-aware table optimization systems that learn from access patterns and reorganize data automatically
- Implement algorithms in compression, representation, layout optimization, and data efficiency
- Contribute to open-source or publish research when appropriate
- Strong engineering depth in distributed systems, storage systems, databases, or data infrastructure
- Production experience with modern data lake or lakehouse technologies such as Iceberg, Delta Lake, Hudi, Spark, Trino, Presto, Flink, Hive Metastore, Unity Catalog, or similar systems
- Hands-on experience with columnar formats such as Parquet or ORC
- Understanding of metadata-driven architectures, table formats, transaction semantics, query planning, and physical data layout
- Experience with table maintenance, compaction, clustering, file sizing, metadata pruning, snapshot expiration, or garbage collection
- Familiarity with cloud object storage systems such as S3, GCS, or ADLS and the performance tradeoffs of building lakehouse systems on top of them
- Strong programming skills in Java, Scala, Go, Rust, C++, or similar systems-oriented languages
- Curiosity about compression, entropy, information theory, and how data representation affects AI efficiency
- A pragmatic builder’s mindset: rigorous, hands-on, and comfortable owning complex systems end to end