Staff Software Engineer
2 г. назад
USALeadHybrid
distributed systemsmetadata managementcachingreplicationconcurrent programmingfault tolerancemulti-cloud
Experienced distributed-systems engineer working on 's data-orchestration engine for AI and analytics at scale.
О компании
- About powers the data layer for modern AI and analytics. Proven in production at eight of the top ten internet companies and seven of the ten highest-valued enterprises globally, ’s data orchestration platform unifies data across storage systems, regions, and clouds providing a high-performance distributed caching layer built for large-scale AI workloads. Spun out of UC Berkeley’s AMPLab by the creators of Tachyon and backed by Andreessen Horowitz, Hillhouse Capital, and Seven Seas Partners, sits at the intersection of data, distributed systems, and AI infrastructure . Our technology is deployed at scale by organizations such as Meta, Uber, Tencent, TikTok, Alibaba, Expedia, Rakuten, Microsoft, and Walmart , orchestrating data for billions of operations per d
Другое
- Cache and metadata consistency - advance ’s intelligent caching framework for multi-tenant environments (TTL policies, write-back consistency, invalidation protocols, and distributed metadata scaling).
- High-throughput data I/O optimization - profile and optimize ’s data path across S3, GCS, HDFS, and POSIX interfaces using adaptive prefetching, async I/O, and tier-aware scheduling.
- Scaling for AI and analytics workloads - evolve the coordination layer to efficiently serve distributed AI training clusters, accelerating model load and shuffle operations across regions and clouds.
- Observability and performance insights - build fine-grained metrics and tracing for cache efficiency, throughput, and latency across storage tiers.
- Open-source leadership - drive design discussions, mentor contributors, and represent ’s core-systems direction within the OSS community.
- Design and implement core components of ’s distributed file and object-access layer.
- Optimize performance for large-scale, high-throughput environments using advanced concurrency and caching techniques.
- Build scalable metadata and coordination systems that ensure strong consistency, high availability, and minimal latency.
- Collaborate cross-functionally with product, solution-engineering, and research teams to drive roadmap and customer success.
- Strong computer-science fundamentals and a passion for large-scale distributed systems.
- Professional experience developing in Java, C++, or Go .
- Deep understanding of concurrency, replication, fault tolerance, and performance optimization .
- Experience with distributed storage, data-access layers, or cloud infrastructure (e.g., Spark, Presto, Hadoop, Kubernetes).
- Bachelor’s or advanced degree in Computer Science or related technical field (or equivalent experience).
- Demonstrated technical leadership: defining architecture, mentoring peers, or driving major projects from design through release.
- Build infrastructure trusted by the world’s largest AI and data-driven companies.
- Join a small, senior engineering team where your designs shape the product’s evolution.
- Work directly with the original creators of open-source .
- A culture of empathy, curiosity, and ownership - where engineers collaborate closely to solve hard problems.