Staff Engineer, Data Platform
2 дн. назад
200k–240k USD / yearCanadaWorldwideLeadOnsite
data engineeringdata platformdata normalizationdata quality validation
Lead engineering role responsible for building the data platform that transforms fragmented public sector data into structured intelligence.
Обязанности
- We’re looking for a Staff Engineer, Data Platform to own one of the most important technical problems at : turning the outside world’s fragmented government information into a proprietary data advantage.
- This is not a traditional data engineering role focused on maintaining a warehouse or internal analytics.
- You’ll own the technical ecosystem that:
- Discovers external data
- Acquires it reliably
- Understands and extracts information from it
- Normalizes and connects it
- Validates its quality
- Makes it available to ’s products and models
- The scope starts with more than 110,000 independent state and local government agencies, but extends to federal data, Canada, and eventually public-sector information globally.
- You’ll work across:
- Data engineering
- Distributed systems
- Information retrieval
- Data modeling
- LLMs and agents
- Applied ML
- Entity resolution
- Knowledge graphs
- You’ll partner closely with Product, ML Research, and Infrastructure to determine both:
- How we acquire data
- What data should have that nobody else does
Другое
- is building the data and intelligence layer for the public sector.
- More than 110,000 state and local government agencies across the U.S. independently publish information about: How they operate
- What they buy
- Who they work with
- What problems they are trying to solve
- That information is fragmented across millions of websites, documents, databases, procurement systems, meeting records, and public records.
- turns that information into structured, connected, actionable intelligence for businesses selling to government.
- Founded in 2024, is dedicated to making uncommon knowledge common , because public data should actually be public.
- Own our external data platform end-to-end Design systems spanning discovery, acquisition, extraction, normalization, entity resolution, validation, storage, serving, and monitoring.
- Establish the architecture and abstractions other engineers build on.
- Map the world of government data Develop a deep understanding of where government information lives.
- Understand how it is published, how it changes, and how information across thousands of institutions can be connected.
- Build systems for messy, real-world data Work across government websites, APIs, procurement systems, PDFs, spreadsheets, meeting records, and public records.
- Build for changing schemas, broken sources, conflicting records, and edge cases.
- Use AI to rethink the traditional data stack Work with our ML Research team to use LLMs, agents, and emerging models to: Discover new sources
- Understand unfamiliar schemas
- Extract structured information
- Resolve entities
- Monitor data quality
- Detect when sources change
- Build proprietary data flywheels Create systems where more data improves our models.
- Use better models to discover and understand more data.
- Continuously expand ’s underlying knowledge graph.
- Set technical direction Define the architecture for how acquires and represents public-sector information.
- Make decisions that will shape the platform over the next several years.
- Help determine which technical investments create the strongest long-term data advantage.
- You’re an unusually strong engineer who genuinely enjoys working with data.
- You’ve owned significant production data systems end-to-end.
- You enjoy the detective work of making sense of unfamiliar, messy datasets.
- You’re strong in Python, Go, or another systems/backend language.
- You’re highly proficient with SQL.
- You understand distributed data systems, including: Orchestration
- Idempotency
- Backfills
- Retries
- Observability
- Lineage
- Failure recovery
- You have experience with one or more of: Large-scale external data
- Crawling
- Information retrieval
- Entity resolution
- Knowledge graphs
- Document processing
- You’re excited about using LLMs and modern ML as components of data infrastructure.
- You care deeply about data quality, correctness, and reliability.
- You have strong product judgment and can reason about what data is actually worth acquiring , not just how to acquire it.
- You thrive in ambiguity and would rather create the architecture than be handed one.
- We’re particularly interested in backgrounds spanning:
- Alternative data
- Quantitative research infrastructure
- Search and crawling
- AI data infrastructure
- Knowledge graphs
- Large-scale document processing
- Data aggregation
- None of these are requirements.
- Backend: Python, Go, PostgreSQL
- Infrastructure: Redis, Docker, Kubernetes
- Frontend: React, TypeScript