Research Operations, Code
1 мес. назад
130k–250k USD / yearUSAMiddleOnsite
data programsai modelsresearch operations
Manage multi-million-dollar data programs for AI frontier model training as part of research operations.
Обязанности
- Silicon Valley's leading AI labs partner with to build the high-quality code and reasoning data that trains their frontier models. As a member of Research Ops on the Code team, you will own multi-million-dollar data programs end to end — and ownership here is technical as much as operational. You will:
- Decompose frontier model capabilities. Reason about where a lab's frontier coding model is weak, and translate that into the data that will make it stronger.
- Design the pipeline. Shape how data gets produced — human-expert workflows, synthetic generation, model-in-the-loop systems — not just run a fixed process.
- Drive execution. Turn that design into delivery through a team of expert contributors, under demanding timelines, at a consistently high quality bar.
- Own the customer. Build deep relationships with lab researchers and become the person they trust to tell them what data they actually need.
- This role sits at the intersection of operator and builder. You'll spend as much time reasoning about designing tasks which fail frontier models in fair ways , designing verifiers with parity to production codebases , and constructing self-contained environments which capture real world tasks, as you will on delivery and quality. We operate with startup intensity: occasionally responsive on weekends, always a high bar. The upside is that performance incentives, high-slope career trajectory, and meaningful equity reflects that intensity.
Требования
- We hire for range: people who can hold a credible technical conversation with a lab researcher in the morning and design the expert workflow that produces the data in the afternoon.
Будет плюсом
- Pipeline building: designing data or automation pipelines; hands-on with LLMs/agents; building synthetic-data or model-in-the-loop systems.
- Research fluency: connecting model/benchmark literature to what data would move a frontier model; having created a benchmark or published analysis of model behavior.
- Backgrounds that often fit: ML/data/software engineers who love operating, technical PMs, research engineers, or strong generalist operators with real technical range. Backgrounds from consulting, finance, or high-growth startups can work if paired with real technical fluency .
Условия
- Bi-annual performance bonus structure
- Generous equity grant vested over 4 years
- Up to $15k Relocation bonus
- $10K housing bonus (if you live within 0.5 miles of our office)
- $1.5K monthly stipend for meals
- Free Equinox membership
- $200 monthly laundry reimbursement
- $200 monthly personal wellness reimbursement
- Health, Dental, Vision insurance
Другое
- 's mission is to organize human intelligence to power the AI economy. We're a leading AI data company, building the layer between human expertise and frontier models. Millions of domain experts on the platform are paid over $4 million per day to train frontier AI models. 's APEX benchmark family measures AI's real-world impact on professional work. Enterprise brings this same infrastructure to Fortune 500 companies: helping companies capture how their best people actually work, translating that expertise directly back into agents.
- is creating a new category of work where expertise powers AI advancement. Achieving this requires an ambitious, fast-paced and deeply committed team. You’ll work alongside researchers, operators, and AI companies at the forefront of shaping the systems that are redefining society. is a profitable Series C company valued at $10 billion. We work in-person five days a week in our San Francisco, NYC, or London offices.
- Develop a working understanding of what our customers' models can and can't do, and where the capability gaps are.
- Evaluate model outputs and benchmark performance, and create targeted loss analysis to identify what data will drive improvement.
- Translate research goals into concrete task designs, difficulty targets, and quality specifications — including novel frontier tasks that don't exist anywhere else.
- Design how data is produced — human-expert, synthetic, and hybrid model-in-the-loop pipelines — and continuously improve them.
- Prototype new generation and validation approaches; bring ideas for new data products, not just improvements to existing ones.
- Balance quality, throughput, and cost as you scale a pipeline from prototype to production.
- Manage end-to-end data pipelines from customer specification to final delivery.
- Diagnose bottlenecks, restructure workflows, and implement solutions — incentive systems, workflow re-sequencing, sharper instructions, scaled review processes, and automated quality assurance.
- Run daily internal syncs ("war rooms") to stay ahead of issues.
- Act as the primary point of contact for leading AI labs; deliver clear, consistent reporting.
- Proactively anticipate researcher needs and identify opportunities for expansion.
- Design the human-expert labeling and evaluation flows on the platform — how tasks are structured, reviewed, and scored.
- Source, vet, train, and performance-manage teams of domain experts (software engineers, competitive programmers, and specialists).
- Maintain a high execution and quality standard across every stage of production.
- Technical judgment. Coding literacy and genuine ML or Model Benchmark familiarity — enough to evaluate model outputs, reason about frontier-model capabilities and failure modes, read benchmark/eval work, and translate research goals into concrete task and data designs. You don't need to be a research scientist, but you do need to go deep on the technical substance.
- Operational ownership. A track record of running complex, high-stakes projects end to end, and genuine energy for large-scale execution and gritty process optimization under pressure.
- Communication & customer instinct. Strong analytical and communication skills; comfortable owning high-profile relationships with technical customers.
- Research Ops team members generate $10M+ in project revenue annually.
- Customer requests are high-pressure and time-sensitive (e.g., standing up a vetted expert team and delivering high-quality data against a lab's spec within days).
- Day-to-day is roughly ~30% customer engagement, ~40–50% data production (pipelines, experts, scaling), and ~15% operations and metrics review.
- This role blends operational ownership with real technical judgment about models and data. It is not a pure project-management seat, nor a pure research seat.
- Location: San Francisco (in person, five days a week).