Software Engineer (Model Inference)
1 мес. назад
400k–800k USD / yearAustraliaWorldwideSeniorOnsite
cudagpumlmodel inference
Responsible for owning and optimizing the model inference stack, serving AI models with high performance and low latency to millions of users.
О компании
- At , we're redefining entertainment with AI. We're one of the fastest-growing startups globally, serving 20 million users in our first year . We're an in-person company based in Sydney, Australia.
- We're guided by a set of principles that are foundational to our work and drive every decision: User-first: We build things that people want. We invest time to understand our users and focus on adding value instead of extracting value.
- High agency, high ownership: We're responsible for the pieces we own, end-to-end. We own every mistake, figure out what went wrong, and fix it. We don't blame anyone or anything else.
- Urgency: This is a once-in-a-lifetime opportunity. We prioritize well, find ways to increase leverage, and move at an inspirational pace.
Обязанности
- You'll own our model inference stack end-to-end — serving our in-house and open-source models to tens of millions of users at low latency and high throughput, and squeezing every bit of performance out of our GPU fleet.
Условия
- Competitive compensation with meaningful upside.
- A company card for food, coffee, tools, and anything that helps you operate at a high level.
- Daily team lunch and dinner at the office.
- Unlimited workspace budget. Build your ideal setup.
- Real ownership and impact from day one.
- Based in Sydney and hiring globally. We sponsor visas and offer relocation assistance to help you make the move.
Как откликнуться
- A quick phone call to learn about you and share more about our company.
- A technical interview, 1hr of system design, no live coding.
- A paid work trial, so you get a feel for what it's like to work with us.
- You get an offer, and we celebrate!
Другое
- Ship a high-throughput inference server for our in-house models, serving millions of generations per day at <200ms latency.
- Squeeze more out of every GPU with batching, quantization, and custom CUDA kernels .
- Test and productionize LoRAs and our in-house models to serve 10s of millions of users.
- 5+ years of experience building software at scale, with a focus on ML inference or GPU-accelerated systems .
- Deep familiarity with GPU inference : batching, quantization, and serving frameworks like vLLM, TensorRT, or Triton.
- The ability to get shit done end-to-end . From understanding users → proposing an idea → implementation → iteration.
- Hunger to win. This is not going to be easy.
- Our overarching philosophy is to raise the ceiling for our best performers, not the floor. We pay based on the value you add to the company.
- The range listed for this role is total comp: cash plus equity combined, with superannuation paid on top.
- We re-evaluate compensation every 6 months. Bonuses are backward-looking; raises are forward-looking.
- A bonus rewards work you've already done, like exceptional output or consistently performing above your level.
- A raise reflects that you've moved up. It resets your baseline to match the higher level you're now operating at.