About the job Machine Learning Platform Engineer
Our client is an AI-native product company built to replace how billions of people manage their digital lives — starting with email, notes, and task tools that were never designed to be AI-native. It's backed by multi-million-dollar investment and built remote-first from day one, with a clear target: cut the time it takes users to get everyday things done by roughly 90%. That means solving the problems most AI products avoid — long-running workflows, persistent context, and reliable behavior under real-world, non-deterministic conditions. The team is small and high-talent-density by design, not by necessity, and moves at a pace that matches the scale of what it's building.
About the Role
As an ML Platform Engineer, you'll build the infrastructure and systems that power our client's AI capabilities — designing and operating everything behind the AI stack, from model training and evaluation to deployment, inference, observability, and continuous improvement. You'll work closely with AI engineers, researchers, and product engineers to turn models into reliable, scalable, cost-efficient production systems, and build the platforms and tooling that let the team experiment quickly and ship AI capabilities with confidence.
What You'll Work On
Build and operate the ML infrastructure and platforms powering the AI products
Design systems for model training, evaluation, deployment, inference, and experimentation
Build and optimize model serving and inference infrastructure for high-throughput, low-latency workloads
Improve reliability, scalability, latency, and cost efficiency of AI systems
Develop reliable pipelines for data preparation, training, evaluation, model release, and continuous improvement
Build platforms and tooling that let AI engineers and researchers experiment, evaluate, and ship models faster
Develop evaluation and benchmarking infrastructure to measure model quality, performance, and regressions
Build production observability, monitoring, tracing, and alerting for AI/ML workloads
Identify bottlenecks across the ML stack and continuously improve system performance
Work closely with AI engineers, researchers, and product teams to turn evolving model requirements into production-ready infrastructure
Requirements
Strong software engineering fundamentals and experience building production systems
Experience building ML infrastructure, platforms, or production machine learning systems
Experience with model deployment, inference, evaluation, or data pipelines
Strong understanding of distributed systems and system reliability
Ability to write clean, maintainable, production-quality code
Comfortable working in ambiguous, fast-moving environments
A bias toward ownership, experimentation, and continuous improvement
Tech Stack
Python · PyTorch / JAX · LLM/ML serving infrastructure (vLLM, SGLang, TensorRT-LLM) · cloud infrastructure · distributed systems · ML/data pipelines and workflow orchestration · GPU infrastructure and performance tooling · vector databases and retrieval infrastructure
What Success Looks Like
AI infrastructure reliably supports production workloads at scale
Models can be trained, evaluated, deployed, and improved efficiently
Inference systems deliver strong latency, throughput, reliability, and cost efficiency
ML pipelines are reproducible, observable, maintainable, and robust
Model and infrastructure regressions are detected quickly and diagnosed efficiently
Common ML infrastructure capabilities become reusable platform primitives rather than being rebuilt per product
The AI stack can evolve rapidly as new models, architectures, and inference techniques emerge
What to Expect
The best products in the world are built by small, world-class teams. Decisions are made collectively, at rapid speed — balancing high-quality shipping with fast learning. You'll be expected to bring structure, exercise judgment, and execute independently.
Compensation and benefits are competitive and vary by location; the package includes base salary and equity, discussed openly with you as part of the process.
If there's a fit, expect 3–4 interviews total, followed by a prompt, transparent decision.