Job Openings Machine Learning Platform Engineer

About the job Machine Learning Platform Engineer

Our client is an AI-native product company built to replace how billions of people manage their digital lives — starting with email, notes, and task tools that were never designed to be AI-native. It's backed by multi-million-dollar investment and built remote-first from day one, with a clear target: cut the time it takes users to get everyday things done by roughly 90%. That means solving the problems most AI products avoid — long-running workflows, persistent context, and reliable behavior under real-world, non-deterministic conditions. The team is small and high-talent-density by design, not by necessity, and moves at a pace that matches the scale of what it's building.

About the Role

As an ML Platform Engineer, you'll build the infrastructure and systems that power our client's AI capabilities — designing and operating everything behind the AI stack, from model training and evaluation to deployment, inference, observability, and continuous improvement. You'll work closely with AI engineers, researchers, and product engineers to turn models into reliable, scalable, cost-efficient production systems, and build the platforms and tooling that let the team experiment quickly and ship AI capabilities with confidence.

What You'll Work On


  • Build and operate the ML infrastructure and platforms powering the AI products

  • Design systems for model training, evaluation, deployment, inference, and experimentation

  • Build and optimize model serving and inference infrastructure for high-throughput, low-latency workloads

  • Improve reliability, scalability, latency, and cost efficiency of AI systems

  • Develop reliable pipelines for data preparation, training, evaluation, model release, and continuous improvement

  • Build platforms and tooling that let AI engineers and researchers experiment, evaluate, and ship models faster

  • Develop evaluation and benchmarking infrastructure to measure model quality, performance, and regressions

  • Build production observability, monitoring, tracing, and alerting for AI/ML workloads

  • Identify bottlenecks across the ML stack and continuously improve system performance

  • Work closely with AI engineers, researchers, and product teams to turn evolving model requirements into production-ready infrastructure


Requirements


  • Strong software engineering fundamentals and experience building production systems

  • Experience building ML infrastructure, platforms, or production machine learning systems

  • Experience with model deployment, inference, evaluation, or data pipelines

  • Strong understanding of distributed systems and system reliability

  • Ability to write clean, maintainable, production-quality code

  • Comfortable working in ambiguous, fast-moving environments

  • A bias toward ownership, experimentation, and continuous improvement


Tech Stack

Python · PyTorch / JAX · LLM/ML serving infrastructure (vLLM, SGLang, TensorRT-LLM) · cloud infrastructure · distributed systems · ML/data pipelines and workflow orchestration · GPU infrastructure and performance tooling · vector databases and retrieval infrastructure

What Success Looks Like


  • AI infrastructure reliably supports production workloads at scale

  • Models can be trained, evaluated, deployed, and improved efficiently

  • Inference systems deliver strong latency, throughput, reliability, and cost efficiency

  • ML pipelines are reproducible, observable, maintainable, and robust

  • Model and infrastructure regressions are detected quickly and diagnosed efficiently

  • Common ML infrastructure capabilities become reusable platform primitives rather than being rebuilt per product

  • The AI stack can evolve rapidly as new models, architectures, and inference techniques emerge


What to Expect

The best products in the world are built by small, world-class teams. Decisions are made collectively, at rapid speed — balancing high-quality shipping with fast learning. You'll be expected to bring structure, exercise judgment, and execute independently.

Compensation and benefits are competitive and vary by location; the package includes base salary and equity, discussed openly with you as part of the process.

If there's a fit, expect 3–4 interviews total, followed by a prompt, transparent decision.