Job Openings Remote | ML Infrastructure & Kernel Optimization Engineer — $65–$105/hour

About the job Remote | ML Infrastructure & Kernel Optimization Engineer — $65–$105/hour

We are sharing a specialised full-time consulting opportunity for US-based MLOps and ML systems engineers with production experience in JAX, PyTorch, distributed training infrastructure, and custom GPU kernel development using Pallas or Triton.

This role supports a high-impact generative AI initiative focused on developing and evaluating advanced ML infrastructure tasks for frontier model training. Selected engineers will design technically challenging problems, produce rigorous solutions, assess model-generated outputs, and help establish evaluation standards across training pipelines, distributed systems, framework-level optimisation, and GPU kernel performance.

Key Responsibilities

ML Infrastructure & Training Systems

  • Analyse and improve machine learning training infrastructure, deployment workflows, and model-development systems
  • Guide research and engineering teams on MLOps, distributed training, and ML framework-level challenges
  • Evaluate training-pipeline architecture, scalability, reliability, and performance
  • Identify technical gaps affecting model training, experimentation, and infrastructure efficiency

Technical Task & Solution Development

  • Design challenging tasks covering MLOps, ML systems, training infrastructure, and framework-level engineering
  • Write accurate, technically rigorous, and well-structured solutions
  • Develop realistic scenarios involving distributed systems, accelerator utilisation, and production ML workflows
  • Ensure tasks reflect practical engineering challenges encountered in advanced AI environments

JAX, PyTorch & GPU Kernel Evaluation

  • Evaluate technical work involving JAX and PyTorch at production scale
  • Review custom GPU kernels written or optimised using Pallas or Triton
  • Assess kernel correctness, memory access patterns, computational efficiency, and hardware utilisation
  • Analyse framework-level implementation choices and identify opportunities for performance improvement

Evaluation Frameworks & Technical Feedback

  • Compare alternative technical solutions and determine which approach is more accurate and effective
  • Provide clear written feedback on correctness, system design, scalability, and optimisation quality
  • Develop detailed rubrics for evaluating training pipelines, distributed systems reasoning, and kernel-level implementations
  • Collaborate with other technical specialists to maintain consistency across evaluation standards and training data

Ideal Profile

Strong candidates may have:

  • At least 2 years of dedicated professional experience in MLOps, ML infrastructure, or ML systems engineering
  • Production experience with JAX, PyTorch, or both at meaningful scale
  • Hands-on experience writing or optimising custom GPU kernels using Pallas or Triton
  • Strong knowledge of model-training pipelines, distributed systems, accelerators, and performance optimisation
  • Experience working within a recognised technology, AI research, or high-performance engineering organisation
  • Demonstrable professional growth and increasing technical responsibility
  • Strong written communication and the ability to explain complex engineering decisions clearly
  • Reliable availability for a full-time, 40-hour weekday schedule

Educational Background

  • A degree in computer science, machine learning, electrical engineering, applied mathematics, or a related technical field is highly relevant
  • Graduate-level education in machine learning systems, distributed computing, or high-performance computing may be helpful
  • Equivalent professional experience in production ML infrastructure may also be considered
  • Advanced technical work involving GPU programming, compiler systems, or large-scale model training is especially valuable

Nice to Have

  • Experience supporting large language model or generative AI training environments
  • Familiarity with distributed training frameworks, accelerator orchestration, and multi-host systems
  • Knowledge of XLA, CUDA, compiler optimisation, or low-level performance engineering
  • Experience benchmarking GPU workloads and diagnosing training-performance bottlenecks
  • Familiarity with model-evaluation pipelines, technical annotation, or structured training-data development
  • Previous involvement in technical review, engineering mentorship, or rubric development
  • Experience collaborating with research scientists and infrastructure engineering teams

Why This Opportunity

  • Contribute to advanced generative AI and large-scale model-training initiatives
  • Apply deep expertise in JAX, PyTorch, Pallas, Triton, and ML infrastructure
  • Work on challenging problems spanning training systems, distributed computing, and GPU optimisation
  • Influence the quality of technical training data used in frontier AI development
  • Join a full-time remote engagement with competitive hourly compensation

Contract Details

  • Full-time W-2 contingent employment arrangement
  • Fully remote role available to candidates based in the United States
  • Expected commitment of 40 hours per week during weekdays
  • This engagement requires full professional availability without conflicting employment or external commitments
  • Competitive rates between $65–$105 per hour depending on expertise and project scope
  • Immediate availability is preferred
  • Work may include onboarding, technical calibration, and ongoing quality-review activities
  • Project scope and duration may be adjusted according to programme requirements and performance

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.