Job Openings Remote | AWS Trainium Kernel Engineer (NKI) — $60–$80/hour

About the job Remote | AWS Trainium Kernel Engineer (NKI) — $60–$80/hour

We are sharing a specialised part-time consulting opportunity for experienced kernel engineers with hands-on expertise in the Neuron Kernel Interface (NKI), AWS Trainium/Inferentia2 hardware, low-level performance optimisation, and CUDA-to-NKI migration.

This role focuses on evaluating NKI kernel-development tasks for technical correctness, hardware appropriateness, numerical fidelity, and performance quality. Selected experts will review Trainium-specific implementations, migration decisions, memory-management strategies, profiling results, and cross-platform numerical behaviour while providing clear, rubric-based technical feedback.

Key Responsibilities

NKI Kernel Development Review

  • Evaluate kernels developed using the Neuron Kernel Interface (NKI)
  • Assess whether implementations appropriately target AWS Trainium and Inferentia2 hardware
  • Review low-level computation patterns for technical correctness
  • Identify inefficient, incorrect, or hardware-inappropriate implementation choices
  • Apply practical judgement grounded in hands-on NKI development experience

CUDA-to-NKI Migration

  • Review migrations of existing CUDA kernels to NKI
  • Assess whether computational semantics are preserved across platforms
  • Identify translation errors, unsupported assumptions, or inefficient migration strategies
  • Evaluate whether NKI implementations appropriately account for Trainium architecture
  • Distinguish faithful migrations from implementations that merely reproduce surface-level CUDA structure

Tile-Based Computation

  • Assess tile decomposition and computation strategies
  • Review partitioning decisions against NKI execution constraints
  • Evaluate whether kernels make effective use of available compute resources
  • Identify inefficient tiling or data-movement patterns
  • Assess whether implementation choices align with NKI programming requirements

Memory Hierarchy Management

  • Review use of SBUF, PSUM, and HBM
  • Assess data placement and movement across Trainium memory hierarchies
  • Evaluate memory-bandwidth utilisation and locality
  • Identify unnecessary transfers or memory bottlenecks
  • Review implementation decisions affecting on-chip memory efficiency

DMA & Data Movement

  • Evaluate DMA orchestration within NKI kernels
  • Review sequencing of computation and data-transfer operations
  • Identify stalls, inefficient transfer patterns, or synchronisation issues
  • Assess whether data movement appropriately overlaps with computation
  • Evaluate implementation choices affecting pipeline utilisation

Trainium Performance Optimisation

  • Review Trainium-specific optimisation strategies
  • Assess NeuronCore pipeline utilisation, tensor-engine throughput, and memory-bandwidth behaviour
  • Identify performance bottlenecks within kernel implementations
  • Evaluate whether optimisation decisions are supported by profiling evidence
  • Review trade-offs affecting throughput, latency, and resource utilisation

Numerical Correctness

  • Evaluate numerical consistency between GPU and Trainium implementations
  • Review differences caused by accumulation order, rounding behaviour, and mixed-precision semantics
  • Assess appropriate tolerances for cross-platform comparisons
  • Identify numerical discrepancies that indicate implementation defects
  • Distinguish expected hardware-level variation from substantive correctness problems

Precision & Data Types

  • Review kernels using supported formats such as FP32, BF16, FP8, and INT8
  • Assess precision choices against computational requirements
  • Evaluate mixed-precision behaviour and numerical stability
  • Identify inappropriate casting or accumulation strategies
  • Review whether performance gains are achieved without compromising required correctness

AWS Neuron Ecosystem

  • Evaluate implementations using the AWS Neuron SDK
  • Review interactions between kernel code, compilation, and Trainium execution
  • Assess compiler-related behaviours where relevant
  • Apply familiarity with NKI kernel libraries and Neuron tooling
  • Identify implementation issues arising from platform-specific constraints

Benchmarking & Validation

  • Review benchmark results for Trainium workloads
  • Assess performance comparisons and experimental methodology
  • Evaluate workloads running on Trn1 or Trn2 instances where applicable
  • Determine whether claimed performance improvements are supported by evidence
  • Identify benchmarking methodologies that could produce misleading conclusions

Rubric-Based Technical Evaluation

  • Assess assigned kernel-development tasks against structured technical criteria
  • Provide clear written explanations supporting evaluation decisions
  • Reference specific implementation, profiling, or numerical evidence
  • Apply evaluation standards consistently across assignments
  • Distinguish valid optimisation alternatives from technically flawed approaches

Ideal Profile

  • 2+ years of hands-on experience developing or optimising kernels using the Neuron Kernel Interface (NKI)
  • Professional experience targeting AWS Trainium or Inferentia2 hardware
  • Strong understanding of tile-based computation
  • Deep familiarity with SBUF, PSUM, and HBM memory management
  • Strong knowledge of partition-dimension constraints and DMA orchestration
  • Demonstrated experience evaluating or performing CUDA-to-NKI migrations
  • Familiarity with Trainium-specific performance profiling
  • Experience assessing NeuronCore pipeline utilisation, tensor-engine throughput, and memory-bandwidth bottlenecks
  • Strong understanding of cross-platform numerical correctness and mixed-precision behaviour
  • Direct experience with the AWS Neuron SDK, Neuron Compiler internals, or NKI kernel libraries is preferred
  • Prior CUDA or Triton kernel development experience is advantageous
  • Familiarity with NeuronCore-v2 architecture and supported numerical formats is preferred
  • Experience benchmarking ML workloads on Trn1 or Trn2 instances is advantageous
  • Strong written communication and ability to provide precise technical feedback

Engagement Details

  • Part-time independent contractor engagement
  • Fully remote within the United States
  • Flexible scheduling based on project requirements
  • Compensation: $60–$80/hour
  • Work focuses on NKI kernel development, Trainium performance optimisation, CUDA migration, numerical correctness, and technical quality evaluation
  • Projects may be extended, shortened, or concluded based on project needs and performance
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
  • H1-B and STEM OPT support is unavailable for this engagement

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.