About the job Remote | AWS Trainium Kernel Engineer (NKI) — $60–$80/hour
We are sharing a specialised part-time consulting opportunity for experienced kernel engineers with hands-on expertise in the Neuron Kernel Interface (NKI), AWS Trainium/Inferentia2 hardware, low-level performance optimisation, and CUDA-to-NKI migration.
This role focuses on evaluating NKI kernel-development tasks for technical correctness, hardware appropriateness, numerical fidelity, and performance quality. Selected experts will review Trainium-specific implementations, migration decisions, memory-management strategies, profiling results, and cross-platform numerical behaviour while providing clear, rubric-based technical feedback.
Key Responsibilities
NKI Kernel Development Review
- Evaluate kernels developed using the Neuron Kernel Interface (NKI)
- Assess whether implementations appropriately target AWS Trainium and Inferentia2 hardware
- Review low-level computation patterns for technical correctness
- Identify inefficient, incorrect, or hardware-inappropriate implementation choices
- Apply practical judgement grounded in hands-on NKI development experience
CUDA-to-NKI Migration
- Review migrations of existing CUDA kernels to NKI
- Assess whether computational semantics are preserved across platforms
- Identify translation errors, unsupported assumptions, or inefficient migration strategies
- Evaluate whether NKI implementations appropriately account for Trainium architecture
- Distinguish faithful migrations from implementations that merely reproduce surface-level CUDA structure
Tile-Based Computation
- Assess tile decomposition and computation strategies
- Review partitioning decisions against NKI execution constraints
- Evaluate whether kernels make effective use of available compute resources
- Identify inefficient tiling or data-movement patterns
- Assess whether implementation choices align with NKI programming requirements
Memory Hierarchy Management
- Review use of SBUF, PSUM, and HBM
- Assess data placement and movement across Trainium memory hierarchies
- Evaluate memory-bandwidth utilisation and locality
- Identify unnecessary transfers or memory bottlenecks
- Review implementation decisions affecting on-chip memory efficiency
DMA & Data Movement
- Evaluate DMA orchestration within NKI kernels
- Review sequencing of computation and data-transfer operations
- Identify stalls, inefficient transfer patterns, or synchronisation issues
- Assess whether data movement appropriately overlaps with computation
- Evaluate implementation choices affecting pipeline utilisation
Trainium Performance Optimisation
- Review Trainium-specific optimisation strategies
- Assess NeuronCore pipeline utilisation, tensor-engine throughput, and memory-bandwidth behaviour
- Identify performance bottlenecks within kernel implementations
- Evaluate whether optimisation decisions are supported by profiling evidence
- Review trade-offs affecting throughput, latency, and resource utilisation
Numerical Correctness
- Evaluate numerical consistency between GPU and Trainium implementations
- Review differences caused by accumulation order, rounding behaviour, and mixed-precision semantics
- Assess appropriate tolerances for cross-platform comparisons
- Identify numerical discrepancies that indicate implementation defects
- Distinguish expected hardware-level variation from substantive correctness problems
Precision & Data Types
- Review kernels using supported formats such as FP32, BF16, FP8, and INT8
- Assess precision choices against computational requirements
- Evaluate mixed-precision behaviour and numerical stability
- Identify inappropriate casting or accumulation strategies
- Review whether performance gains are achieved without compromising required correctness
AWS Neuron Ecosystem
- Evaluate implementations using the AWS Neuron SDK
- Review interactions between kernel code, compilation, and Trainium execution
- Assess compiler-related behaviours where relevant
- Apply familiarity with NKI kernel libraries and Neuron tooling
- Identify implementation issues arising from platform-specific constraints
Benchmarking & Validation
- Review benchmark results for Trainium workloads
- Assess performance comparisons and experimental methodology
- Evaluate workloads running on Trn1 or Trn2 instances where applicable
- Determine whether claimed performance improvements are supported by evidence
- Identify benchmarking methodologies that could produce misleading conclusions
Rubric-Based Technical Evaluation
- Assess assigned kernel-development tasks against structured technical criteria
- Provide clear written explanations supporting evaluation decisions
- Reference specific implementation, profiling, or numerical evidence
- Apply evaluation standards consistently across assignments
- Distinguish valid optimisation alternatives from technically flawed approaches
Ideal Profile
- 2+ years of hands-on experience developing or optimising kernels using the Neuron Kernel Interface (NKI)
- Professional experience targeting AWS Trainium or Inferentia2 hardware
- Strong understanding of tile-based computation
- Deep familiarity with SBUF, PSUM, and HBM memory management
- Strong knowledge of partition-dimension constraints and DMA orchestration
- Demonstrated experience evaluating or performing CUDA-to-NKI migrations
- Familiarity with Trainium-specific performance profiling
- Experience assessing NeuronCore pipeline utilisation, tensor-engine throughput, and memory-bandwidth bottlenecks
- Strong understanding of cross-platform numerical correctness and mixed-precision behaviour
- Direct experience with the AWS Neuron SDK, Neuron Compiler internals, or NKI kernel libraries is preferred
- Prior CUDA or Triton kernel development experience is advantageous
- Familiarity with NeuronCore-v2 architecture and supported numerical formats is preferred
- Experience benchmarking ML workloads on Trn1 or Trn2 instances is advantageous
- Strong written communication and ability to provide precise technical feedback
Engagement Details
- Part-time independent contractor engagement
- Fully remote within the United States
- Flexible scheduling based on project requirements
- Compensation: $60–$80/hour
- Work focuses on NKI kernel development, Trainium performance optimisation, CUDA migration, numerical correctness, and technical quality evaluation
- Projects may be extended, shortened, or concluded based on project needs and performance
- Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
- H1-B and STEM OPT support is unavailable for this engagement
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.