Job Openings Remote | QA Test Engineer — $55–$85/hour

About the job Remote | QA Test Engineer — $55–$85/hour

We are sharing a specialised full-time consulting opportunity for experienced QA and test engineers with strong expertise in test-case design, end-to-end debugging, quality assurance, Python, and complex technical evaluation workflows.

This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will review complex multi-step tasks, test reference solutions, identify ambiguity and grading gaps, debug technical environments, and develop repeatable quality processes that keep benchmark results accurate and trustworthy.

Key Responsibilities

Test Case Design

  • Create comprehensive test cases confirming that benchmark tasks function as intended
  • Design positive, negative, boundary, and edge-case tests
  • Validate task requirements, expected outputs, reference solutions, and grading logic
  • Identify scenarios that may produce incorrect or misleading evaluation results
  • Ensure tests measure the intended technical capability accurately

Benchmark Task Review

  • Review complex multi-step tasks and reference solutions before finalisation
  • Identify ambiguous instructions, inconsistent requirements, missing assumptions, and incomplete acceptance criteria
  • Run tasks independently to confirm reproducibility and expected behaviour
  • Assess whether grading standards are clear, fair, and technically defensible
  • Provide actionable feedback to task authors and researchers

Technical Debugging

  • Investigate failures across Python scripts, test harnesses, repositories, and task environments
  • Diagnose unexpected behaviour within unfamiliar codebases
  • Reproduce reported issues and isolate their underlying causes
  • Correct or document environment, dependency, logic, and validation problems
  • Use Git-based workflows to support structured review and collaboration

Quality Process Development

  • Develop practical checklists and repeatable review procedures for benchmark quality
  • Improve consistency across task validation, testing, and approval workflows
  • Document findings clearly so authors can resolve issues efficiently
  • Track recurring defects and recommend preventive quality measures
  • Collaborate closely with researchers, task authors, and other technical reviewers

Benchmark Integrity & Shortcut Detection

  • Examine AI agent runs for unintended shortcuts, loopholes, and grading weaknesses
  • Identify cases where models can receive credit without completing the intended reasoning or technical work
  • Test whether benchmark tasks remain robust across alternative approaches
  • Strengthen evaluation criteria to maintain reliable and meaningful benchmark scores
  • Distinguish valid solution diversity from unintended task exploitation

Ideal Profile

Strong candidates may have:

  • At least 1 year of experience in test engineering, quality assurance, software engineering, research engineering, or a related technical role
  • Demonstrated experience designing test cases and quality-review processes
  • Strong end-to-end debugging skills across complex technical systems
  • Working proficiency in Python and Git
  • Comfort navigating unfamiliar codebases, repositories, and execution environments
  • Exceptional attention to detail and strong written documentation habits
  • Ability to identify ambiguity, edge cases, hidden assumptions, and quality gaps
  • Capacity to work independently through open-ended technical problems
  • Reliable availability for approximately 35 hours per week

Educational Background

  • A master's degree or PhD in a STEM field is highly relevant
  • Equivalent practical experience in an engineering-intensive or research-intensive domain may also be considered
  • Academic or professional experience involving computer science, software engineering, machine learning, mathematics, statistics, or a related technical field may strengthen an application
  • Technical research, open-source contributions, testing projects, or substantial engineering work may also be valuable

Nice to Have

  • Experience with AI training, model evaluation, or quality review of AI-generated work
  • Familiarity with agentic systems and multi-step AI benchmarks
  • Background testing machine learning, research, or data-processing workflows
  • Experience developing automated test suites or validation scripts
  • Familiarity with CI/CD systems, test harnesses, containers, or reproducible environments
  • Experience reviewing reference solutions, grading logic, or technical rubrics
  • Knowledge of adversarial testing, failure-mode analysis, or benchmark design
  • Prior collaboration with AI research or evaluation teams

Why This Opportunity

  • Serve as the quality backbone for advanced agentic AI benchmarks
  • Ensure complex evaluation tasks are accurate, reproducible, and resistant to shortcuts
  • Apply testing and debugging expertise to technically challenging AI research workflows
  • Work closely with researchers and task authors on benchmark improvement
  • Help protect the reliability of evaluation results for frontier AI systems
  • Participate in a structured full-time remote role with competitive hourly compensation

Contract Details

  • Full-time W-2 contingent employment opportunity
  • Fully remote within the United States
  • Expected commitment of approximately 35 hours per week
  • Competitive rates between $55–$85 per hour depending on expertise and project scope
  • Individual task reviews may require one to two days of focused technical work
  • Work may include test design, benchmark review, Python debugging, quality-process development, and shortcut detection
  • Close collaboration with research and task-development teams
  • Engagement scope and duration may evolve according to project requirements and performance

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.