About the job Remote | QA Test Engineer — $55–$85/hour
We are sharing a specialised full-time consulting opportunity for experienced QA and test engineers with strong expertise in test-case design, end-to-end debugging, quality assurance, Python, and complex technical evaluation workflows.
This role supports the development of advanced agentic evaluation benchmarks for frontier AI models. Selected professionals will review complex multi-step tasks, test reference solutions, identify ambiguity and grading gaps, debug technical environments, and develop repeatable quality processes that keep benchmark results accurate and trustworthy.
Key Responsibilities
Test Case Design
- Create comprehensive test cases confirming that benchmark tasks function as intended
- Design positive, negative, boundary, and edge-case tests
- Validate task requirements, expected outputs, reference solutions, and grading logic
- Identify scenarios that may produce incorrect or misleading evaluation results
- Ensure tests measure the intended technical capability accurately
Benchmark Task Review
- Review complex multi-step tasks and reference solutions before finalisation
- Identify ambiguous instructions, inconsistent requirements, missing assumptions, and incomplete acceptance criteria
- Run tasks independently to confirm reproducibility and expected behaviour
- Assess whether grading standards are clear, fair, and technically defensible
- Provide actionable feedback to task authors and researchers
Technical Debugging
- Investigate failures across Python scripts, test harnesses, repositories, and task environments
- Diagnose unexpected behaviour within unfamiliar codebases
- Reproduce reported issues and isolate their underlying causes
- Correct or document environment, dependency, logic, and validation problems
- Use Git-based workflows to support structured review and collaboration
Quality Process Development
- Develop practical checklists and repeatable review procedures for benchmark quality
- Improve consistency across task validation, testing, and approval workflows
- Document findings clearly so authors can resolve issues efficiently
- Track recurring defects and recommend preventive quality measures
- Collaborate closely with researchers, task authors, and other technical reviewers
Benchmark Integrity & Shortcut Detection
- Examine AI agent runs for unintended shortcuts, loopholes, and grading weaknesses
- Identify cases where models can receive credit without completing the intended reasoning or technical work
- Test whether benchmark tasks remain robust across alternative approaches
- Strengthen evaluation criteria to maintain reliable and meaningful benchmark scores
- Distinguish valid solution diversity from unintended task exploitation
Ideal Profile
Strong candidates may have:
- At least 1 year of experience in test engineering, quality assurance, software engineering, research engineering, or a related technical role
- Demonstrated experience designing test cases and quality-review processes
- Strong end-to-end debugging skills across complex technical systems
- Working proficiency in Python and Git
- Comfort navigating unfamiliar codebases, repositories, and execution environments
- Exceptional attention to detail and strong written documentation habits
- Ability to identify ambiguity, edge cases, hidden assumptions, and quality gaps
- Capacity to work independently through open-ended technical problems
- Reliable availability for approximately 35 hours per week
Educational Background
- A master's degree or PhD in a STEM field is highly relevant
- Equivalent practical experience in an engineering-intensive or research-intensive domain may also be considered
- Academic or professional experience involving computer science, software engineering, machine learning, mathematics, statistics, or a related technical field may strengthen an application
- Technical research, open-source contributions, testing projects, or substantial engineering work may also be valuable
Nice to Have
- Experience with AI training, model evaluation, or quality review of AI-generated work
- Familiarity with agentic systems and multi-step AI benchmarks
- Background testing machine learning, research, or data-processing workflows
- Experience developing automated test suites or validation scripts
- Familiarity with CI/CD systems, test harnesses, containers, or reproducible environments
- Experience reviewing reference solutions, grading logic, or technical rubrics
- Knowledge of adversarial testing, failure-mode analysis, or benchmark design
- Prior collaboration with AI research or evaluation teams
Why This Opportunity
- Serve as the quality backbone for advanced agentic AI benchmarks
- Ensure complex evaluation tasks are accurate, reproducible, and resistant to shortcuts
- Apply testing and debugging expertise to technically challenging AI research workflows
- Work closely with researchers and task authors on benchmark improvement
- Help protect the reliability of evaluation results for frontier AI systems
- Participate in a structured full-time remote role with competitive hourly compensation
Contract Details
- Full-time W-2 contingent employment opportunity
- Fully remote within the United States
- Expected commitment of approximately 35 hours per week
- Competitive rates between $55–$85 per hour depending on expertise and project scope
- Individual task reviews may require one to two days of focused technical work
- Work may include test design, benchmark review, Python debugging, quality-process development, and shortcut detection
- Close collaboration with research and task-development teams
- Engagement scope and duration may evolve according to project requirements and performance
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.