Job Openings Remote | Open Source Software Engineer — $60–$80/hour

About the job Remote | Open Source Software Engineer — $60–$80/hour

We are sharing a specialised part-time consulting opportunity for experienced Software Engineers with hands-on open-source contribution or maintainer experience and strong expertise in repository-level code review, testing, debugging, and software quality evaluation.

This role focuses on reviewing software-engineering benchmark tasks for correctness, reproducibility, and grading integrity. Selected engineers will audit repository-level assignments, reference patches, test harnesses, containerised environments, and evaluation logic while identifying technical flaws, unintended shortcuts, and weaknesses in task design.

Key Responsibilities

Repository-Level Code Review

  • Review software engineering tasks built around real code repositories
  • Assess whether task requirements are technically clear, complete, and reproducible
  • Evaluate repository state, dependencies, configuration, and expected behaviour
  • Identify ambiguities or implementation issues that could affect task validity
  • Apply practical engineering judgement to realistic codebase-level problems

Reference Patch Auditing

  • Review reference patches for correctness and completeness
  • Determine whether proposed solutions appropriately address the underlying software issue
  • Identify unintended behavioural changes, incomplete fixes, or unsupported assumptions
  • Compare reference implementations against task requirements and expected outcomes
  • Assess whether alternative valid implementations are treated fairly

Test Harness & Grading Review

  • Audit test runners and automated evaluation logic
  • Assess whether tests accurately measure the intended behaviour
  • Identify missing coverage, brittle assertions, or grading inconsistencies
  • Verify that evaluation criteria appropriately distinguish correct from incorrect solutions
  • Review benchmark tasks for reliable and repeatable scoring

Reproducibility & Environment Validation

  • Evaluate whether tasks can be reproduced consistently across clean environments
  • Review dependency installation, build processes, configuration, and runtime requirements
  • Assess Docker-based isolation and containerised execution
  • Identify environmental dependencies or hidden assumptions affecting reproducibility
  • Verify that tasks execute reliably under their intended setup

Benchmark Integrity

  • Identify potential answer leakage, unintended shortcuts, or reward-hacking opportunities
  • Evaluate whether benchmark structure exposes information that makes tasks artificially easy
  • Review task and grading design for loopholes or exploitable behaviours
  • Assess whether successful completion genuinely demonstrates the intended engineering capability
  • Recommend improvements where benchmark integrity is compromised

Software Testing & Debugging

  • Investigate failing or inconsistent benchmark tasks
  • Review stack traces, logs, test failures, and repository behaviour
  • Identify root causes of technical issues
  • Distinguish task defects from legitimate implementation failures
  • Assess whether debugging and validation processes follow sound engineering practices

Open-Source Engineering

  • Apply experience from contributing to or maintaining open-source software
  • Evaluate repository conventions, contribution patterns, and realistic development workflows
  • Review patches with the perspective of an experienced contributor or maintainer
  • Assess whether proposed changes would meet reasonable code-review expectations
  • Apply practical judgement derived from real-world pull request and repository experience

Multi-Language Code Evaluation

  • Review software written in Python
  • Evaluate tasks involving at least one additional ecosystem such as Java, Go, TypeScript, or C++
  • Assess code structure, tests, implementation choices, and repository conventions across languages
  • Identify language-specific implementation or testing issues
  • Apply consistent engineering standards across different technology stacks

Rubric-Based Evaluation

  • Assess benchmark tasks against structured technical criteria
  • Provide clear written explanations supporting evaluation decisions
  • Reference specific code, tests, patches, or execution behaviour when identifying issues
  • Apply grading standards consistently across assignments
  • Distinguish substantive benchmark defects from minor implementation differences

Ideal Profile

  • 3+ years of professional software engineering experience
  • Demonstrated open-source contribution or maintainer experience, such as merged pull requests, committer responsibilities, or maintainer roles
  • Strong ability to review repository-level software changes
  • Experience auditing reference patches, test runners, and automated test suites
  • Comfortable evaluating Docker-based isolation and reproducible development environments
  • Strong understanding of software testing, debugging, and code-review practices
  • Ability to identify answer leakage, reward hacking, or other benchmark-integrity issues
  • Strong proficiency in Python
  • Professional fluency in at least one additional language such as Java, Go, TypeScript, or C++
  • Familiarity with SWE-Bench Verified or similar repository-level software engineering benchmarks is preferred
  • Maintainer or contributor history with established Python open-source projects is highly valued
  • Previous code-review, software evaluation, or task-grading experience is advantageous
  • Strong written communication and ability to provide precise technical feedback

Engagement Details

  • Part-time independent contractor engagement
  • Fully remote within the United States
  • Flexible scheduling based on project requirements
  • Compensation: $60–$80/hour
  • Work focuses on repository-level software evaluation, reference-patch review, testing, reproducibility, benchmark integrity, and technical quality assessment
  • Projects may be extended, shortened, or concluded based on project needs and performance
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
  • H1-B and STEM OPT support is unavailable for this engagement

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.