About the job Remote | Open Source Software Engineer — $60–$80/hour
We are sharing a specialised part-time consulting opportunity for experienced Software Engineers with hands-on open-source contribution or maintainer experience and strong expertise in repository-level code review, testing, debugging, and software quality evaluation.
This role focuses on reviewing software-engineering benchmark tasks for correctness, reproducibility, and grading integrity. Selected engineers will audit repository-level assignments, reference patches, test harnesses, containerised environments, and evaluation logic while identifying technical flaws, unintended shortcuts, and weaknesses in task design.
Key Responsibilities
Repository-Level Code Review
- Review software engineering tasks built around real code repositories
- Assess whether task requirements are technically clear, complete, and reproducible
- Evaluate repository state, dependencies, configuration, and expected behaviour
- Identify ambiguities or implementation issues that could affect task validity
- Apply practical engineering judgement to realistic codebase-level problems
Reference Patch Auditing
- Review reference patches for correctness and completeness
- Determine whether proposed solutions appropriately address the underlying software issue
- Identify unintended behavioural changes, incomplete fixes, or unsupported assumptions
- Compare reference implementations against task requirements and expected outcomes
- Assess whether alternative valid implementations are treated fairly
Test Harness & Grading Review
- Audit test runners and automated evaluation logic
- Assess whether tests accurately measure the intended behaviour
- Identify missing coverage, brittle assertions, or grading inconsistencies
- Verify that evaluation criteria appropriately distinguish correct from incorrect solutions
- Review benchmark tasks for reliable and repeatable scoring
Reproducibility & Environment Validation
- Evaluate whether tasks can be reproduced consistently across clean environments
- Review dependency installation, build processes, configuration, and runtime requirements
- Assess Docker-based isolation and containerised execution
- Identify environmental dependencies or hidden assumptions affecting reproducibility
- Verify that tasks execute reliably under their intended setup
Benchmark Integrity
- Identify potential answer leakage, unintended shortcuts, or reward-hacking opportunities
- Evaluate whether benchmark structure exposes information that makes tasks artificially easy
- Review task and grading design for loopholes or exploitable behaviours
- Assess whether successful completion genuinely demonstrates the intended engineering capability
- Recommend improvements where benchmark integrity is compromised
Software Testing & Debugging
- Investigate failing or inconsistent benchmark tasks
- Review stack traces, logs, test failures, and repository behaviour
- Identify root causes of technical issues
- Distinguish task defects from legitimate implementation failures
- Assess whether debugging and validation processes follow sound engineering practices
Open-Source Engineering
- Apply experience from contributing to or maintaining open-source software
- Evaluate repository conventions, contribution patterns, and realistic development workflows
- Review patches with the perspective of an experienced contributor or maintainer
- Assess whether proposed changes would meet reasonable code-review expectations
- Apply practical judgement derived from real-world pull request and repository experience
Multi-Language Code Evaluation
- Review software written in Python
- Evaluate tasks involving at least one additional ecosystem such as Java, Go, TypeScript, or C++
- Assess code structure, tests, implementation choices, and repository conventions across languages
- Identify language-specific implementation or testing issues
- Apply consistent engineering standards across different technology stacks
Rubric-Based Evaluation
- Assess benchmark tasks against structured technical criteria
- Provide clear written explanations supporting evaluation decisions
- Reference specific code, tests, patches, or execution behaviour when identifying issues
- Apply grading standards consistently across assignments
- Distinguish substantive benchmark defects from minor implementation differences
Ideal Profile
- 3+ years of professional software engineering experience
- Demonstrated open-source contribution or maintainer experience, such as merged pull requests, committer responsibilities, or maintainer roles
- Strong ability to review repository-level software changes
- Experience auditing reference patches, test runners, and automated test suites
- Comfortable evaluating Docker-based isolation and reproducible development environments
- Strong understanding of software testing, debugging, and code-review practices
- Ability to identify answer leakage, reward hacking, or other benchmark-integrity issues
- Strong proficiency in Python
- Professional fluency in at least one additional language such as Java, Go, TypeScript, or C++
- Familiarity with SWE-Bench Verified or similar repository-level software engineering benchmarks is preferred
- Maintainer or contributor history with established Python open-source projects is highly valued
- Previous code-review, software evaluation, or task-grading experience is advantageous
- Strong written communication and ability to provide precise technical feedback
Engagement Details
- Part-time independent contractor engagement
- Fully remote within the United States
- Flexible scheduling based on project requirements
- Compensation: $60–$80/hour
- Work focuses on repository-level software evaluation, reference-patch review, testing, reproducibility, benchmark integrity, and technical quality assessment
- Projects may be extended, shortened, or concluded based on project needs and performance
- Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
- H1-B and STEM OPT support is unavailable for this engagement
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.