Job Openings Remote | Computer Science Research Expert (AI Benchmarking) — $55–$75/hour

About the job Remote | Computer Science Research Expert (AI Benchmarking) — $55–$75/hour

We are sharing a specialised part-time consulting opportunity for computer scientists, software engineers, and technical researchers with advanced expertise in computer systems, algorithms, Machine Learning, infrastructure, or software engineering.

This role supports an AI research initiative focused on developing rigorous computer science benchmarks. Selected experts will author or verify challenging multiple-choice assessment content, evaluate technical accuracy and solution quality, develop well-supported reference solutions, and help establish high-quality evaluation standards for advanced AI systems.

Key Responsibilities

Computer Science Question Authoring

  • Create original multiple-choice questions that test deep technical understanding rather than surface-level recall
  • Develop challenging problems within relevant areas of computer science expertise
  • Ensure questions are self-contained, precise, unambiguous, and independently solvable
  • Design questions appropriate for undergraduate, advanced undergraduate, and postgraduate difficulty levels
  • Produce one correct answer alongside nine plausible and carefully constructed distractors

Question Verification & Technical Review

  • Review pre-written computer science questions for correctness, clarity, rigour, and completeness
  • Identify technical errors, hidden assumptions, ambiguity, or solvability issues
  • Edit questions and answer choices where required
  • Verify algorithms, system behaviour, calculations, implementation assumptions, and designated correct answers
  • Document substantive changes and explain the technical reasoning behind them

GPU, Accelerators & Computer Architecture

  • Develop assessment content involving GPU kernel engineering and accelerator programming
  • Create problems addressing parallelism, memory hierarchy, performance optimisation, and computational efficiency
  • Evaluate computer architecture concepts involving processors, accelerators, and specialised hardware
  • Develop questions involving hardware–software interaction and performance trade-offs
  • Apply systems-level reasoning to realistic compute-intensive workloads

Formal Methods & Automated Reasoning

  • Create rigorous problems involving formal verification, logic, theorem proving, and automated reasoning
  • Evaluate correctness properties, specifications, invariants, and proof strategies
  • Develop questions involving model checking and formal system behaviour
  • Apply mathematical reasoning to software and hardware verification scenarios
  • Distinguish formally valid conclusions from plausible but incorrect alternatives

Distributed Systems, Cloud & Infrastructure

  • Develop questions involving distributed systems, cloud architecture, and infrastructure engineering
  • Create scenarios addressing consistency, replication, fault tolerance, scalability, and availability
  • Evaluate distributed-system trade-offs and failure modes
  • Develop content involving DevOps, Site Reliability Engineering, and production infrastructure
  • Assess architecture and operational decisions across large-scale systems

Operating Systems & Systems Engineering

  • Create problems involving operating systems, kernels, concurrency, memory management, and scheduling
  • Evaluate low-level system behaviour and resource-management decisions
  • Develop questions covering processes, threads, synchronisation, storage, and I/O
  • Apply systems reasoning to performance, reliability, and security scenarios
  • Assess implementation trade-offs within systems software

Data Engineering & Databases

  • Develop assessment content involving database systems, data modelling, and large-scale data engineering
  • Create problems involving transactions, indexing, query optimisation, storage, and distributed databases
  • Evaluate data pipelines, processing architectures, and reliability considerations
  • Apply database theory and engineering principles to realistic system scenarios
  • Assess trade-offs involving consistency, performance, scalability, and maintainability

Machine Learning Engineering

  • Develop rigorous questions involving Machine Learning systems and engineering workflows
  • Create problems addressing training, inference, evaluation, optimisation, and deployment
  • Evaluate model-development decisions and production ML architecture
  • Develop questions involving data quality, model performance, scalability, and system constraints
  • Apply technical judgement to realistic Machine Learning engineering scenarios

Software, Web & API Engineering

  • Develop assessment content involving software architecture, web applications, and API development
  • Create problems covering interface design, backend systems, application architecture, and software reliability
  • Evaluate implementation decisions, performance constraints, and maintainability
  • Apply software engineering principles to production-oriented scenarios
  • Develop questions requiring understanding of modern application development practices

Embedded, Graphics, Game & Mobile Systems

  • Create problems involving embedded systems, hardware interfaces, and resource-constrained computing
  • Develop assessment content involving computer graphics, rendering, and game-development systems
  • Evaluate real-time performance, graphics pipelines, and computational trade-offs
  • Create questions involving mobile application architecture and platform constraints
  • Apply domain-specific engineering judgement across specialised computing environments

Academic Research & Benchmark Quality

  • Provide 1–5 academic references per question where required
  • Use reputable sources such as peer-reviewed publications, academic texts, technical documentation, and university repositories
  • Develop clear, structured solution explanations with appropriate intermediate reasoning
  • Rate question difficulty according to defined Medium, Hard, and Expert standards
  • Help maintain consistent technical and academic quality across benchmark content
  • Contribute to gold-standard evaluation materials used to assess advanced AI capabilities

Ideal Profile

  • PhD or current doctoral candidacy in Computer Science, Electrical Engineering, Computer Engineering, or a closely related technical discipline
  • Master's degree may be considered for candidates with exceptional expertise in a relevant subdomain
  • Strong command of graduate-level computer science theory, algorithms, systems design, and/or Machine Learning
  • Expertise in one or more areas including GPU Kernel Engineering, Computer Architecture, Formal Methods, Distributed Systems, DevOps, Site Reliability Engineering, Data Engineering, Databases, Cloud Infrastructure, Operating Systems, Machine Learning Engineering, Web Development, APIs, Embedded Systems, Computer Graphics, Game Development, or Mobile Engineering
  • Strong ability to analyse complex technical systems and verify rigorous solutions
  • Excellent written English and ability to communicate complex technical concepts clearly and precisely
  • Academic research publications are highly valued
  • Substantial software engineering or technical industry experience is highly valued
  • Experience at major technology companies, research organisations, or advanced engineering teams is advantageous
  • Competitive programming, examination development, or advanced technical problem-design experience is highly valued
  • Strong attention to technical precision, logical consistency, and problem solvability

Engagement Details

  • Part-time independent contractor engagement
  • Fully remote and asynchronous
  • Expected availability of at least 10 hours per week
  • Flexible scheduling based on project requirements
  • Compensation: $55–$75/hour
  • Contributors may be assigned either question-authoring or question-verification workstreams
  • Projects may be extended, shortened, or concluded based on project needs and performance
  • Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
  • H1-B and STEM OPT support is unavailable for this engagement

About the Platform

This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.

By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.