About the job Remote | Member of Technical Staff, Enterprise AI — $300,000–$700,000/year
We are sharing a specialised full-time opportunity for experienced technical professionals to operate at the intersection of enterprise AI, applied research, machine-learning evaluation, and real-world AI system performance.
Selected professionals will work directly within enterprise AI workflows to identify real-world failure modes, design high-signal datasets and evaluation frameworks, and run rapid experimental cycles that improve system performance. The role combines forward-deployed research, ML-oriented data design, agentic workflow evaluation, technical analysis, and close collaboration across research, product, domain, and enterprise teams.
Key Responsibilities
Enterprise AI Research & Failure Analysis
- Embed within enterprise AI workflows as a technical research collaborator
- Work alongside domain experts and enterprise teams to understand real-world system behaviour
- Identify, formalise, and prioritise failure modes emerging from deployed AI systems
- Translate operational issues into structured research questions and measurable technical problems
- Produce clear analyses of system behaviour, limitations, and opportunities for improvement
ML-Oriented Data & Evaluation Design
- Design high-signal datasets targeting identified model and system weaknesses
- Develop evaluation protocols, quality criteria, and structured assessment frameworks
- Apply strong judgement to data selection, evaluation design, and research-signal quality
- Identify gaps in existing datasets and evaluation coverage
- Structure research workflows to support measurable improvements in model performance
Experimentation & Agentic Workflow Evaluation
- Run rapid experimental cycles to test hypotheses and quantify system improvements
- Develop and benchmark agentic workflows with a focus on robustness, reliability, and scalability
- Evaluate AI systems operating across complex enterprise workflows
- Analyse experimental results and determine whether observed improvements are meaningful and reproducible
- Iterate on datasets, evaluations, and system configurations based on research findings
Research Tooling & Cross-Functional Collaboration
- Build lightweight tooling to support evaluation, data curation, experimentation, and rapid iteration
- Collaborate across research, engineering, product, domain, and enterprise-facing teams
- Translate research findings into clear, decision-oriented recommendations
- Contribute to research artifacts including reports, benchmarks, evaluation documentation, and technical analyses
- Communicate complex findings clearly to both technical and non-technical stakeholders
Ideal Profile
- Master's degree in Computer Science, Machine Learning, Artificial Intelligence, or a closely related technical discipline
- Strong judgement regarding research-signal quality, data selection, and evaluation design
- Experience designing datasets, evaluation frameworks, or QA processes for machine-learning systems
- Ability to translate ambiguous operational issues into structured research and evaluation problems
- Familiarity with reinforcement-learning environments, agentic systems, or AI-system evaluation
- Strong analytical skills and ability to produce concise, actionable technical insights
- Proven ability to execute effectively within rapid iteration cycles and high-ambiguity environments
- Strong written and verbal communication skills
- Collaborative experience across research, product, engineering, and domain teams
- Client-facing experience within technical or research-focused environments is advantageous
- Experience building internal research or evaluation tooling is beneficial
- Contributions to benchmarks, research publications, or open research initiatives are advantageous
- Exposure to enterprise AI deployments or forward-deployed research environments is strongly valued
Engagement Details
- Full-time engagement
- Fully remote
- Compensation: $300,000–$700,000/year
- Work will span enterprise AI research, evaluation design, ML-oriented data systems, experimentation, and agentic workflow analysis
- Responsibilities will involve direct collaboration with research, product, technical, domain, and enterprise stakeholders
- Research priorities, datasets, evaluation frameworks, and system requirements may evolve based on experimental findings and deployment needs
- Work must be completed without using confidential or proprietary information belonging to any employer, client, institution, or other third party
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy