About the job Remote | AI Safety & Adversarial Testing Specialist — $65–$80/hour
We are sharing a specialised part-time consulting opportunity for experienced AI safety, red-teaming, trust and safety, cybersecurity, investigative research, and scientific-risk professionals with expertise in adversarial evaluation of advanced AI systems.
This role supports a frontier AI safety initiative focused on identifying vulnerabilities, unsafe behaviours, policy failures, and robustness gaps through structured adversarial testing. Selected professionals will design challenging prompts, evaluate model behaviour across high-risk and ambiguous scenarios, document findings, and contribute to safety benchmarks used to improve model alignment and reliability.
Key Responsibilities
Adversarial Prompt Design
- Design sophisticated prompts that stress-test advanced AI systems
- Develop realistic scenarios intended to expose model limitations, policy weaknesses, and inconsistent behaviour
- Test direct, indirect, multi-turn, and context-dependent adversarial strategies
- Create evaluations covering both clearly unsafe requests and complex grey-area situations
Safety & Robustness Evaluation
- Identify jailbreak vulnerabilities, hallucinations, unsafe outputs, and instruction-following failures
- Assess model robustness across misinformation, cybersecurity, biosecurity, fraud, political content, and scientific safety
- Evaluate whether responses appropriately balance safety, usefulness, factual accuracy, and policy adherence
- Distinguish isolated failures from broader or reproducible behavioural patterns
Vulnerability Analysis & Reporting
- Document identified vulnerabilities through clear, structured, and reproducible reports
- Explain testing methods, observed behaviours, severity, and potential impact
- Classify failures according to established safety taxonomies and evaluation frameworks
- Contribute findings to red-teaming reports, benchmark datasets, and model-improvement workflows
Research Collaboration & Calibration
- Collaborate with AI researchers, safety specialists, and other domain experts
- Participate in calibration exercises to maintain consistent evaluation standards
- Review adversarial tasks and findings developed by other contributors
- Support improvements to testing methodologies, safety rubrics, and vulnerability classifications
Ideal Profile
Strong candidates may have:
- At least 5 years of professional experience in AI safety, AI red teaming, trust and safety, cybersecurity, investigative journalism, life sciences, public policy, or a related field
- Hands-on experience designing adversarial prompts or evaluating advanced AI systems
- Strong analytical reasoning and the ability to identify subtle safety and policy failures
- Experience working with complex, high-risk, or ambiguous subject matter
- Excellent written communication and the ability to produce clear technical findings
- Ability to work independently while applying structured evaluation standards
- Professional residence in one of the eligible countries listed below
Educational Background
- A bachelor's degree or higher in computer science, cybersecurity, journalism, communications, psychology, biology, chemistry, public policy, or a related discipline is required
- Graduate-level education in artificial intelligence, security, behavioural science, life sciences, or policy may be valuable
- Equivalent specialist experience in adversarial testing, safety research, or high-risk investigations may also be considered
- Research publications, safety evaluations, or relevant technical portfolios may strengthen an application
Nice to Have
- Experience with AI red teaming, reinforcement learning from human feedback, supervised fine-tuning, AI alignment, or trust and safety
- Familiarity with jailbreak testing, prompt engineering, model-behaviour analysis, or adversarial evaluation methodologies
- Expertise in cybersecurity, biosecurity, political content, misinformation, fraud, or scientific safety
- Experience developing safety benchmarks, evaluation rubrics, or structured testing frameworks
- Knowledge of responsible disclosure, threat modelling, or vulnerability-severity assessment
- Previous collaboration with AI researchers, policy teams, security engineers, or scientific specialists
- Familiarity with frontier-model safety policies and model-alignment workflows
Why This Opportunity
- Help strengthen the safety and robustness of advanced AI systems
- Work on complex adversarial testing alongside experienced researchers and safety specialists
- Apply domain expertise to realistic high-risk and grey-area scenarios
- Influence how AI systems respond to sensitive and consequential real-world requests
- Participate in flexible remote work with competitive hourly compensation
Contract Details
- Independent contractor role
- Fully remote with flexible scheduling
- Competitive rates between $65–$80 per hour depending on expertise and project scope
- Weekly payments via Stripe or Wise
- Eligible locations include Albania, Austria, Belgium, Bosnia and Herzegovina, Bulgaria, Croatia, the Czech Republic, Denmark, Estonia, Finland, France, Germany, Greece, Hungary, Iceland, Ireland, Italy, Kosovo, Latvia, Liechtenstein, Lithuania, Luxembourg, Malta, Moldova, Monaco, the Netherlands, North Macedonia, Norway, Poland, Portugal, Romania, San Marino, Serbia, Slovakia, Slovenia, Spain, Sweden, Switzerland, the United Kingdom, and the United States
- Projects may be extended, shortened, or adjusted depending on scope and performance
- Work will not involve access to confidential or proprietary information from any employer, client, or institution
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.