About the job Remote | AI Safety Specialist (English & Swedish) — $45–$60/hour
We are sharing a specialised consulting opportunity for AI safety and red teaming professionals with native-level fluency in both English and Swedish and experience evaluating conversational AI systems through adversarial testing.
This role supports an AI safety initiative focused on identifying vulnerabilities, failure modes, and systemic risks in conversational models and agents. Selected professionals will conduct structured red teaming, generate high-quality evaluation data, classify model failures, and document reproducible attack scenarios that can be used to improve model robustness and safety.
Key Responsibilities
AI Red Teaming
- Conduct adversarial testing of conversational AI models and agents
- Develop jailbreaks, prompt-injection scenarios, misuse cases, and multi-turn manipulation strategies
- Probe models for vulnerabilities that may be missed by automated evaluation systems
- Test model behaviour across diverse conversational and adversarial scenarios
- Apply systematic testing frameworks and established evaluation methodologies
Safety & Vulnerability Evaluation
- Identify and classify model failures and safety vulnerabilities
- Evaluate issues involving bias, misinformation, misuse, and potentially harmful model behaviours
- Identify recurring or systemic patterns across model responses
- Assess the severity, reproducibility, and practical significance of discovered vulnerabilities
- Apply established taxonomies, benchmarks, and testing playbooks consistently
Human Data & Annotation
- Produce high-quality human evaluation data from red teaming activities
- Annotate model failures and categorise identified vulnerabilities
- Create structured attack cases and supporting evaluation materials
- Maintain consistency across repeated assessments and datasets
- Provide clear rationale for classifications and safety judgments
Reporting & Documentation
- Document adversarial scenarios in a clear and reproducible format
- Produce reports, datasets, and structured findings that technical teams can act upon
- Explain identified risks to both technical and non-technical stakeholders
- Record testing methodology, model behaviour, and relevant failure patterns
- Contribute to broader evaluation coverage across models and use cases
Ideal Profile
Strong candidates may have:
- Native-level fluency in both English and Swedish
- Prior experience with AI red teaming, adversarial AI evaluation, cybersecurity, or socio-technical system testing
- Strong understanding of conversational AI systems and model failure modes
- Experience developing structured adversarial tests and evaluation frameworks
- Ability to identify subtle vulnerabilities and recurring behavioural patterns
- Strong analytical reasoning and written communication skills
- Ability to document findings clearly and reproducibly
- Comfort working across changing projects, scenarios, and evaluation frameworks
Educational Background
- A background in computer science, cybersecurity, artificial intelligence, machine learning, linguistics, behavioural science, or a related discipline may be helpful
- Equivalent professional experience in AI safety, adversarial testing, security, or structured model evaluation may also be considered
- Practical red teaming experience is particularly valuable for this engagement
Nice to Have
- Experience creating jailbreak or prompt-injection datasets
- Familiarity with adversarial machine learning
- Knowledge of RLHF, DPO, model extraction, or related AI training and evaluation concepts
- Cybersecurity experience involving penetration testing, exploit development, or reverse engineering
- Background analysing abuse, harassment, misinformation, or other socio-technical risks
- Experience testing conversational AI systems
- Strong creative writing, psychology, or behavioural-analysis skills applicable to adversarial testing
- Previous experience producing structured human data for AI evaluation
Why This Opportunity
- Apply adversarial thinking directly to advanced AI safety work
- Identify vulnerabilities that automated testing may overlook
- Generate human evaluation data that can improve model robustness
- Work across jailbreaks, prompt injections, misuse scenarios, bias, and conversational manipulation
- Contribute to safer and more reliable AI systems
- Participate in flexible remote consulting work with competitive hourly compensation
Contract Details
- Independent contractor role
- Fully remote with flexible scheduling
- Competitive rates between $45–$60 per hour depending on expertise and project scope
- Native-level English and Swedish proficiency is required
- Work may include AI red teaming, adversarial testing, vulnerability classification, annotation, and structured reporting
- Some projects may involve reviewing text-based AI outputs concerning sensitive topics such as bias, misinformation, or harmful behaviours
- Participation in higher-sensitivity projects is optional, with the relevant subject matter communicated before exposure
- Weekly payments via Stripe or Wise
- Projects may be extended, shortened, or adjusted depending on scope and performance
- Work will not involve access to confidential or proprietary information from any employer, client, or institution
About the Platform
This opportunity is available through 24-MAG LLC. We connect experienced professionals with remote consulting opportunities across technical, evaluation, and project-based workstreams.
By submitting this application, you acknowledge that your information may be processed by 24-MAG LLC for recruitment and opportunity matching in accordance with our Privacy Policy: https://www.24-mag.com/privacy-policy.