AI Safety Expert — English & Malay (Part-Time AI Model Training)
Remote · Full Time
Be the first to apply
- Experience
- Any
- Salary
- —
- Openings
- 1
- Posted
- 1 hour ago
- Work mode
- Work from home
- Resume
- Required to apply
Job description
About the Role
Mercor is partnering with hackajob to find skilled professionals for a vital role involving AI safety by performing adversarial testing on AI models. The position is remote and requires native-level fluency in both English and Malay. The core mission is to create safer AI systems by identifying potential vulnerabilities through rigorous human-led testing.
Key Responsibilities
- Conduct adversarial testing on conversational AI models including methods like jailbreaks, prompt injections, misuse scenarios, and bias identification through multi-turn interactions.
- Produce precise human-generated data by annotating model failures, classifying vulnerabilities, and highlighting systemic risks.
- Utilize standardized taxonomies, benchmarks, and playbooks to ensure consistency in testing approaches.
- Document findings with reproducible reports, datasets, and attack cases that help clients mitigate risks effectively.
Candidate Profile
- Experience in red teaming involving AI adversarial techniques, cybersecurity, or socio-technical probing.
- Inherently curious and adversarial mindset focused on pushing systems to their limits.
- Methodical and structured in approach, favoring frameworks and benchmarks over random attempts.
- Excellent communicator capable of conveying risks clearly to both technical and non-technical audiences.
- Adaptable and comfortable working across various projects and client environments.
Preferred Expertise
- Knowledge of adversarial machine learning concepts like jailbreak datasets, prompt injection attacks, RLHF/DPO strategies, and model extraction.
- Cybersecurity skills such as penetration testing, exploit development, and reverse engineering.
- Experience analyzing socio-technical risks including harassment, disinformation, abuse, and testing conversational AI.
- Creative probing skills drawing on psychology, acting, and writing to devise unconventional adversarial approaches.
Measures of Success
- Identifying vulnerabilities that automated tests fail to detect.
- Producing repeatable outputs that reinforce the security of customer AI systems.
- Expanding evaluation coverage by testing a broader range of scenarios to minimize surprises in deployment.
- Fostering customer confidence in AI safety due to extensive adversarial assessments.
Benefits of Joining Mercor
- Gain hands-on experience in human-driven AI red teaming at the cutting edge of AI safety research.
- Contribute directly to enhancing the reliability, safety, and trustworthiness of AI technology.
Skills
Work styles they’re looking for
Adaptability
Technical Communication
Clear Communication
Structured Thinking
Curiosity