hackajob

AI Safety Expert — English & Malay (Part-Time AI Model Training)

hackajob

Remote · Full Time

Be the first to apply

Experience
Any
Salary
Openings
1
Posted
1 hour ago
Work mode
Work from home
Resume
Required to apply

Job description

About the Role

Mercor is partnering with hackajob to find skilled professionals for a vital role involving AI safety by performing adversarial testing on AI models. The position is remote and requires native-level fluency in both English and Malay. The core mission is to create safer AI systems by identifying potential vulnerabilities through rigorous human-led testing.

Key Responsibilities

  • Conduct adversarial testing on conversational AI models including methods like jailbreaks, prompt injections, misuse scenarios, and bias identification through multi-turn interactions.
  • Produce precise human-generated data by annotating model failures, classifying vulnerabilities, and highlighting systemic risks.
  • Utilize standardized taxonomies, benchmarks, and playbooks to ensure consistency in testing approaches.
  • Document findings with reproducible reports, datasets, and attack cases that help clients mitigate risks effectively.

Candidate Profile

  • Experience in red teaming involving AI adversarial techniques, cybersecurity, or socio-technical probing.
  • Inherently curious and adversarial mindset focused on pushing systems to their limits.
  • Methodical and structured in approach, favoring frameworks and benchmarks over random attempts.
  • Excellent communicator capable of conveying risks clearly to both technical and non-technical audiences.
  • Adaptable and comfortable working across various projects and client environments.

Preferred Expertise

  • Knowledge of adversarial machine learning concepts like jailbreak datasets, prompt injection attacks, RLHF/DPO strategies, and model extraction.
  • Cybersecurity skills such as penetration testing, exploit development, and reverse engineering.
  • Experience analyzing socio-technical risks including harassment, disinformation, abuse, and testing conversational AI.
  • Creative probing skills drawing on psychology, acting, and writing to devise unconventional adversarial approaches.

Measures of Success

  • Identifying vulnerabilities that automated tests fail to detect.
  • Producing repeatable outputs that reinforce the security of customer AI systems.
  • Expanding evaluation coverage by testing a broader range of scenarios to minimize surprises in deployment.
  • Fostering customer confidence in AI safety due to extensive adversarial assessments.

Benefits of Joining Mercor

  • Gain hands-on experience in human-driven AI red teaming at the cutting edge of AI safety research.
  • Contribute directly to enhancing the reliability, safety, and trustworthiness of AI technology.

Work styles they’re looking for

Adaptability Technical Communication Clear Communication Structured Thinking Curiosity

Leave it if you'd like a reply — we won't use it for anything else.

Click to browse, drag & drop, or paste a screenshot

PNG, JPG, GIF, MP4, WebM, MOV · Max 20MB each · Up to 5 files

🤖
Online · instant AI help