Back to jobs

LLM Red Team Specialist - AI
MercorAnywherePosted 3w ago
AI Evaluation Engineer
Full-time
Remote
Apply on company siteCraft my tailored resume free
Free to start · No card · 5 credits the moment you sign up
Job description
About The Job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco, our investors include Benchmark, General Catalyst, Peter Thiel, Adam D'Angelo, Larry Summers, and Jack Dorsey.
Position: LLM Red Team Specialist — Failure Modes & Edge Cases
Type: Contract
Compensation: $60–$90/hour
Location: Remote
Commitment: 35 hours/week
Role Responsibilities
Evaluate frontier AI models on coding, ML, and analysis tasks to identify vulnerabilities and failure modes.Design complex tasks that challenge models and are fair for grading.Document findings with clear evidence and reproducible steps.Collaborate with task authors to close loopholes and improve grading.Share insights with researchers to enhance benchmark quality.Work independently and asynchronously to meet deadlines and improve AI model performance.
Qualifications
Must-Have
MSc or PhD in a STEM field or equivalent experience.1+ years in research, research-engineering, security, or AI evaluation.Experience identifying vulnerabilities in LLMs or ML systems.Proficiency in Python and Git.Familiarity with LLM capabilities and evaluation techniques.Ability to work 35 hours/week.
Preferred
Experience in AI training, model evaluation, or benchmark/task authoring.
Resources & Support
For details about the interview process and platform information, please check: https://talent.docs.mercor.com/welcomeFor any help or support, reach out to: support@mercor.com
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.