← Back to results

AI Red Teamer (Remote)

Join Handshake as an AI Red Teamer to creatively stress-test large language models for safety and robustness.

Location
Remote (USA)
Compensation
Not disclosed
Level
Not specified
Type
contract · Remote

Posted by employer 4 days ago

First seen on Joblaze 2 days ago

Last verified on the company career page 17 hours ago

Apply at Handshake → Save job Scanned from joinhandshake.com

What you'll build

  • Craft creative prompts to stress-test AI guardrails
  • Discover ways around safety filters and defenses
  • Evaluate and score model responses
  • Document experiments clearly
  • Collaborate with engineers and researchers

Must have

  • Strong hands-on experience using multiple LLMs
  • Intuition for crafting adversarial prompts
  • Clear and thoughtful written communication
  • Strong ethical judgment

Nice to have

  • Familiarity with Python or other scripting languages
  • Experience working with LLM APIs
  • Comfort with structured data annotation
  • Prior work in trust and safety or security research

Practical constraints

  • Unable to hire candidates residing in California

AI in the day-to-day

You will stress-test large language models by designing adversarial prompts to expose vulnerabilities.

Requirements

Visa
No sponsorship (stated in posting)

Not disclosed in this posting: compensation, seniority, years of experience.

Joblaze summary

In the role of AI Red Teamer, the individual will focus on stress-testing large language models by crafting adversarial prompts to uncover vulnerabilities and assess model responses across various risk categories. Key skills include hands-on experience with multiple LLMs, creativity in problem-solving, and strong ethical judgment. This position is suited for those with a background in trust and safety, security research, or related fields, particularly those who thrive in collaborative and feedback-rich environments. The team at Handshake AI emphasizes the importance of ethical considerations while working with potentially disturbing content.

Joblaze insights

  • Listed 2 days ago — first seen on Joblaze September 17, 2026. Last confirmed on Handshake's careers page September 19, 2026.
  • AI/ML appears in 47.2% of 1593 comparable ai/ml roles in United States; ChatGPT appears in 0.5% of 1593 comparable ai/ml roles in United States.

Quick facts

Is the AI Red Teamer (Remote) role remote?
Yes — Handshake lists this as a fully remote position.
What's the tech stack?
Joblaze extracted these technologies from the posting: AI/ML, ChatGPT, Claude, Gemini.
Is this full-time or contract?
Contract for this AI Red Teamer (Remote) role at Handshake.

From the original posting

AI Red Teamer (LLM Generalist)

Location: Remote (USA)
Type: Contract, 40 hours per week

About the Role

As an AI Red Teamer, you will stress-test large language models by intentionally trying to break them. Rather than checking whether an answer is correct, you will design creative, adversarial prompts that expose vulnerabilities: unsafe content, bias, broken guardrails, hallucinations, prompt injection weaknesses, and unexpected behaviors. Your work directly supports AI safety and model robustness for leading research labs.

This is a generalist red teaming role. You will probe models across the full spectrum of risk categories, including content safety, CBRN (chemical, biological, radiological, nuclear), cybersecurity, persuasion and influence operations, child safety, self-harm, over-companionship, and regulatory compliance. Red teaming may span text, image, voice, and agentic model capabilities depending on project needs.

This role requires creativity, curiosity, and an ability to think like an adversary while operating with strong ethical judgment.

Day-to-Day Responsibilities

  • Craft creative prompts and multi-turn scenarios to stress-test AI guardrails across diverse risk categories

  • Discover ways around safety filters, restrictions, and defenses using jailbreak, evasion, and prompt injection techniques

  • Explore edge cases to provoke disallowed, harmful, or incorrect outputs

  • Evaluate and score model responses against structured harm taxonomies and severity rubrics

  • Document experiments clearly, including what you tried, why you tried it, and what it revealed

  • Review and refine adversarial prompts generated by other team members

  • Contribute to harm taxonomy development, calibration exercises, and inter-rater reliability work

  • Collaborate with engineers, data scientists, and researchers to share findings and strengthen defenses

  • Work with potentially disturbing content on a regular basis (see Content Warning below)

  • Stay current on jailbreaks, attack methods, and evolving model behaviors

Desired Capabilities

Core

  • Strong hands-on experience using multiple LLMs (ChatGPT, Claude, Gemini, open-source models, etc.)

  • Intuition for crafting adversarial prompts; familiarity with jailbreak or evasion techniques is a strong plus

  • Creative, adversarial problem-solving skills

  • Clear and thoughtful written communication

  • Strong ethical judgment and the ability to separate adversarial thinking from personal values

  • Self-directed, collaborative, and comfortable in feedback-heavy environments

  • Curiosity, persistence, and comfort with frequent failure in experimentation

Nice to Have

  • Familiarity with Python or other scripting languages

  • Experience working with LLM APIs or evaluation tooling

  • Comfort with structured data annotation and rubric-based scoring

  • Prior work in trust and safety, content moderation, QA, or security research

  • Subject matter expertise in any high-risk domain (cybersecurity, chemistry, biology, medicine, law, finance, etc.)

You Will Thrive Here If

  • You treat every model response as a hypothesis to challenge

  • You can switch between creative free-association and rigorous documentation in the same session

  • You go deep into unusual interests (fandoms, niche internet cultures, gaming exploits, Wikipedia rabbit holes, etc.)

  • You come from a creative background: writing, visual art, improv, puzzle design, or similar

  • You are energized by finding the thing nobody else thought to try

  • You are genuinely passionate about AI and follow the space closely

Content Warning

This role involves regular and deliberate exposure to harmful content. You will encounter and intentionally generate content involving violence, self-harm, hate speech, sexually explicit material, child safety scenarios, and other categories of harmful output as part of structured adversarial testing. Candidates must be able to engage with this material professionally and sustainably. Support resources are available.

About Handshake AI

Handshake AI partners with leading AI research labs to make models safer and more robust. Our red teaming operations help identify vulnerabilities before they reach users, contributing directly to the responsible development of frontier AI systems.

California eligibility: We are unable to hire candidates residing in California for this role.

Similar positions

Handshake
AI Red Teamer (Seattle)
Handshake · Seattle, WA
Handshake
AI Red Teamer, CBRNE (Remote)
Handshake · Remote (USA)
Handshake
AI Red Teamer, Cybersecurity (Remote)
Handshake · Remote (USA)
TRM Labs
AI Agent Engineer - US Remote
TRM Labs · United States
Horizon3.ai
Staff Attack Engineer, AI/LLM
Horizon3.ai · US, Remote