AI Red Team Engineer
WFA Digital Insight
White Circle’s AI Red Team Engineer sits at the intersection of security research and cutting‑edge LLM development. Unlike generic pen‑testing gigs, this position demands deep familiarity with prompt‑injection, jailbreaks, and token‑cost abuse—vectors unique to large language models. The role isn’t just about finding bugs; it’s about turning each discovery into reproducible regression tests and concrete sales collateral that fuels the company’s go‑to‑market narrative. Candidates will work hands‑on with the company’s own fine‑tuned models, write lightweight Python automation, and maintain a living attack library that feeds both product improvement and customer demos. It’s a rare chance to influence AI safety practice from inside a funded startup backed by AI heavyweights.
Job Description
TLDR: We're looking for an AI Red Team Engineer to break LLM-powered systems responsibly, automate the repetitive attacks, and turn their findings into clear evidence that powers customer demos, security reviews, and sales conversations. You'll own hands-on adversarial testing end to end: find the failure, prove it, script it, and write it up. About us White Circle https://whitecircle.ai/ is an AI Safety company building the safety, reliability, and optimization layer for AI systems. At the core of our platform are policies – simple natural-language rules that define what an AI model should and shouldn’t do. We automatically test, enforce, and continuously improve these policies at scale.
- We’ve raised $11M from top funds, founders, and senior leaders at OpenAI, Anthropic, HuggingFace, Mistral, DeepMind, Datadog, Sentry, and others - We process over one hundred million API calls every month - We fine-tune and train our own LLMs so they run faster and cheaper than any open or proprietary model We’re a small, highly focused team.
- Test for jailbreaks, prompt injection, system-prompt and tool leakage, sensitive-data and context leakage, unsafe outputs, policy bypass, tool misuse, excessive agency, resource and token-cost abuse, and business-logic abuse.
- Write lightweight Python to automate attacks, run prompt sets, call model APIs, collect and score responses, and generate repeatable reports.
- Build and maintain an internal attack library: prompts, scenarios, test cases, regression tests, scoring rubrics, and reusable demo cases.
- Turn model failures into clear reports: what happened, why it matters, how to reproduce it, how severe it is, and how to fix it.
- Convert successful attacks into regression tests and product requirements.
- Track new red-team and safety techniques and fold the useful ones into our tests.
- Support GTM by producing strong, credible evidence for customer demos, security reviews, and sales conversations.
- Have a background in QA automation, AppSec, API/security/pen testing, or bug bounty.
- Have strong Python scripting skills.
- Have experience testing APIs, web apps, backends, or SaaS products.
- Are hands-on with LLMs, prompts, system instructions, RAG, agents, and tool/function calling.
- Understand LLM-specific abuse vectors (prompt injection, jailbreaks, data leakage, tool misuse, excessive agency, token-cost exhaustion).
- Can find bypasses, abuse edge cases, chain failures, and reason about real-world impact.
- Can separate real customer risk from low-impact prompt tricks.
- Write clear, reproducible bug reports in clear English.
- Can move fast without perfect requirements.
- Hold a firm ethical line: you red-team to make systems safer, operate within scope and the law, and don't produce or traffic in genuinely harmful material.
- Experience with modern LLM red-teaming automated agents and pipelines.
- Familiarity with LangChain, LangGraph, LlamaIndex, RAG pipelines, AI agents, tool/function calling, and LLM-as-judge evaluation.
- Familiarity with OWASP LLM Top 10, OWASP Web Top 10, MITRE ATLAS, or other AI security taxonomies.
- Experience testing RAG systems, AI agents, tool-calling workflows, browser agents, or internal copilots.
- Experience writing customer-facing security reports.
- Experience with trust & safety, abuse prevention, fraud, moderation, or platform security.
- Experience building eval pipelines, regression suites, dashboards, or CI-friendly security tests.
- A track record in CTFs, red-team competitions, or responsible-disclosure / bounty programs.
How to Stand Out
- Showcase a portfolio of Python scripts that automate security testing against APIs or LLMs; concrete examples speak louder than a resume.
- Highlight any bug bounty or CTF achievements, especially those involving prompt injection or model exploitation.
- Be prepared to discuss specific LLM abuse vectors you’ve mitigated and the reasoning behind severity assessments.
- During interviews, demonstrate your ability to write clear, reproducible bug reports on the spot.
- Emphasize ethical red‑team practices—explain how you stay within scope and legal boundaries.
- If asked about GTM support, illustrate how technical evidence can be translated into persuasive sales material.
- Watch for vague promises about “flexible hours” without clear remote‑work policies; ask about communication cadence and team overlap expectations.
This is a remote position listed on WFA Digital, the platform for professionals who work from anywhere. Browse more remote jobs across all categories.