A

Software Engineer, Safeguards Evals

Anthropic

remote, remote, Canada Full-time July 20, 2026
Apply Now

Vacancy Description

About Anthropic

Anthropic’s mission is to create reliable, interpretable, and steerable AI systems. We want AI to be safe and beneficial for our users and for society as a whole. Our team is a quickly growing group of committed researchers, engineers, policy experts, and business leaders working together to build beneficial AI systems.

About the role

How do we know our safety systems actually catch misuse? Anthropic increasingly uses AI to investigate potential misuse of Claude — analyzing real‑world traffic to surface bad actors, policy violations, and emerging threats. Its findings inform enforcement actions and model launch decisions, which means we need rigorous, trustworthy answers to questions like: Does the monitoring agent catch what it should? Where does it fail? Does it stay reliable as adversaries adapt, as models improve, and as the agent itself changes?

This role builds the evaluation infrastructure that answers those questions. You’ll...

Ready to Apply?

अभी आवेदन करें

Submit your application for Software Engineer, Safeguards Evals at Anthropic

Apply for this Position