Get matched free

AI Red Team Engineer

Apollo Research · London & San Francisco

No salary listed. Estimated from 42 salary-disclosed Policy & Safety roles on this board: $232K–$300K (interquartile range; estimate, not the employer's figure).

Apply at Apollo ResearchFind your warm intro on LinkedIn
About the role

THE OPPORTUNITY

We are currently building Watcher ,  a monitoring tool for coding agents. Our monitoring research agenda attempts to translate compute into safety at scale. Red-teaming previously sat inside the RS (Control) role as a partial responsibility. As it's grown from a single pilot into a recurring need, it now needs a dedicated owner.

As the AI Red Team Engineer, you will help build the practice of red-teaming AI monitors (both Watcher's own defenses and frontier labs' monitoring systems (see our pilot campaign red-teaming Anthropic's auto mode ). You will hunt for attack surfaces monitors that haven't been tested against yet and turn what you find into fixes.You'll work closely with Marius (CEO & currently leads the monitoring efforts), control researchers and product engineers. 

You will like this opportunity if you think like an attacker and want your adversarial findings to directly strengthen AI monitoring systems. You will join a small team and will have significant ability to shape the team & tech, and have the ability to earn responsibility quickly. 

More like this

More Policy & Safety roles · All Apollo Research jobs

Is this your role?
Claim it free with your work email to see how it’s performing here: views, apply-clicks and saves from candidates browsing AI roles. You can also keep it up to date or mark it filled.
We create your account on first use. Only an email on the company’s own domain can claim a role.