AI Evals Jobs
91 open roles matching this title right now, live from company career pages.

Software Engineer - Evals
xAI · Palo Alto, California, United States · Late · raised 2026

Manager, Test and Evaluation Engineering - Intel Systems
Anduril Industries · Reston, Virginia, United States

Manager, Test & Evaluation Engineer, CW
Anduril Industries · Costa Mesa, California, United States

Senior Manager, Test and Evaluation, Mission Systems
Anduril Industries · Costa Mesa, California, United States

Senior Project Engineer, Test and Evaluation
Anduril Industries · Costa Mesa, California, United States; San Clemente, California, United States

Senior Test & Evaluation Engineer
Anduril Industries · Quincy, Massachusetts, United States

Senior Test & Evaluation Engineer, Titan
Anduril Industries · Costa Mesa, California, United States

Test and Evaluation Engineer, Air Defense
Anduril Industries · Costa Mesa, California, United States; San Clemente, California, United States

Test & Evaluation Engineer
Anduril Industries · Quincy, Massachusetts, United States

Test & Evaluation Engineer, Imaging
Anduril Industries · Boulder, Colorado, United States

Machine Learning Engineer, Driver Understanding and Evaluation
Waymo · Mountain View, CA, USA · Late

Machine Learning Engineer (Infra), Driver Understanding and Evaluation
Waymo · Mountain View, CA, USA · Late

Product Manager, New Geo Evaluation
Waymo · Mountain View, CA, US; San Francisco, CA, US · Late

Senior Machine Learning Engineer, Driver Understanding and Evaluation
Waymo · Mountain View, CA, USA · Late

Senior Machine Learning Engineer (Infra), Driver Understanding and Evaluation
Waymo · Mountain View, CA, USA · Late

Senior Machine Learning Engineer, Simulation Evaluation
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late

Senior Software Engineer, Eval Authoring APIs
Waymo · Mountain View,CA, USA; San Francisco, CA, USA; New York, NY, USA · Late

Senior Software Engineer, ML/Eval Data Platforms & Infrastructure
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late

Senior Software Engineer, ML Evaluation Infra and Efficiency
Waymo · Mountain View, California · Late

Senior Software Engineer, Planner Evaluation
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late

Senior Software Engineer, Quantitative Evaluations
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late

Senior Software Engineer, Simulator Evaluation
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late

Senior Software Engineer, Statistical Evaluation and Sampling
Waymo · Mountain View, CA, USA; San Francisco, CA, USA; New York, NY, USA · Late

Senior Staff ML Engineer, Driver Understanding and Evaluation
Waymo · Mountain View, CA, United States · Late

Senior Staff TLM, Data Mining and Sampling for ML and Evaluation
Waymo · Mountain View, California, United States; San Francisco, California, United States; New York City, New York, United States. · Late

Software Engineer, Perception Scaling Evaluation
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late

Software Engineer, Quantitative Evaluations
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late

Software Engineer, Statistical Evaluation and Sampling
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late

Software Quality Operations Specialist, Safety Evaluation
Waymo · Hyderabad, India · Late

Staff Machine Learning Engineer, Driver Understanding and Evaluation
Waymo · Mountain View, CA, USA · Late

Staff Machine Learning Engineer (Infra), Driver Understanding and Evaluation
Waymo · Mountain View, CA, USA · Late

Staff Machine Learning Engineer (TLM), Driver Understanding and Evaluation
Waymo · London, UK · Late

Staff Machine Learning Engineer – VLM/LLM Evaluation
Waymo · Mountain View, CA, USA; San Francisco, CA, USA; Kirkland, WA, USA; New York City, NY, USA · Late

Staff Software Engineer, Planner Evaluation
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late

Staff Software Engineer, Quantitative Evaluations
Waymo · Mountain View, CA, USA; San Francisco, CA, USA; New York, NY, USA · Late

Staff Software Engineer, Simulator Evaluation
Waymo · Mountain View, California, United States; San Francisco, California, United States. · Late

Tech Lead, Self Driving Eval Infrastructure
Waymo · Mountain View, CA, USA · Late

Technical Lead Manager, Prediction, ML Evaluation
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late

Applied Scientist, Price Perception and Evaluation Science
Amazon (AI roles) · Seattle, Washington, USA · Public

Director of Platform Management for Simulation, Evaluation & Validation
Wayve · London; Sunnyvale

Software Engineer, Evals
Glean · Bangalore, India · Late

Technical Program Manager, Safeguards (Infrastructure & Evals)
Anthropic · San Francisco, CA | New York City, NY | Seattle, WA · Late · raised 2026

2026 Early Career Test & Evaluation Engineer
Anduril Industries · Costa Mesa, California, United States

Director of Engineering, Eval Platform
Nuro · Mountain View, California (HQ)

Technical Lead, Evaluation Infrastructure
Nuro · Mountain View, California (HQ)

Technical Lead Manager, Autonomy Evaluation and Intelligence
Nuro · Mountain View, California (HQ)

Engineering Manager, Agent Prompts & Evals
Anthropic · San Francisco, CA | New York City, NY · Late · raised 2026

Research Engineer, Model Evaluations
Anthropic · Remote-Friendly (Travel-Required) | San Francisco, CA | New York City, NY · Late · raised 2026Remote

Safeguards Enforcement Analyst, Safety Evaluations
Anthropic · Remote-Friendly (Travel-Required) | San Francisco, CA | Washington, DC; San Francisco, CA | New York City, NY · Late · raised 2026Remote

Staff+ Software Engineer, Safeguards Evals
Anthropic · San Francisco, CA | New York City, NY · Late · raised 2026

National Security Cyber Evaluation Lead
OpenAI · Washington, DC · Late · raised 2026Remote$252K – $342K • Offers Equity

Senior Software Engineer, Metrics and Evaluation - Autonomous Vehicles
NVIDIA · 5 Locations · Public

Research Scientist, Frontier Risk Evaluations
Scale AI · San Francisco, CA; New York, NY · Late

Director, Research - Evaluation & Training
Snorkel AI · San Francisco, CA (Hybrid) · Late

Member of Technical Staff (Data Scientist, Evals)
Perplexity · San Francisco · Late · raised 2026Remote$200K – $300K

Agent Post-Training, Frontier Evals and Environments Research
OpenAI · San Francisco · Late · raised 2026$295K – $445K

Evaluation Engineer, Applied AI
Mistral AI · Paris · Late

Machine Learning Engineer, LLM Evals & Observability
Glean · Mountain View, CA · Late

Machine Learning Engineer, LLM Evals & Observability
Glean · San Francisco, CA · Late

Senior Program Manager, Eval Operations
Nuro · Mountain View, California (HQ)

Product Lead, AI/ML (Evals)
Abridge · SF Office · Late · raised 2026Remote$250K – $290K • Offers Equity

Senior Deep Learning Engineer - Model Evaluation & AI Systems
NVIDIA · US, CA, Santa Clara · Public

Senior Research Manager, World Model Evaluation
NVIDIA · US, CA, Santa Clara · Public

AI Engineer, Evaluation
Distyl AI · San Francisco · GrowthRemote

Senior Product Operations Manager, Evaluation
Harvey · San Francisco · Late · raised 2026Remote$150K – $210K • Offers Equity • Offers Bonus

Software Engineering Manager, AI Observability & Evals Platform (New York, NY)
LangChain · New York, NY · Growth

Engineering Manager, Evals
Cursor · San Francisco · Late

Member of Technical Staff, Evals
Magic · San Francisco · Growth$200K – $550K

Applied Scientist II, GenAI Evaluation Media (GEM)
Amazon (AI roles) · Seattle, Washington, USA · Public

Principal Software Engineer, AI Observability & Evals Platform
LangChain · Boston, MA · Growth

Research Program Manager - Model Evals and Safety
Reflection AI · New York · Late · raised 2026

Member of Engineering (Evaluations)
Poolside · Remote (EMEA/East Coast) · GrowthRemote

Member of Engineering (Evaluations / Engineering)
Poolside · Remote (EMEA/East Coast) · GrowthRemote

Software Engineer, Agent Evaluation and Quality
Cursor · San Francisco · Late

Research Engineer – Benchmarking, Evals & Failure Analysis
Mercor · San Francisco · Late$130K – $500K • Offers Equity

Senior Applied Scientist, HST Health Evaluation
Amazon (AI roles) · Bengaluru, Karnataka, IND · Public

Applied Scientist II, HST Health Evaluation
Amazon (AI roles) · Bengaluru, Karnataka, IND · Public

Senior Software Engineer, Evaluation Infrastructure
Waabi · Toronto, ON

Research Scientist (Measurement and Evaluation)
Abridge · NYC Office · Late · raised 2026Remote$188K – $277K • Offers Equity

Software Engineering Manager, AI Observability & Evals Platform (San Francisco, CA)
LangChain · San Francisco, CA · Growth

Senior Backend Software Engineer, AI Observability & Evals Platform (LangSmith)
LangChain · San Francisco, CA · Growth

Backend Software Engineer (Evals)
OpenAI · San Francisco · Late · raised 2026$230K – $385K • Offers Equity

Member of Technical Staff - Evaluations
Reflection AI · San Francisco · Late · raised 2026

Member of Technical Staff, Data Analysis and Evaluation
Cohere · London · Late · raised 2026Remote

Member of Technical Staff, Evals & Post-Training Product
Fireworks AI · San Mateo · Late · raised 2026

Senior Research Scientist, Model Evaluation
Cohere · Toronto · Late · raised 2026Remote

Researcher, Evals
Cartesia · *HQ - San Francisco, CA$220K – $350K • Offers Equity

Research, Evals
Exa · San Francisco, California · Growth · raised 2026$180K – $350K • Offers Equity

FullStack Engineer, AI Observability & Evals Platform (LangSmith)
LangChain · San Francisco, CA · Growth

Research Engineer, Frontier Evals & Environments
OpenAI · San Francisco · Late · raised 2026$205K – $380K • Offers Equity

Senior Fullstack Engineer, AI Observability & Evals Platform
LangChain · San Francisco, CA · Growth