Post a job

AI Evals Jobs

91 open roles matching this title right now, live from company career pages.

Software Engineer - Evals
xAI · Palo Alto, California, United States · Late · raised 2026
Manager, Test and Evaluation Engineering - Intel Systems
Anduril Industries · Reston, Virginia, United States
Manager, Test & Evaluation Engineer, CW
Anduril Industries · Costa Mesa, California, United States
Senior Manager, Test and Evaluation, Mission Systems
Anduril Industries · Costa Mesa, California, United States
Senior Project Engineer, Test and Evaluation
Anduril Industries · Costa Mesa, California, United States; San Clemente, California, United States
Senior Test & Evaluation Engineer
Anduril Industries · Quincy, Massachusetts, United States
Senior Test & Evaluation Engineer, Titan
Anduril Industries · Costa Mesa, California, United States
Test and Evaluation Engineer, Air Defense
Anduril Industries · Costa Mesa, California, United States; San Clemente, California, United States
Test & Evaluation Engineer
Anduril Industries · Quincy, Massachusetts, United States
Test & Evaluation Engineer, Imaging
Anduril Industries · Boulder, Colorado, United States
Machine Learning Engineer, Driver Understanding and Evaluation
Waymo · Mountain View, CA, USA · Late
Machine Learning Engineer (Infra), Driver Understanding and Evaluation
Waymo · Mountain View, CA, USA · Late
Product Manager, New Geo Evaluation
Waymo · Mountain View, CA, US; San Francisco, CA, US · Late
Senior Machine Learning Engineer, Driver Understanding and Evaluation
Waymo · Mountain View, CA, USA · Late
Senior Machine Learning Engineer (Infra), Driver Understanding and Evaluation
Waymo · Mountain View, CA, USA · Late
Senior Machine Learning Engineer, Simulation Evaluation
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late
Senior Software Engineer, Eval Authoring APIs
Waymo · Mountain View,CA, USA; San Francisco, CA, USA; New York, NY, USA · Late
Senior Software Engineer, ML/Eval Data Platforms & Infrastructure
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late
Senior Software Engineer, ML Evaluation Infra and Efficiency
Waymo · Mountain View, California · Late
Senior Software Engineer, Planner Evaluation
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late
Senior Software Engineer, Quantitative Evaluations
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late
Senior Software Engineer, Simulator Evaluation
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late
Senior Software Engineer, Statistical Evaluation and Sampling
Waymo · Mountain View, CA, USA; San Francisco, CA, USA; New York, NY, USA · Late
Senior Staff ML Engineer, Driver Understanding and Evaluation
Waymo · Mountain View, CA, United States · Late
Senior Staff TLM, Data Mining and Sampling for ML and Evaluation
Waymo · Mountain View, California, United States; San Francisco, California, United States; New York City, New York, United States. · Late
Software Engineer, Perception Scaling Evaluation
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late
Software Engineer, Quantitative Evaluations
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late
Software Engineer, Statistical Evaluation and Sampling
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late
Software Quality Operations Specialist, Safety Evaluation
Waymo · Hyderabad, India · Late
Staff Machine Learning Engineer, Driver Understanding and Evaluation
Waymo · Mountain View, CA, USA · Late
Staff Machine Learning Engineer (Infra), Driver Understanding and Evaluation
Waymo · Mountain View, CA, USA · Late
Staff Machine Learning Engineer (TLM), Driver Understanding and Evaluation
Waymo · London, UK · Late
Staff Machine Learning Engineer – VLM/LLM Evaluation
Waymo · Mountain View, CA, USA; San Francisco, CA, USA; Kirkland, WA, USA; New York City, NY, USA · Late
Staff Software Engineer, Planner Evaluation
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late
Staff Software Engineer, Quantitative Evaluations
Waymo · Mountain View, CA, USA; San Francisco, CA, USA; New York, NY, USA · Late
Staff Software Engineer, Simulator Evaluation
Waymo · Mountain View, California, United States; San Francisco, California, United States. · Late
Tech Lead, Self Driving Eval Infrastructure
Waymo · Mountain View, CA, USA · Late
Technical Lead Manager, Prediction, ML Evaluation
Waymo · Mountain View, CA, USA; San Francisco, CA, USA · Late
Applied Scientist, Price Perception and Evaluation Science
Amazon (AI roles) · Seattle, Washington, USA · Public
Director of Platform Management for Simulation, Evaluation & Validation
Wayve · London; Sunnyvale
Software Engineer, Evals
Glean · Bangalore, India · Late
Technical Program Manager, Safeguards (Infrastructure & Evals)
Anthropic · San Francisco, CA | New York City, NY | Seattle, WA · Late · raised 2026
2026 Early Career Test & Evaluation Engineer
Anduril Industries · Costa Mesa, California, United States
Director of Engineering, Eval Platform
Nuro · Mountain View, California (HQ)
Technical Lead, Evaluation Infrastructure
Nuro · Mountain View, California (HQ)
Technical Lead Manager, Autonomy Evaluation and Intelligence
Nuro · Mountain View, California (HQ)
Engineering Manager, Agent Prompts & Evals
Anthropic · San Francisco, CA | New York City, NY · Late · raised 2026
Research Engineer, Model Evaluations
Anthropic · Remote-Friendly (Travel-Required) | San Francisco, CA | New York City, NY · Late · raised 2026Remote
Safeguards Enforcement Analyst, Safety Evaluations
Anthropic · Remote-Friendly (Travel-Required) | San Francisco, CA | Washington, DC; San Francisco, CA | New York City, NY · Late · raised 2026Remote
Staff+ Software Engineer, Safeguards Evals
Anthropic · San Francisco, CA | New York City, NY · Late · raised 2026
National Security Cyber Evaluation Lead
OpenAI · Washington, DC · Late · raised 2026Remote$252K – $342K • Offers Equity
Senior Software Engineer, Metrics and Evaluation - Autonomous Vehicles
NVIDIA · 5 Locations · Public
Research Scientist, Frontier Risk Evaluations
Scale AI · San Francisco, CA; New York, NY · Late
Director, Research - Evaluation & Training
Snorkel AI · San Francisco, CA (Hybrid) · Late
Member of Technical Staff (Data Scientist, Evals)
Perplexity · San Francisco · Late · raised 2026Remote$200K – $300K
Agent Post-Training, Frontier Evals and Environments Research
OpenAI · San Francisco · Late · raised 2026$295K – $445K
Evaluation Engineer, Applied AI
Mistral AI · Paris · Late
Machine Learning Engineer, LLM Evals & Observability
Glean · Mountain View, CA · Late
Machine Learning Engineer, LLM Evals & Observability
Glean · San Francisco, CA · Late
Senior Program Manager, Eval Operations
Nuro · Mountain View, California (HQ)
Product Lead, AI/ML (Evals)
Abridge · SF Office · Late · raised 2026Remote$250K – $290K • Offers Equity
Senior Deep Learning Engineer - Model Evaluation & AI Systems
NVIDIA · US, CA, Santa Clara · Public
Senior Research Manager, World Model Evaluation
NVIDIA · US, CA, Santa Clara · Public
AI Engineer, Evaluation
Distyl AI · San Francisco · GrowthRemote
Senior Product Operations Manager, Evaluation
Harvey · San Francisco · Late · raised 2026Remote$150K – $210K • Offers Equity • Offers Bonus
Software Engineering Manager, AI Observability & Evals Platform (New York, NY)
LangChain · New York, NY · Growth
Engineering Manager, Evals
Cursor · San Francisco · Late
Member of Technical Staff, Evals
Magic · San Francisco · Growth$200K – $550K
Applied Scientist II, GenAI Evaluation Media (GEM)
Amazon (AI roles) · Seattle, Washington, USA · Public
Principal Software Engineer, AI Observability & Evals Platform
LangChain · Boston, MA · Growth
Research Program Manager - Model Evals and Safety
Reflection AI · New York · Late · raised 2026
Member of Engineering (Evaluations)
Poolside · Remote (EMEA/East Coast) · GrowthRemote
Member of Engineering (Evaluations / Engineering)
Poolside · Remote (EMEA/East Coast) · GrowthRemote
Software Engineer, Agent Evaluation and Quality
Cursor · San Francisco · Late
Research Engineer – Benchmarking, Evals & Failure Analysis
Mercor · San Francisco · Late$130K – $500K • Offers Equity
Senior Applied Scientist, HST Health Evaluation
Amazon (AI roles) · Bengaluru, Karnataka, IND · Public
Applied Scientist II, HST Health Evaluation
Amazon (AI roles) · Bengaluru, Karnataka, IND · Public
Senior Software Engineer, Evaluation Infrastructure
Waabi · Toronto, ON
Research Scientist (Measurement and Evaluation)
Abridge · NYC Office · Late · raised 2026Remote$188K – $277K • Offers Equity
Software Engineering Manager, AI Observability & Evals Platform (San Francisco, CA)
LangChain · San Francisco, CA · Growth
Senior Backend Software Engineer, AI Observability & Evals Platform (LangSmith)
LangChain · San Francisco, CA · Growth
Backend Software Engineer (Evals)
OpenAI · San Francisco · Late · raised 2026$230K – $385K • Offers Equity
Member of Technical Staff - Evaluations
Reflection AI · San Francisco · Late · raised 2026
Member of Technical Staff, Data Analysis and Evaluation
Cohere · London · Late · raised 2026Remote
Member of Technical Staff, Evals & Post-Training Product
Fireworks AI · San Mateo · Late · raised 2026
Senior Research Scientist, Model Evaluation
Cohere · Toronto · Late · raised 2026Remote
Researcher, Evals
Cartesia · *HQ - San Francisco, CA$220K – $350K • Offers Equity
Research, Evals
Exa · San Francisco, California · Growth · raised 2026$180K – $350K • Offers Equity
FullStack Engineer, AI Observability & Evals Platform (LangSmith)
LangChain · San Francisco, CA · Growth
Research Engineer, Frontier Evals & Environments
OpenAI · San Francisco · Late · raised 2026$205K – $380K • Offers Equity
Senior Fullstack Engineer, AI Observability & Evals Platform
LangChain · San Francisco, CA · Growth
Search all 16,773 AI jobs