Why This Job is Featured on The SaaS Jobs
In SaaS, agentic AI features are increasingly shipped as product capabilities rather than standalone research projects, which raises the bar for measurable quality. This AI Evaluation Engineer role sits in that critical layer between model behavior and customer-facing reliability, focusing on how voice and chat agents are assessed before release. The emphasis on LLM-judge metrics, benchmark design, and release-readiness decisions reflects a maturing approach to AI within a production SaaS environment.
For a long-term SaaS career, evaluation work builds durable leverage because it connects experimentation to outcomes through repeatable measurement. Ownership of regression suites, A/B comparisons, and red teaming develops an operator’s view of model risk, iteration cadence, and quality gates that many SaaS teams struggle to formalize. The tooling and pipeline components mentioned also translate well across SaaS organizations adopting LLMs, where scalable evaluation becomes a shared platform need.
This role tends to fit professionals who prefer structured problem-solving and evidence-driven decision support over feature delivery. It also aligns with engineers comfortable collaborating across applied science, product QA, and engineering, and who enjoy turning ambiguous model failures into actionable test coverage and clear escalation paths.
The section above is editorial commentary from The SaaS Jobs, provided to help SaaS professionals understand the role in a broader industry context.
Job Description
Your role
As an AI Evaluation Engineer, you'll be an integral part of our AI Evaluation team, owning evaluation coverage for Dialpad's Agentic AI systems alongside our existing evaluation lead. A key focus will be co-owning LLM-judge metric development and calibration, scenario and benchmark dataset curation, and structured error analysis to support release-readiness decisions for our agentic voice and chat solutions.
This position reports to the manager of the AI Evaluation team and has the opportunity to be based in our Vancouver office.
What you’ll do
- You will design and execute validation strategies for agentic, NLP, and speech workflows across staging, beta, and release candidates.
- You will build, run, and improve regression evaluations, A/B comparisons, and red teaming analyses to determine whether product and model changes are ready to move forward.
- You will co-own LLM-judge metric development, calibration, and prompt refinement across evaluation dimensions.
- You will create, configure, and monitor data annotation jobs to keep evaluation and calibration datasets fed on schedule.
- You will develop and maintain QA tooling, notebooks, and pipeline components that make recurring evaluations scalable and reusable across teams.
- You will investigate bugs, triage issues, and decide whether problems should become engineering escalations, test set additions, or follow-up analysis.
- You will collaborate with cross-functional teams, including applied science, engineering, and Product QA.
Skills you’ll bring
- Bachelor's or Master's degree in Computer Science, Software Engineering, Computational Linguistics, or a related field.
- 3+ years of experience in QA, test engineering, model evaluation, or applied ML quality for AI-driven products.
- Experience designing structured test strategies across manual and automated workflows.
- Comfort working with complex AI systems such as speech, NLP, LLM, or agentic products.
- Experience working with evaluation datasets, gold sets, adversarial test sets, or benchmark creation for AI systems.
- Strong analytical skills for investigating failures, comparing outputs, and identifying actionable quality patterns.
- Experience collaborating with cross-functional technical teams and communicating clearly through documentation and reporting.