Staff Test Engineer - AI
getro
Job Description
- Own the AI Quality Strategy. Define and lead the end-to-end testing strategy for Outreach’s GenAI platform, including agentic workflows, LLM tool calls, LangGraph orchestration, and supporting ML pipelines.
- Build Evaluation Frameworks. Design and implement evaluation systems that handle both deterministic and non-deterministic outputs — combining rule-based assertions, golden dataset testing, and LLM-as-Judge approaches to grade agent responses at scale.
- Test Agents End-to-End. Own testing across Outreach’s suite of AI agents — Revenue Agent, Research Agent, Meeting Agent, Personalisation Agent, and Ask Outreach — covering functional correctness, tool selection accuracy, context handling, and response quality.
- Partner with DS and Engineering. Work closely with Data Science, MLOps, and platform engineers to ensure testability is designed in from the start — not bolted on after.
- Drive CI/CD for AI. Integrate evaluation pipelines into CI/CD workflows so that regressions in agent behavior are caught before they reach production.
- Define Quality Metrics. Establish and track metrics that matter for AI systems: answer quality scores, tool invocation accuracy, hallucination rates, latency, and regression trends over model and prompt changes.
- Champion Best Practices. Define standards for AI testing across the org — including prompt regression testing, retrieval quality evaluation, and agent behavior contracts.
- Mentor and Influence. Raise the quality bar across engineering teams by mentoring engineers, reviewing designs for testability, and advocating for quality-driven development practices.
- Stay Current. Actively track developments in AI evaluation tooling, LLM benchmarking, and testing research — and bring relevant advances into our practice.