The role
Job description
Assess model trustworthiness of evaluators; Build evaluation pipeline for AI agents; Collect eval data from production traces; Construct automated eval loop for every change; Design task suites that reflect real distribution; Design verifier guided selection and escalation strategies; Prevent task suite gaming; Read frontier research and convert into live experiments; Run automated hill climbing for prompt context tools and routing; Run test time scaling experiments; Set noise floors and accepta
Index terms