Skills & Keywords
indeed
Job Description
**Overview** How do you know an AI product is actually getting better, and how do you prove it, at scale, before millions of people feel the difference? We build and operate the offline evaluation platform that gates Microsoft Copilot’s quality: teams across Copilot depend on us to run their scenarios against the product, score the responses, and produce the scorecards that decide what ships. We do this at the scale of one of the world’s largest AI products, inside a strict enterprise compli
View full posting