Human role

Mathematical Scientist for AI Safety Research

LawZero

Montréal, Québec, CAN builtin 1w ago
Apply now

The role

Job description

Frontier AI companies are throwing billions of dollars into scaling existing architectures and methods such as next-token-prediction, direct preference optimization (DPO), reinforcement learning with human feedback (RLHF), and reinforcement learning with verified rewards (RLVR). These methods are very powerful, yet fundamentally flawed, resulting in misalignment, sycophancy, systematic biases, and other forms of harmful behavior that are already having severely negative consequences in our socie

View full posting

Index terms

Skills & keywords

AI SafetyMidbuiltin

Get roles like this in your inbox

New agentic AI jobs, curated every Thursday. No spam.

Explore more

All AI Safety jobs
LawZeroMathematical Scientist for AI Safety Research
Apply