The role
Job description
Frontier AI companies are throwing billions of dollars into scaling existing architectures and methods such as next-token-prediction, direct preference optimization (DPO), reinforcement learning with human feedback (RLHF), and reinforcement learning with verified rewards (RLVR). These methods are very powerful, yet fundamentally flawed, resulting in misalignment, sycophancy, systematic biases, and other forms of harmful behavior that are already having severely negative consequences in our socie
Index terms