Skills & Keywords
Distributed TrainingFine TuningHuman FeedbackLanguage ModelsLarge Language ModelsLearning from Human FeedbackModel EvaluationPreference optimizationPyTorchPythonReinforcement LearningReinforcement Learning from Human Feedback
Job Description
Advance post training techniques; Collaborate with cross functional teams to ship to production; Design evaluations and benchmarks; Improve model performance efficiency and scalability; Research and build agentic AI systems; Track frontier research and publish findings;
View full posting