Human role

Engineering Manager, LLM Inference & Deployment at Scale

NVIDIA

US, CA, Santa Clara workday 1d ago
Apply now

The role

Job description

NVIDIA is looking for an Engineering Manager to lead the team responsible for deploying and serving Large Language Models (LLMs) and Vision-Language Models (VLMs) at scale. Our team builds and operates an AI inference platform that enables customers to deploy and run brand new generative AI models efficiently across NVIDIA GPU platforms. The platform operates at the intersection of model optimization, inference systems, distributed computing, and production infrastructure. In this role, you will

View full posting

Index terms

Skills & keywords

LLMGenerative AI

Get roles like this in your inbox

New agentic AI jobs, curated every Thursday. No spam.

Explore more

All LLM Engineer jobs All jobs at NVIDIA
NVIDIAEngineering Manager, LLM Inference & Deployment at Scale
Apply