The role
Job description
NVIDIA is looking for an Engineering Manager to lead the team responsible for deploying and serving Large Language Models (LLMs) and Vision-Language Models (VLMs) at scale. Our team builds and operates an AI inference platform that enables customers to deploy and run brand new generative AI models efficiently across NVIDIA GPU platforms. The platform operates at the intersection of model optimization, inference systems, distributed computing, and production infrastructure. In this role, you will
Index terms