Human role

Senior Deep Learning Architect, LLM Inference

NVIDIA

US, CA, Santa Clara workday 1d ago
Apply now

The role

Job description

We are now looking for a Senior Deep Learning Architect, LLM Inference! NVIDIA is at the forefront of the generative AI revolution. The Inference Benchmarking (IB) team specifically focuses on inference server performance optimization for Large Language Models (LLMs). If you're passionate about pushing the boundaries of GPU hardware and software performance and understand terms like disaggregated serving, data parallel attention, MoE, Qwen3.5, DeepSeek, GPT-OSS, then this is a great role for you

View full posting

Index terms

Skills & keywords

LLMGenerative AI

Get roles like this in your inbox

New agentic AI jobs, curated every Thursday. No spam.

Explore more

All LLM Engineer jobs All jobs at NVIDIA
NVIDIASenior Deep Learning Architect, LLM Inference
Apply