Skills & Keywords
LLM
Job Description
Location: In person Experience: • 4+ years • GPU: 3 years Key Skills: • vLLM • TensorRT-LLM • Triton Inference Server • CUDA • cuDNN • NCCL • GPU Performance Engineering • Multi-GPU Architecture • Model Quantization • INT8 • FP16 • GPTQ • AWQ • AWS GPU Instances • AWS Inferentia • Performance Profiling • Nsight • DCGM • nvidia-smi • Kubernetes • KServe • Docker • NVIDIA Triton • Neuron SDK • Prometheus • Locust Nice to have: • Kubernetes • KServe • Docker Other: • Tech stack includes vLLM,...
View full posting