The role
Job description
Conduct performance benchmarking and production rollout; Design AI training and online inference infrastructure; Develop distributed training and serving systems; Identify and resolve AI stack performance bottlenecks; Implement LLM training and inference infrastructure; Improve latency throughput and GPU utilization; Mentor junior engineers; Optimize model performance across GPU and runtime; Productionize new model architectures and algorithms;
Index terms