Skills & Keywords
Agent systemsCache CompressionDPODistillationHybrid AttentionInference OptimizationInstruction TuningKV Cache CompressionKV cacheKernel optimizationLlama.cppLoRA
Job Description
Create experimental evaluation pipelines; Develop KV cache compression approaches; Evaluate compression and pruning methods; Implement backpropagation free optimization methods; Implement gradient free model merging methods; Optimize efficient LLM inference performance; Prototype AI model compression approaches; Prototype speculative decoding for low latency generation; Publish technical reports and IP disclosures; Research model efficiency methods;
View full posting