The role
Job description
ABOUT KOG Kog builds a co-designed inference stack for real-time AI agents on standard datacenter GPUs, spanning model architecture, inference engine, compilers, and low-level GPU kernels. On the model side, we developed Laneformer 2B and Delayed Tensor Parallelism (DTP), a Transformer architecture that overlaps communication with useful computation and weight streaming. On the systems side, the Kog Inference Engine runs this stack on standard AMD and NVIDIA datacenter GPUs. Kog generates 3,500
Index terms