Job Overview
Build ultra-scale inference engines and developer tooling powering the open-source Llama foundation model family across global datacenters.
Key Responsibilities
- Optimize FlashAttention, vLLM, and speculative decoding kernels on thousands of GPUs.
- Develop open-source PyTorch libraries and edge runtime adapters for low-latency Llama deployments.
- Lead scalability tests for next-generation 400B+ parameter multimodal architectures.
Requirements
- Strong C++/CUDA/Python skills and experience optimizing deep learning compilers.
- Deep familiarity with distributed training (FSDP, Megatron-LM) and high-throughput inference.