Job Overview
Develop the next generation of on-device language models and Private Cloud Compute runtimes powering Apple Intelligence across 1+ billion devices.
Key Responsibilities
- Optimize transformer architectures for Apple Silicon Neural Engine (ANE) and Metal Performance Shaders.
- Design privacy-preserving Private Cloud Compute attestation protocols and sub-10ms token generators.
- Implement memory-efficient 2-bit and 4-bit quantization kernels for low-power on-device execution.
Requirements
- Strong systems programming background in C++, Swift, or Rust.
- Deep knowledge of low-level GPU/NPU acceleration, quantization, and linear algebra libraries.