Job Overview
Build, optimize, and scale Copilot runtime services and Azure AI Studio infrastructure for millions of enterprise developers worldwide.
Key Responsibilities
- Design low-latency orchestration for small language models (Phi series) and large foundation models.
- Implement high-throughput caching and fine-tuning pipelines using Semantic Kernel and ONNX Runtime.
- Collaborate with OpenAI research teams on model alignment and multi-tenant security boundaries.
Requirements
- 5+ years software engineering experience with C#, Python, or Go.
- Expertise in distributed systems, vector search, and cloud-native Kubernetes/Azure microservices.