Site Reliability Engineer, Inference Infrastructure
Cohere · Toronto, San Francisco, London, New York, Montreal
Cohere is hiring a Site Reliability Engineer, Inference Infrastructure in Toronto. The role focuses on operating systems & kernel, databases & storage, networking, graphics, gpu & compute, low-latency & performance.
Key responsibilities
- Experience designing large, highly available distributed systems with Kubernetes, and GPU workloads on those clusters.
- Experience in compute/storage/network resource and cost management.
- In this role, you will work closely with many teams to deploy optimized NLP models to production in low latency, high throughput, and high availability environments.
Source and corrections
NearMetal classified and summarized this role from an official company source.
Open official source