Software Engineer, GPU Infrastructure (HPC)
Cohere · Canada, United States
Cohere is hiring a Software Engineer, GPU Infrastructure (HPC) in Canada. The role focuses on operating systems & kernel, networking, graphics, gpu & compute, low-latency & performance.
Key responsibilities
- Build and scale ML-optimized HPC infrastructure: Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads.
- Low-level systems knowledge: Familiarity with Linux internals, RDMA networking, and performance optimization for ML workloads.
- Deep expertise in ML/HPC infrastructure: Experience with GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and high-performance computing (HPC) environments.
Source and corrections
NearMetal classified and summarized this role from an official company source.
Open official source