PRF·0017
Software Engineer, GPU Infrastructure (HPC)
Summary
Cohere is hiring a Software Engineer, GPU Infrastructure (HPC) in Canada. The role focuses on operating systems & kernel, networking, graphics, gpu & compute, low-latency & performance.
Responsibilities as stated
- Build and scale ML-optimized HPC infrastructure: Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads.
- Low-level systems knowledge: Familiarity with Linux internals, RDMA networking, and performance optimization for ML workloads.
- Deep expertise in ML/HPC infrastructure: Experience with GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and high-performance computing (HPC) environments.
Location, employment, and compensation
- Location. Canada, United States — hybrid
- Employment. unknown
- Compensation. Not stated on the source. NearMetal does not estimate compensation.
- How to apply. On the company’s own posting. NearMetal does not accept applications.
Why this is filed as systems software
- Build and scale ML-optimized HPC infrastructure: Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads.
- Low-level systems knowledge: Familiarity with Linux internals, RDMA networking, and performance optimization for ML workloads.
- Deep expertise in ML/HPC infrastructure: Experience with GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and high-performance computing (HPC) environments.
Source and corrections
This summary was written from the company’s own posting. NearMetal does not reproduce the full description and does not accept applications.