Operating Systems & KernelNetworkingGraphics, GPU & ComputeLow-Latency & Performance

Software Engineer, GPU Infrastructure (HPC)

Cohere · Canada, United States

Cohere is hiring a Software Engineer, GPU Infrastructure (HPC) in Canada. The role focuses on operating systems & kernel, networking, graphics, gpu & compute, low-latency & performance.

Key responsibilities

  • Build and scale ML-optimized HPC infrastructure: Deploy and manage Kubernetes-based GPU/TPU superclusters across multiple clouds, ensuring high throughput and low-latency performance for AI workloads.
  • Low-level systems knowledge: Familiarity with Linux internals, RDMA networking, and performance optimization for ML workloads.
  • Deep expertise in ML/HPC infrastructure: Experience with GPU/TPU clusters, distributed training frameworks (JAX, PyTorch, TensorFlow), and high-performance computing (HPC) environments.

Source and corrections

NearMetal classified and summarized this role from an official company source.

Open official source