Operating Systems & KernelCompilers & RuntimesDatabases & StorageVirtualization & ContainersGraphics, GPU & ComputeLow-Latency & Performance

Training Performance Engineer

OpenAI · San Francisco

OpenAI is hiring a Training Performance Engineer in San Francisco. The role focuses on operating systems & kernel, compilers & runtimes, databases & storage, virtualization & containers, graphics, gpu & compute, low-latency & performance.

Key responsibilities

  • This role blends deep systems understanding with practical performance engineering — analyzing GPU kernel performance, collective communication throughput, investigating I/O bottlenecks, and sharding our models so we can train them at massive scale.
  • Collaborate with runtime and systems engineers to improve kernel efficiency, scheduling, and collective communication performance.
  • Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs.

Source and corrections

NearMetal classified and summarized this role from an official company source.

Open official source