Training Performance Engineer
OpenAI · San Francisco
OpenAI is hiring a Training Performance Engineer in San Francisco. The role focuses on operating systems & kernel, compilers & runtimes, databases & storage, virtualization & containers, graphics, gpu & compute, low-latency & performance.
Key responsibilities
- This role blends deep systems understanding with practical performance engineering — analyzing GPU kernel performance, collective communication throughput, investigating I/O bottlenecks, and sharding our models so we can train them at massive scale.
- Collaborate with runtime and systems engineers to improve kernel efficiency, scheduling, and collective communication performance.
- Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs.
Source and corrections
NearMetal classified and summarized this role from an official company source.
Open official source