PRF·0053
Training Performance Engineer
Summary
OpenAI is hiring a Training Performance Engineer in San Francisco. The role focuses on operating systems & kernel, compilers & runtimes, databases & storage, virtualization & containers, graphics, gpu & compute, low-latency & performance.
Responsibilities as stated
- This role blends deep systems understanding with practical performance engineering — analyzing GPU kernel performance, collective communication throughput, investigating I/O bottlenecks, and sharding our models so we can train them at massive scale.
- Collaborate with runtime and systems engineers to improve kernel efficiency, scheduling, and collective communication performance.
- Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs.
Location, employment, and compensation
- Location. San Francisco — hybrid
- Employment. unknown
- Compensation. Not stated on the source. NearMetal does not estimate compensation.
- How to apply. On the company’s own posting. NearMetal does not accept applications.
Why this is filed as systems software
- This role blends deep systems understanding with practical performance engineering — analyzing GPU kernel performance, collective communication throughput, investigating I/O bottlenecks, and sharding our models so we can train them at massive scale.
- Collaborate with runtime and systems engineers to improve kernel efficiency, scheduling, and collective communication performance.
- Training Runtime designs the core distributed machine-learning training runtime that powers everything from early research experiments to frontier-scale model runs.
Source and corrections
This summary was written from the company’s own posting. NearMetal does not reproduce the full description and does not accept applications.