PRF·0124
Senior Software Engineer - GPU Kernel Authoring & Optimization
Summary
CoreWeave is hiring a Senior Software Engineer - GPU Kernel Authoring & Optimization in Sunnyvale, CA / Bellevue, WA. The role focuses on operating systems & kernel, compilers & runtimes, networking, virtualization & containers, graphics, gpu & compute, low-latency & performance.
Responsibilities as stated
- You'll partner with product, orchestration, and hardware teams to turn kernel-level wins into end-to-end gains and meet strict P99 SLAs at scale.
- Author, profile, and optimize CUDA kernels—GEMMs, attention, MoE routing, quantization, KV-cache, and fused epilogues—on the critical path of LLM inference.
- Benchmark rigorously: build reproducible microbenchmarks and roofline analyses, and validate that kernel-level wins translate to end-to-end latency/throughput gains across model-serving stacks (vLLM, TensorRT-LLM, llm-d, SGLang).
Location, employment, and compensation
- Location. Sunnyvale, CA / Bellevue, WA
- Employment. unknown
- Compensation. Not stated on the source. NearMetal does not estimate compensation.
- How to apply. On the company’s own posting. NearMetal does not accept applications.
Why this is filed as systems software
- You'll partner with product, orchestration, and hardware teams to turn kernel-level wins into end-to-end gains and meet strict P99 SLAs at scale.
- Author, profile, and optimize CUDA kernels—GEMMs, attention, MoE routing, quantization, KV-cache, and fused epilogues—on the critical path of LLM inference.
- Benchmark rigorously: build reproducible microbenchmarks and roofline analyses, and validate that kernel-level wins translate to end-to-end latency/throughput gains across model-serving stacks (vLLM, TensorRT-LLM, llm-d, SGLang).
Source and corrections
This summary was written from the company’s own posting. NearMetal does not reproduce the full description and does not accept applications.