PRF·0174

Software Engineer - Model Performance

Baseten · San Francisco, Toronto, New York, Montreal

Networking & ProtocolsPerformance & ObservabilityC++

Summary

Baseten is hiring a Software Engineer - Model Performance in San Francisco. The role focuses on networking, graphics, gpu & compute, low-latency & performance.

Responsibilities as stated

  • Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure.
  • Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues.
  • Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer.

Location, employment, and compensation

  • Location. San Francisco, Toronto, New York, Montreal — hybrid
  • Employment. unknown
  • Compensation. Not stated on the source. NearMetal does not estimate compensation.
  • How to apply. On the company’s own posting. NearMetal does not accept applications.

Why this is filed as systems software

  • Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure.
  • Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues.
  • Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer.

Source and corrections

This summary was written from the company’s own posting. NearMetal does not reproduce the full description and does not accept applications.

Open the company posting ↗

LAST UPDATED 2h263 roles80 companiesnext check 10h263 roles · updated 2hindex status ↗