NetworkingGraphics, GPU & ComputeLow-Latency & Performance

Software Engineer - Model Performance

Baseten · San Francisco, Toronto, New York, Montreal

Baseten is hiring a Software Engineer - Model Performance in San Francisco. The role focuses on networking, graphics, gpu & compute, low-latency & performance.

Key responsibilities

  • Implement, refine, and productionize cutting-edge techniques (quantization, speculative decoding, kv cache reuse, chunked prefill and LoRA) for ML model inference and infrastructure.
  • Deep dive into underlying codebases of TensorRT, PyTorch, TensorRT-LLM, vllm, sglang, CUDA, and other libraries to debug ML performance issues.
  • Baseten powers mission-critical inference for the world's most dynamic AI companies, like Cursor, Notion, OpenEvidence, Abridge, Clay, Gamma and Writer.

Source and corrections

NearMetal classified and summarized this role from an official company source.

Open official source