Member of Technical Staff - Research, Inference
Modal · New York, San Francisco
Modal is hiring a Member of Technical Staff - Research, Inference in New York. The role focuses on databases & storage, virtualization & containers, graphics, gpu & compute, low-latency & performance.
Key responsibilities
- They rely on Modal for instant GPU access, sub-second container starts, and native storage, so it's simple to serve low-latency inference, fine-tune models, and access production-ready sandboxes at scale.
- Own end-to-end inference research bets: speculative decoding, disaggregated prefill/decode, quantization (FP8, INT4), KV-cache and memory management, autoscaling for spiky serverless traffic, and whatever else the research agenda calls for.
- We already run elastic inference, sandboxes, distributed volumes, and multi-node training, and we control the infrastructure underneath, so the serving stack is ours to shape rather than something we resell.
Source and corrections
NearMetal classified and summarized this role from an official company source.
Open official source