Member of Technical Staff (AI Inference Engineer)
Perplexity · San Francisco, Palo Alto, New York City
Perplexity is hiring a Member of Technical Staff (AI Inference Engineer) in San Francisco. The role focuses on operating systems & kernel, compilers & runtimes, databases & storage, networking, virtualization & containers, graphics, gpu & compute, low-latency & performance.
Key responsibilities
- Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.
- We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets.
- Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving.
Source and corrections
NearMetal’s Agent classified and summarized this role from an official company source. Human review is shown separately when present.
Open official source