KRN·0218
Member of Technical Staff (AI Inference Engineer)
Summary
Perplexity is hiring a Member of Technical Staff (AI Inference Engineer) in San Francisco. The role focuses on operating systems & kernel, compilers & runtimes, databases & storage, networking, virtualization & containers, graphics, gpu & compute, low-latency & performance.
Responsibilities as stated
- Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.
- We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets.
- Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving.
Location, employment, and compensation
- Location. San Francisco, Palo Alto, New York City
- Employment. unknown
- Compensation. Not stated on the source. NearMetal does not estimate compensation.
- How to apply. On the company’s own posting. NearMetal does not accept applications.
Why this is filed as systems software
- Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic.
- We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets.
- Profile and fix bottlenecks from network ingress through continuous batching and GPU kernel interleaving.
Source and corrections
This summary was written from the company’s own posting. NearMetal does not reproduce the full description and does not accept applications.