CMP·0007
Distributed LLM Inference Engineer
Summary
Anyscale is hiring a Distributed LLM Inference Engineer in San Francisco. The role focuses on compilers & runtimes, databases & storage, graphics, gpu & compute, low-latency & performance.
Responsibilities as stated
- Familiarity with running ML inference at large scale with high throughput and low latency.
- With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
- As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale.
Location, employment, and compensation
- Location. San Francisco, Palo Alto — hybrid
- Employment. unknown
- Compensation. Not stated on the source. NearMetal does not estimate compensation.
- How to apply. On the company’s own posting. NearMetal does not accept applications.
Why this is filed as systems software
- Familiarity with running ML inference at large scale with high throughput and low latency.
- With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
- As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale.
Source and corrections
This summary was written from the company’s own posting. NearMetal does not reproduce the full description and does not accept applications.