CMP·0007

Distributed LLM Inference Engineer

Anyscale · San Francisco, Palo Alto

Compilers & RuntimesDatabases & Query EnginesPerformance & ObservabilityStorage & Filesystems

Summary

Anyscale is hiring a Distributed LLM Inference Engineer in San Francisco. The role focuses on compilers & runtimes, databases & storage, graphics, gpu & compute, low-latency & performance.

Responsibilities as stated

  • Familiarity with running ML inference at large scale with high throughput and low latency.
  • With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
  • As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale.

Location, employment, and compensation

  • Location. San Francisco, Palo Alto — hybrid
  • Employment. unknown
  • Compensation. Not stated on the source. NearMetal does not estimate compensation.
  • How to apply. On the company’s own posting. NearMetal does not accept applications.

Why this is filed as systems software

  • Familiarity with running ML inference at large scale with high throughput and low latency.
  • With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
  • As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale.

Source and corrections

This summary was written from the company’s own posting. NearMetal does not reproduce the full description and does not accept applications.

Open the company posting ↗

LAST UPDATED 1h263 roles80 companiesnext check 10h263 roles · updated 1hindex status ↗