Compilers & RuntimesDatabases & StorageGraphics, GPU & ComputeLow-Latency & Performance

Distributed LLM Inference Engineer

Anyscale · San Francisco, Palo Alto

Anyscale is hiring a Distributed LLM Inference Engineer in San Francisco. The role focuses on compilers & runtimes, databases & storage, graphics, gpu & compute, low-latency & performance.

Key responsibilities

  • Familiarity with running ML inference at large scale with high throughput and low latency.
  • With Anyscale, we’re building the best place to run Ray, so that any developer or data scientist can scale an ML application from their laptop to the cluster without needing to be a distributed systems expert.
  • As a Distributed LLM Inference Engineer, you will help systems and optimizations that push the boundaries of performance for inference at large scale.

Source and corrections

NearMetal classified and summarized this role from an official company source.

Open official source