Databases & StorageGraphics, GPU & ComputeLow-Latency & Performance

Inference Engineer

Cartesia · *HQ - San Francisco, CA

Cartesia is hiring a Inference Engineer in *HQ - San Francisco, CA. The role focuses on databases & storage, graphics, gpu & compute, low-latency & performance.

Key responsibilities

  • Design and build low latency, scalable, and reliable model inference and serving stack for our cutting edge foundation models using Transformers, SSMs and hybrid models.
  • Design and build robust inference infrastructure and monitoring for our products.
  • Experience building large-scale distributed systems with high demands on performance, reliability, and observability.

Source and corrections

NearMetal classified and summarized this role from an official company source.

Open official source