A high-performance serving framework for LLMs and multimodal models, from a single GPU to large distributed clusters.
SGLang serves language and multimodal models with low latency and high throughput. The post quotes a production footprint of 400,000 GPUs and trillions of tokens a day.
Follow the installation guide linked from the README for your GPU and model.
vLLM project · Tool / CLI
A high-throughput, memory-efficient engine for serving LLMs, with PagedAttention, continuous batching and prefix caching.
lyogavin · Tool / CLI
Runs 70B models on a single 4 GB GPU without quantization, by keeping only one layer on the GPU at a time.
Modular · Tool / CLI
Modular's open platform for AI development and deployment: the MAX framework, the Mojo language and an OpenAI-compatible server.