One command to find which open-source LLMs actually run on your hardware, scored for fit, speed, quality and context.
uvx llmfitllmfit inspects your CPU, RAM, GPU and VRAM, then scores hundreds of models on memory fit, speed, quality and context. You skip the trial and error of downloading models that will not load or will crawl.
uvx llmfit
Homebrew and an install script are also available.
Unsloth · Tool / CLI
Run, fine-tune and serve LLMs and diffusion models on your own hardware, with an OpenAI-compatible API for your agents.
lyogavin · Tool / CLI
Runs 70B models on a single 4 GB GPU without quantization, by keeping only one layer on the GPU at a time.
antirez · Tool / CLI
antirez's focused engine for running DeepSeek V4 Flash and a few other large open models locally on Metal, CUDA and ROCm.