Qortora · Search · Indexed page

vllm.aiFetched 2026-08-15T04:48:41Z

vLLM

vLLM is a high-throughput and memory-efficient inference and serving engine for Large Language Models (LLMs). Deploy AI models faster with state-of-the-art performance. Easy, fast, and cost-efficient LLM serving for everyone.

Open original source · Full cached text

vLLM Menu Theme The High-Throughput and Memory-Efficient inference and serving engine for LLMs Easy, fast, and cost-efficient LLM serving for everyone. Get StartedDocumentation Easy Deploy the widest range of open-source models on any hardware. Includes a drop-in OpenAI-compatible API for instant integration. Fast Maximize throughput with PagedAttention. Advanced scheduling and continuous batching ensure peak GPU utilization. Cost Efficient Slash inference costs by maximizing hardware efficiency. We make high-performance LLMs affordable and accessible to everyone. Quick Start Select your preferences and run the install command. Stable represents the most currently tested and supported version of vLLM. Nightly is available if you want the latest builds. 📦 Requires Python 3.10+. Python 3.12+ recommended. ⚡ We recommend uv for faster and more reliable installation. 🔧 For other platforms, see docs.vllm.ai 🎉 See what's new in 🔍 Find which release contains a PR BuildStableNightly PlatformCUDAROCmXPUCPU PackagePython (uv)PythonDocker CUDA VersionCUDA 13.0CUDA 12.9 Run this Command: uv pip install vllm --torch-backend auto 💡 Compatible with all CUDA 13.x versions (13.0 - 13.1) · Troubleshooting Looking for older versions? Sponsors vLLM is a community project. Our compute resources for development and testing are supported by the following organizations. Thank you for your support! Cash Donations a16z Sequoia Capital Skywork AI ZhenFund Compute Resources Alibaba Cloud AMD Anyscale AWS Crusoe Cloud Google Cloud IBM Intel Lambda Lab Nebius Novita AI NVIDIA Red Hat Roblox RunPod UC Berkeley Slack Sponsor Inferact — Stars— ⭐— Contributors— 👥PyTorch Foundation We collect donation through GitHub and OpenCollective. We plan to use the fund to support the development, maintenance, and adoption of vLLM. Universal Compatibility One engine, endless possibilities. Run any model on any hardware. Hardware Unified API across platforms NVIDIACUDA GPU Popular AMDROCm GPU HuaweiAscend NPU AWSNeuron Accelerator GoogleCloud TPU IBMSpyre Accelerator IntelGaudiXPUCPU AppleApple Silicon BaiduKunlun XPU CambriconMLU View all supported hardware Open Models Latest trending open-source models, optimized & production-ready DeepSeek DeepSeek V4DeepSeek V3.2DeepSeek R1 Google Gemma 4Gemma 3 Meta Muse GlimmerLlama 4 ScoutLlama 4 Maverick Minimax MiniMax M3MiniMax M2.7MiniMax M2.5 Mistral AI Mistral Small 4Mistral Large 3 MoonshotAI Kimi K3Kimi K2.6Kimi K2.5 NVIDIA Nemotron 3 UltraNemotron 3 SuperNemotron 3 Nano Qwen Qwen3.8Qwen3.6Qwen3.5 StepFun Step-3.7-FlashStep-3.5-Flash Z-AI GLM 5.2GLM 5.1GLM 5 View all supported models Everyone welcome! Got questions? We're here to help. Whether you're just getting started or debugging a complex deployment, our community is open to everyone. No question is too basic! Fast & friendly responses Active maintainers Join Slack Real-time help & discussions Visit Forum Searchable Q&A knowledge base GitHub Issues Bug reports & feature requests Resources Explore recipes, benchmarks, and roadmap Recipes Example notebooks and tutorials recipes.vllm.ai Performance Benchmarks and comparisons perf.vllm.ai Roadmap Project roadmap and milestones roadmap.vllm.ai