vllm is by vllm-project — not by us.
We indexed this project so people can find it. We are not selling it, we host no copy of the code, and we are not affiliated with or endorsed by its authors — get it from them.
Indexed Aug 3, 2026 · 88,047 stars at index time. Maintainers: claiming verifies your identity and unlocks a higher assurance tier. Removal requests are honored.
Application · for humans & their agents
vllm
unclaimed listingactively maintainedFreeApache-2.0
A high-throughput and memory-efficient inference and serving engine for LLMs
by Open Source Community · New publisher
Every claim on this page is refundable if it is untrue — refund policy.
vllm-project/vllm is an open-source project by vllm-project: A high-throughput and memory-efficient inference and serving engine for LLMs. Indexed here so it can be found — not resold.
It is free. Get it from the upstream repository: https://github.com/vllm-project/vllm
From the project's own README (excerpt, reproduced for discovery under its Apache-2.0 license):
Easy, fast, and cheap LLM serving for everyone
Documentation Blog Paper Twitter/X User Forum Developer Slack
🔥 We have built a vLLM website to help you get started with vLLM. Please visit vllm.ai to learn more. For events, please visit vllm.ai/events to join us.
---
About
vLLM is a fast and easy-to-use library for LLM inference and serving.
Originally developed in the Sky Computing Lab at UC Berkeley, vLLM has grown into one of the most active open-source AI projects built and maintained by a diverse community of many dozens of academic institutions and companies from over 2000 contributors.
vLLM is fast with: • State-of-the-art serving throughput • Efficient management of attention key and value memory with PagedAttention • Continuous batching of incoming requests, chunked prefill, prefix caching • Fast and flexible model execution with piecewise and full CUDA/HIP graphs • Quantization: FP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ, GGUF, compressed-tensors, ModelOpt, TorchAO, and more • Optimized attention kernels including FlashAttention, FlashInfer, TRTLLM-GEN, FlashMLA, and Triton • Optimized GEMM/MoE kernels for various precisions using CUTLASS, TRTLLM-GEN, CuTeDSL • Speculative decoding including n-gram, suffix, EAGLE, DFlash • Automatic kernel generation and graph-level transformations using torch.compile • Disaggregated prefill, decode, and encode
vLLM is flexible and easy to use with: • Seamless integration with popular Hugging Face models • High-throughput serving with various decoding algorithms, including parallel sampling, beam search, and more • Tensor, pipeline, data, expert, and context parallelism for distributed inference • Streaming outputs • Generation of structured outputs using xgrammar or guidance • Tool calling and reasoning…
Preview
See what it does before you commit. Previews show behavior, never source code.
01Capabilities
Does
- + Model usage tracking
- + Developer tooling
- + Resource exhaustion limits
Doesn’t
- No exclusions declared
02Requirements & stack
Depends on
No declared dependencies
Credentials needed
None declared
Stack
03Community
No endorsements yetNo verified confirmations yet — be the first.
Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.
Sign in to confirm — weight comes from verified usage, not vote count.
Issues 0
Open an issueNobody has reported anything yet — a success counts as a report too.
04Trust Passport
Full passport →0/0 automated components pass. An automated score is never a security guarantee.
This listing is unclaimed, so publisher identity cannot be verified and it stays below the “verified” tier by design — that is a statement about the listing, not about the project’s quality. Our automated scans still ran; a maintainer who claims it unlocks identity verification.
- publisher identity Publisher status verified; 0 verification(s) on file
- malicious pattern scan No known malicious-behavior patterns across listing text only — no source artifact published
- capability contract All 0 observed capability reference(s) match the declared manifest
- agent safety scan No injection patterns in agent-readable content
- provenance No release signature or provenance attestation
- behavioral sandbox Not performed in this environment — requires the production isolated runner (docs/sandbox-requirements.md). No untrusted code is ever executed on the application host.
Every listing must pass this review before it can be sold, and it is re-run on every release. Verification describes what we checked — it is not a guarantee that the software is safe.
05Versions
Full history →| Version | Channel | Released | Notes |
|---|---|---|---|
| 0.0.0 | stable | Aug 3, 2026 | Indexed listing — see the upstream repository for real release history. |