Skip to content
Code Recycle
Free · open sourceApache-2.0Unclaimed listing

vllm is by vllm-project — not by us.

We indexed this project so people can find it. We are not selling it, we host no copy of the code, and we are not affiliated with or endorsed by its authors — get it from them.

Indexed Aug 3, 2026 · 88,047 stars at index time. Maintainers: claiming verifies your identity and unlocks a higher assurance tier. Removal requests are honored.

Application · for humans & their agents

vllm

unclaimed listingactively maintainedFreeApache-2.0

A high-throughput and memory-efficient inference and serving engine for LLMs

by Open Source Community · New publisher

Go to the project ↗Open live demo ↗

Every claim on this page is refundable if it is untrue — refund policy.

vllm.ai

vllm-project/vllm is an open-source project by vllm-project: A high-throughput and memory-efficient inference and serving engine for LLMs. Indexed here so it can be found — not resold.

It is free. Get it from the upstream repository: https://github.com/vllm-project/vllm

From the project's own README (excerpt, reproduced for discovery under its Apache-2.0 license):

Easy, fast, and cheap LLM serving for everyone

Documentation Blog Paper Twitter/X User Forum Developer Slack

🔥 We have built a vLLM website to help you get started with vLLM. Please visit vllm.ai to learn more. For events, please visit vllm.ai/events to join us.

---

About

vLLM is a fast and easy-to-use library for LLM inference and serving.

Originally developed in the Sky Computing Lab at UC Berkeley, vLLM has grown into one of the most active open-source AI projects built and maintained by a diverse community of many dozens of academic institutions and companies from over 2000 contributors.

vLLM is fast with: • State-of-the-art serving throughput • Efficient management of attention key and value memory with PagedAttention • Continuous batching of incoming requests, chunked prefill, prefix caching • Fast and flexible model execution with piecewise and full CUDA/HIP graphs • Quantization: FP8, MXFP8/MXFP4, NVFP4, INT8, INT4, GPTQ/AWQ, GGUF, compressed-tensors, ModelOpt, TorchAO, and more • Optimized attention kernels including FlashAttention, FlashInfer, TRTLLM-GEN, FlashMLA, and Triton • Optimized GEMM/MoE kernels for various precisions using CUTLASS, TRTLLM-GEN, CuTeDSL • Speculative decoding including n-gram, suffix, EAGLE, DFlash • Automatic kernel generation and graph-level transformations using torch.compile • Disaggregated prefill, decode, and encode

vLLM is flexible and easy to use with: • Seamless integration with popular Hugging Face models • High-throughput serving with various decoding algorithms, including parallel sampling, beam search, and more • Tensor, pipeline, data, expert, and context parallelism for distributed inference • Streaming outputs • Generation of structured outputs using xgrammar or guidance • Tool calling and reasoning…

Preview

See what it does before you commit. Previews show behavior, never source code.

01Capabilities

Does

  • + Model usage tracking
  • + Developer tooling
  • + Resource exhaustion limits

Doesn’t

  • No exclusions declared

02Requirements & stack

Depends on

No declared dependencies

Credentials needed

None declared

Stack

python

03Community

No endorsements yet

No verified confirmations yet — be the first.

Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.

Open an issue

Sign in to confirm — weight comes from verified usage, not vote count.

Nobody has reported anything yet — a success counts as a report too.

04Trust Passport

Full passport →
–/100

0/0 automated components pass. An automated score is never a security guarantee.

This listing is unclaimed, so publisher identity cannot be verified and it stays below the “verified” tier by design — that is a statement about the listing, not about the project’s quality. Our automated scans still ran; a maintainer who claims it unlocks identity verification.

Unclaimed · identity unverifiedreviewed Aug 12, 2026 · re-verification due Nov 10, 2026
  • publisher identity Publisher status verified; 0 verification(s) on file
  • malicious pattern scan No known malicious-behavior patterns across listing text only — no source artifact published
  • capability contract All 0 observed capability reference(s) match the declared manifest
  • agent safety scan No injection patterns in agent-readable content
  • provenance No release signature or provenance attestation
  • behavioral sandbox Not performed in this environment — requires the production isolated runner (docs/sandbox-requirements.md). No untrusted code is ever executed on the application host.

Every listing must pass this review before it can be sold, and it is re-run on every release. Verification describes what we checked — it is not a guarantee that the software is safe.

VersionChannelReleasedNotes
0.0.0stableAug 3, 2026Indexed listing — see the upstream repository for real release history.