llama.cpp is by ggml-org — not by us.
We indexed this project so people can find it. We are not selling it, we host no copy of the code, and we are not affiliated with or endorsed by its authors — get it from them.
Indexed Aug 3, 2026 · 122,530 stars at index time. Maintainers: claiming verifies your identity and unlocks a higher assurance tier. Removal requests are honored.
Application · for humans & their agents
llama.cpp
unclaimed listingactively maintainedFreeMIT
LLM inference in C/C++
by Open Source Community · New publisher
Every claim on this page is refundable if it is untrue — refund policy.
ggml-org/llama.cpp is an open-source project by ggml-org: LLM inference in C/C++. Indexed here so it can be found — not resold.
It is free. Get it from the upstream repository: https://github.com/ggml-org/llama.cpp
From the project's own README (excerpt, reproduced for discovery under its MIT license):
llama.cpp
LLM inference in C/C++
manifesto / ggml / ops / maintainer PRs%20sort%3Aupdated-desc) / dev branches / compile times / lib llama API / llama-server REST API
Quick start
A few options to get llama.cpp installed on your machine: • Visit https://llama.app and follow the instructions • Run with Docker - see our Docker documentation • Download pre-built binaries from the releases page • Build from source by cloning this repository - check out our build guide
Once installed:
VLM session with llama cli Built-in web UI against llama serve
Description
The main goal of llama.cpp is to enable LLM (and VLM) inference with minimal setup and state-of-the-art performance on a wide range of hardware - locally and in the cloud. • Plain C/C++ implementation without any dependencies • Apple silicon is a first-class citizen - optimized via ARM NEON, Accelerate and Metal frameworks • AVX, AVX2, AVX512 and AMX support for x86 architectures • RVV, ZVFH, ZFH, ZICBOP and ZIHINTPAUSE support for RISC-V architectures • 1.5-bit, 2-bit, 3-bit, 4-bit, 5-bit, 6-bit, and 8-bit integer quantization for faster inference and reduced memory use • Custom CUDA kernels for running LLMs on NVIDIA GPUs (support for AMD GPUs via HIP and Moore Threads GPUs via MUSA) • Vulkan and SYCL backend support • CPU+GPU hybrid inference to partially accelerate models larger than the total VRAM capacity
The llama.cpp project is build on top of the ggml library.
Supported backends
Backend Target devices --- --- BLAS All BLIS All CANN Ascend NPU CUDA Nvidia GPU HIP AMD GPU [Hexagon [In Progress]](docs/backend/snapdragon/README.md) Snapdragon IBM zDNN IBM Z & LinuxONE MUSA Moore Threads GPU Metal Apple Silicon OpenCL Adreno GPU…
Preview
See what it does before you commit. Previews show behavior, never source code.
01Capabilities
Does
- + Context budgeting
- + Developer tooling
- + Resource exhaustion limits
Doesn’t
- No exclusions declared
02Requirements & stack
Depends on
No declared dependencies
Credentials needed
None declared
Stack
03Community
No endorsements yetNo verified confirmations yet — be the first.
Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.
Sign in to confirm — weight comes from verified usage, not vote count.
Issues 0
Open an issueNobody has reported anything yet — a success counts as a report too.
04Trust Passport
Full passport →0/0 automated components pass. An automated score is never a security guarantee.
This listing is unclaimed, so publisher identity cannot be verified and it stays below the “verified” tier by design — that is a statement about the listing, not about the project’s quality. Our automated scans still ran; a maintainer who claims it unlocks identity verification.
- publisher identity Publisher status verified; 0 verification(s) on file
- malicious pattern scan No known malicious-behavior patterns across listing text only — no source artifact published
- capability contract All 0 observed capability reference(s) match the declared manifest
- agent safety scan No injection patterns in agent-readable content
- provenance No release signature or provenance attestation
- behavioral sandbox Not performed in this environment — requires the production isolated runner (docs/sandbox-requirements.md). No untrusted code is ever executed on the application host.
Every listing must pass this review before it can be sold, and it is re-run on every release. Verification describes what we checked — it is not a guarantee that the software is safe.
05Versions
Full history →| Version | Channel | Released | Notes |
|---|---|---|---|
| 0.0.0 | stable | Aug 3, 2026 | Indexed listing — see the upstream repository for real release history. |