Skip to content
Code Recycle
Free · open sourceApache-2.0Unclaimed listing

PaddleOCR is by PaddlePaddle — not by us.

We indexed this project so people can find it. We are not selling it, we host no copy of the code, and we are not affiliated with or endorsed by its authors — get it from them.

Indexed Sep 21, 2026 · 89,930 stars at index time. Maintainers: claiming verifies your identity and unlocks a higher assurance tier. Removal requests are honored.

Application · for humans & their agents

PaddleOCR

unclaimed listingactively maintainedFreeApache-2.0

Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the ga

by Open Source Community · New publisher

Go to the project ↗Open live demo ↗

Every claim on this page is refundable if it is untrue — refund policy.

www.paddleocr.com

PaddlePaddle/PaddleOCR is an open-source project by PaddlePaddle: Turn any PDF or image document into structured data for your AI. A powerful, lightweight OCR toolkit that bridges the gap between images/PDFs and LLMs. Supports 100+ languages.. Indexed here so it can be found — not resold.

It is free. Get it from the upstream repository: https://github.com/PaddlePaddle/PaddleOCR

From the project's own README (excerpt, reproduced for discovery under its Apache-2.0 license):

Global Leading OCR Toolkit & Document AI Engine

English 简体中文 繁體中文 日本語 한국어 Français Русский Español العربية

PaddleOCR converts PDF documents and images into structured, LLM-ready data (JSON/Markdown) with industry-leading accuracy. With 70k+ Stars and trusted by top-tier projects like Dify, RAGFlow, and Cherry Studio, PaddleOCR is the bedrock for building intelligent RAG and Agentic applications.

🚀 Key Features

📄 Intelligent Document Parsing (LLM-Ready) > Transforming messy visuals into structured data for the LLM era.

SOTA Document VLM: Featuring PaddleOCR-VL-1.6 (0.9B), the industry's leading lightweight vision-language model for document parsing. It achieves 96.3% accuracy on OmniDocBench v1.6, leads in text, formula, and table recognition, and shows significantly enhanced capabilities in ancient documents, rare characters, seals, and charts, with structured outputs in Markdown and JSON formats. Structure-Aware Conversion: Powered by PP-StructureV3, seamlessly convert complex PDFs and images into Markdown or JSON. Unlike the PaddleOCR-VL series models, it provides more fine-grained coordinate information, including table cell coordinates, text coordinates, and more. Production-Ready Efficiency: Achieve commercial-grade accuracy with an ultra-small footprint. Outperforms numerous closed-source solutions in public benchmarks while remaining resource-efficient for edge/cloud deployment.

🔍 Universal Text Recognition (Scene OCR) > The global gold standard for high-speed, multilingual text spotting.

100+ Languages Supported: Native recognition for a vast global library. PP-OCRv6 supports 50 languages with a single unified model (Chinese, English, Japanese, and 46 Latin-script languages) — no model switching needed for multilingual…

Preview

See what it does before you commit. Previews show behavior, never source code.

01Capabilities

Does

  • No capabilities recorded

Doesn’t

  • No exclusions declared

02Requirements & stack

Depends on

No declared dependencies

Credentials needed

None declared

Stack

python

03Community

No endorsements yet

No verified confirmations yet — be the first.

Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.

Open an issue

Sign in to confirm — weight comes from verified usage, not vote count.

Nobody has reported anything yet — a success counts as a report too.

04Trust Passport

Full passport →
–/100

0/0 automated components pass. An automated score is never a security guarantee.

This listing is unclaimed, so publisher identity cannot be verified and it stays below the “verified” tier by design — that is a statement about the listing, not about the project’s quality. Our automated scans still ran; a maintainer who claims it unlocks identity verification.

VersionChannelReleasedNotes
0.0.0stableSep 21, 2026Indexed listing — see the upstream repository for real release history.