unstructured is by Unstructured-IO — not by us.
We indexed this project so people can find it. We are not selling it, we host no copy of the code, and we are not affiliated with or endorsed by its authors — get it from them.
Indexed Aug 3, 2026 · 15,246 stars at index time. Maintainers: claiming verifies your identity and unlocks a higher assurance tier. Removal requests are honored.
Application · for humans & their agents
unstructured
unclaimed listingactively maintainedFreeApache-2.0
Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex doc
by Open Source Community · New publisher
Every claim on this page is refundable if it is untrue — refund policy.
Unstructured-IO/unstructured is an open-source project by Unstructured-IO: Convert documents to structured data effortlessly. Unstructured is open-source ETL solution for transforming complex documents into clean, structured formats for language models. Visit our website to learn more about our enterprise grade Platform product for production grade workflows, partitioning, enrichments, chunking and embedding.. Indexed here so it can be found — not resold.
It is free. Get it from the upstream repository: https://github.com/Unstructured-IO/unstructured
From the project's own README (excerpt, reproduced for discovery under its Apache-2.0 license):
Open-Source Pre-Processing Tools for Unstructured Data
The unstructured library provides open-source components for ingesting and pre-processing images and text documents, such as PDFs, HTML, Word docs, and many more. The use cases of unstructured revolve around streamlining and optimizing the data processing workflow for LLMs. unstructured modular functions and connectors form a cohesive system that simplifies data ingestion and pre-processing, making it adaptable to different platforms and efficient in transforming unstructured data into structured outputs.
Unstructured Transform MCP — Document Processing for your Agents
Unstructured Transform brings production-grade document processing to your agents as an MCP server. It gives them the ability to turn 60+ file types into structured data that's ready for your applications, vector databases, and any downstream processes by parsing, enriching, chunking, and embedding files directly inside their current session.
Setup Steps for Your Agent
1. Pick your MCP client. Transform works with virtually any MCP-compatible host or agent framework — Claude Code, Cursor, Codex CLI and more.
2. Add the Transform MCP server to your client's MCP configuration (via the CLI mcp add command or the client's MCP settings/config file, depending on the tool).
3. Authenticate once when your client prompts you. Sign in, and the Transform tools become available to your agent on its next message.
4. Point your agent at a file. Drag and drop or reference a local file or URL. Transform handles 60+ formats (PDFs, emails, images, scanned files, and more).
5. Describe what you need in plain language. Tell the agent your intent (e.g. "parse and chunk this contract for a vector store") and Transform partitions, enriches, chunks, and embeds…
Preview
See what it does before you commit. Previews show behavior, never source code.
01Capabilities
Does
- + Structured document extraction
- + OCR
- + Content extraction verification
Doesn’t
- No exclusions declared
02Requirements & stack
Depends on
No declared dependencies
Credentials needed
None declared
Stack
03Community
No endorsements yetNo verified confirmations yet — be the first.
Confirmations come from verified purchasers, installers, vetted reviewers, or an installation outcome your org reported through the agent tools. They grade quality — security is verified separately, and community votes can never override the security gate.
Sign in to confirm — weight comes from verified usage, not vote count.
Issues 0
Open an issueNobody has reported anything yet — a success counts as a report too.
04Trust Passport
Full passport →0/0 automated components pass. An automated score is never a security guarantee.
This listing is unclaimed, so publisher identity cannot be verified and it stays below the “verified” tier by design — that is a statement about the listing, not about the project’s quality. Our automated scans still ran; a maintainer who claims it unlocks identity verification.
- publisher identity Publisher status verified; 0 verification(s) on file
- malicious pattern scan No known malicious-behavior patterns across listing text only — no source artifact published
- capability contract All 0 observed capability reference(s) match the declared manifest
- agent safety scan No injection patterns in agent-readable content
- provenance No release signature or provenance attestation
- behavioral sandbox Not performed in this environment — requires the production isolated runner (docs/sandbox-requirements.md). No untrusted code is ever executed on the application host.
Every listing must pass this review before it can be sold, and it is re-run on every release. Verification describes what we checked — it is not a guarantee that the software is safe.
05Versions
Full history →| Version | Channel | Released | Notes |
|---|---|---|---|
| 0.0.0 | stable | Aug 3, 2026 | Indexed listing — see the upstream repository for real release history. |