Document Processing

Best AI Tools for Document Processing (2026)

OCR engines, PDF parsers, document understanding, and data extraction pipelines. Turn unstructured documents into structured, searchable data.

30 tools
PyMuPDF — High-Performance PDF Processing for Python logo

PyMuPDF — High-Performance PDF Processing for Python

PyMuPDF is a Python binding for the MuPDF library that provides fast, comprehensive PDF (and other document) processing. It supports text extraction, rendering, annotation, merging, form filling, and OCR across PDF, XPS, EPUB, and image formats.

Script Depot 144Scripts
MonkeyOCR — Lightweight Document Parsing Model logo

MonkeyOCR — Lightweight Document Parsing Model

A lightweight large multimodal model optimized for accurate document parsing extracting text tables and structure from PDFs and images.

AI Open Source 102Configs
MinerU — Extract LLM-Ready Data from Any Document logo

MinerU — Extract LLM-Ready Data from Any Document

Convert PDFs, scans, and complex documents into clean Markdown or JSON for RAG and LLM pipelines. 57K+ GitHub stars.

Script Depot 636Scripts
Claude Official Skill: PDF — Read, Create & Edit PDFs logo

Claude Official Skill: PDF — Read, Create & Edit PDFs

Claude Code skill for PDF files. Read content, extract data, create new PDFs, merge documents, and convert formats. Activates automatically.

Anthropic 590Skills
Zerox — Zero-Shot PDF OCR for AI Pipelines logo

Zerox — Zero-Shot PDF OCR for AI Pipelines

Extract text from any PDF using vision models as OCR. Zerox converts PDF pages to images then uses GPT-4o or Claude to extract clean markdown without training.

Script Depot 522Skills
Kreuzberg — Polyglot Document Intelligence Framework with a Rust Core logo

Kreuzberg — Polyglot Document Intelligence Framework with a Rust Core

An open-source document extraction framework that pulls text, metadata, images, and structured data from PDFs, Office files, images, and 97+ formats, with bindings for 11 programming languages.

Script Depot 519Skills
OpenDataLoader PDF — AI-Ready Document Parser logo

OpenDataLoader PDF — AI-Ready Document Parser

An open-source PDF parser that automates document accessibility and extracts structured, AI-ready data including tables, text, bounding boxes, and tagged content.

AI Open Source 470Skills
DeepSeek-OCR — High-Accuracy Optical Context Compression logo

DeepSeek-OCR — High-Accuracy Optical Context Compression

An OCR model and toolkit from DeepSeek AI that extracts text from images and documents with high accuracy, designed for feeding structured content into LLM pipelines.

AI Open Source 296Configs
LiteParse — Fast Open-Source Document Parser in Rust logo

LiteParse — Fast Open-Source Document Parser in Rust

A fast, helpful, and open-source document parser by LlamaIndex that extracts structured text from PDFs and other documents with high speed and accuracy for RAG and AI pipelines.

Script Depot 292Scripts
Xberg — Polyglot Document Intelligence Framework in Rust logo

Xberg — Polyglot Document Intelligence Framework in Rust

A cross-language document extraction framework with a Rust core that parses PDFs, Office files, images, and 97+ formats into structured text and metadata.

AI Open Source 219Configs
Papermerge — Self-Hosted Document Management for Digital Archives logo

Papermerge — Self-Hosted Document Management for Digital Archives

Papermerge is a self-hosted, open-source document management system with OCR, full-text search, and hierarchical folder organization for scanned documents and PDFs.

AI Open Source 200Configs
Umi-OCR — Free Offline OCR Tool for Screenshots, Images & PDFs logo

Umi-OCR — Free Offline OCR Tool for Screenshots, Images & PDFs

Open-source, privacy-first OCR software that runs entirely offline. Supports batch image import, PDF recognition, QR code scanning, and multi-language text extraction without sending data to external servers.

Script Depot 178Scripts
DeepSeek OCR — Context-Aware Document Optical Compression logo

DeepSeek OCR — Context-Aware Document Optical Compression

High-accuracy document OCR system by DeepSeek that converts scanned documents and PDFs into structured text with layout-aware compression.

AI Open Source 109Configs
Surya — Document OCR for 90+ Languages logo

Surya — Document OCR for 90+ Languages

Surya is a document OCR toolkit with 19.5K+ GitHub stars. Text recognition in 90+ languages, layout analysis, table detection, reading order, and LaTeX OCR. Benchmarks favorably against cloud OCR serv

Script Depot 933Skills
RAGFlow — Deep Document Understanding RAG Engine logo

RAGFlow — Deep Document Understanding RAG Engine

Open-source RAG engine with deep document understanding. Parses complex PDFs, tables, images. Agent-powered Q&A with citations. Multi-model. 77K+ stars.

Script Depot 724Skills
Paperless-ngx — Self-Hosted Document Management with OCR logo

Paperless-ngx — Self-Hosted Document Management with OCR

Paperless-ngx is an open-source document management system that scans, OCRs, indexes, and archives all your physical and digital documents for full-text search.

Script Depot 636Skills
Documenso — Open Source Document Signing Platform logo

Documenso — Open Source Document Signing Platform

Documenso is an open-source DocuSign alternative for self-hosted document signing with PDF e-signatures, audit trails, and Next.js stack.

AI Open Source 606Skills
Kotaemon — Open-Source RAG Document Chat logo

Kotaemon — Open-Source RAG Document Chat

Clean, open-source RAG tool for chatting with your documents. Supports PDF, DOCX, web pages. Multi-model, citation, and multi-user. Self-hostable. 25K+ stars.

Script Depot 603Skills
Stirling PDF — Self-Hosted PDF Editor & Toolkit logo

Stirling PDF — Self-Hosted PDF Editor & Toolkit

Stirling PDF is the #1 open-source PDF tool on GitHub. Merge, split, convert, compress, OCR, sign, and edit PDFs — all self-hosted with no data leaving your server.

Script Depot 593Skills
Docling — Document Parsing for AI logo

Docling — Document Parsing for AI

IBM document parsing library. Converts PDFs, DOCX, PPTX, images, and HTML into structured markdown or JSON. Built for RAG pipelines and LLM ingestion.

Script Depot 494SkillsCLI Tools
Gotenberg — API-Driven Document Conversion and PDF Generation Server logo

Gotenberg — API-Driven Document Conversion and PDF Generation Server

Docker-powered API server for converting HTML, Markdown, Office documents, and URLs into PDFs using Chromium and LibreOffice.

Script Depot 478Skills
Tesseract OCR — Open Source Text Recognition Engine for 100+ Languages logo

Tesseract OCR — Open Source Text Recognition Engine for 100+ Languages

Tesseract is an open-source OCR engine maintained by Google, supporting over 100 languages. It converts images and scanned documents into machine-readable text with high accuracy across multiple output formats.

Script Depot 471Skills
BentoPDF — Privacy-First Self-Hosted PDF Toolkit logo

BentoPDF — Privacy-First Self-Hosted PDF Toolkit

BentoPDF is a self-hosted web application that provides a comprehensive set of PDF tools including merging, splitting, converting, and OCR without sending files to external services.

AI Open Source 470Skills
PaddleOCR — AI-Powered OCR Toolkit for 100+ Languages logo

PaddleOCR — AI-Powered OCR Toolkit for 100+ Languages

A lightweight, production-ready OCR system supporting 100+ languages. Bridges documents and images to structured data for LLM pipelines.

Script Depot 433Skills
Pandoc — Universal Document Format Converter logo

Pandoc — Universal Document Format Converter

Pandoc is a universal document converter that reads and writes dozens of markup formats. It converts between Markdown, LaTeX, HTML, DOCX, EPUB, PDF, and many more with a single command.

Script Depot 432Skills
Claude Office Skills — Docs/PDF/Sheets Skill Set logo

Claude Office Skills — Docs/PDF/Sheets Skill Set

A curated repo of office-focused skills (docs, PDF, spreadsheets) and an Office MCP server; copy skills into Claude Code to standardize document workflows.

Skill Factory 427Skills
KOReader — Document Viewer for E-Ink Devices and Beyond logo

KOReader — Document Viewer for E-Ink Devices and Beyond

KOReader is a free, open-source document viewer optimized for e-ink readers like Kindle, Kobo, and PocketBook. It supports PDF, EPUB, DJVU, and many other formats with fine-grained rendering controls.

AI Open Source 390Skills
Nougat — Neural Optical Understanding for Academic Documents logo

Nougat — Neural Optical Understanding for Academic Documents

Nougat is a visual transformer model from Meta that converts academic PDF pages into structured Markdown, accurately preserving mathematical equations, tables, and text formatting.

AI Open Source 341Skills
jsPDF — Generate PDF Documents in JavaScript logo

jsPDF — Generate PDF Documents in JavaScript

A client-side JavaScript library for generating PDF documents programmatically in the browser and Node.js.

AI Open Source 282Configs
Grimmory — Self-Hosted eBook and Comics Library Server logo

Grimmory — Self-Hosted eBook and Comics Library Server

Grimmory is a self-hosted digital library server for managing and reading eBooks, comics, and documents. It supports EPUB, PDF, CBR, CBZ, and MOBI formats with metadata management, OPDS feeds, and a responsive web reader.

AI Open Source 282Configs

AI Document Intelligence

AI Document Intelligence

AI document processing has leapfrogged traditional OCR. Modern tools don't just recognize characters — they understand document layout, hierarchy, tables, and semantic structure. OCR & Text Extraction — Surya delivers state-of-the-art multilingual OCR with layout detection. Marker converts PDFs to clean Markdown preserving structure. MinerU handles complex scientific papers with equations and diagrams.

Document ETL — DocETL and Unstructured build production pipelines that ingest PDFs, Word docs, scanned images, and HTML into normalized, chunked output ready for RAG or database storage. Translation & Accessibility — PDFMathTranslate preserves mathematical notation while translating academic papers across 100+ languages.

Knowledge Extraction — RAGFlow and Kotaemon combine document parsing with retrieval, letting you ask natural language questions over your document collection with source citations. MarkItDown converts any Office format to Markdown for AI processing.

The world's knowledge is trapped in PDFs — AI document tools are the key that unlocks it.

Frequently Asked Questions

What is the best AI tool for extracting text from PDFs?+

For general PDFs: Marker converts to clean Markdown with excellent layout preservation. For scanned documents: Surya OCR handles 90+ languages with superior accuracy on complex layouts. For scientific papers: MinerU specializes in equations, tables, and figure extraction. For production pipelines: Unstructured and DocETL provide end-to-end document processing with chunking and metadata extraction.

Can AI extract tables from PDFs accurately?+

Yes. Modern tools like Surya, Marker, and MinerU use vision models that understand table structure — headers, merged cells, spanning rows — not just grid lines. Accuracy exceeds 95% on well-formatted tables. For complex or inconsistent tables, combining multiple tools (OCR + layout detection + LLM post-processing) produces the best results.

How do I process thousands of documents with AI?+

Use pipeline tools like DocETL or Unstructured that handle batching, parallel processing, and error recovery. They normalize different formats (PDF, DOCX, images, HTML) into a single output format, extract metadata, chunk content for RAG, and store results in your database or vector store. TokRepo hosts pre-configured pipeline configs for common document processing workflows.

Explore Related Categories