Document Processing

Meilleurs outils IA pour le traitement de documents (2026)

Moteurs OCR, parseurs PDF, compréhension de documents et pipelines d'extraction. Transformez vos documents non structurés en données exploitables et recherchables.

30 outils
PyMuPDF — High-Performance PDF Processing for Python logo

PyMuPDF — High-Performance PDF Processing for Python

PyMuPDF is a Python binding for the MuPDF library that provides fast, comprehensive PDF (and other document) processing. It supports text extraction, rendering, annotation, merging, form filling, and OCR across PDF, XPS, EPUB, and image formats.

Script Depot 199Scripts
MonkeyOCR — Lightweight Document Parsing Model logo

MonkeyOCR — Lightweight Document Parsing Model

A lightweight large multimodal model optimized for accurate document parsing extracting text tables and structure from PDFs and images.

AI Open Source 152Configs
Claude Official Skill: PDF — Read, Create & Edit PDFs logo

Claude Official Skill: PDF — Read, Create & Edit PDFs

Claude Code skill for PDF files. Read content, extract data, create new PDFs, merge documents, and convert formats. Activates automatically.

Anthropic 708Skills
Zerox — Zero-Shot PDF OCR for AI Pipelines logo

Zerox — Zero-Shot PDF OCR for AI Pipelines

Extract text from any PDF using vision models as OCR. Zerox converts PDF pages to images then uses GPT-4o or Claude to extract clean markdown without training.

Script Depot 579Skills
OpenDataLoader PDF — AI-Ready Document Parser logo

OpenDataLoader PDF — AI-Ready Document Parser

An open-source PDF parser that automates document accessibility and extracts structured, AI-ready data including tables, text, bounding boxes, and tagged content.

AI Open Source 524Skills
LiteParse — Fast Open-Source Document Parser in Rust logo

LiteParse — Fast Open-Source Document Parser in Rust

A fast, helpful, and open-source document parser by LlamaIndex that extracts structured text from PDFs and other documents with high speed and accuracy for RAG and AI pipelines.

Script Depot 350Scripts
DeepSeek-OCR — High-Accuracy Optical Context Compression logo

DeepSeek-OCR — High-Accuracy Optical Context Compression

An OCR model and toolkit from DeepSeek AI that extracts text from images and documents with high accuracy, designed for feeding structured content into LLM pipelines.

AI Open Source 349Configs
Xberg — Polyglot Document Intelligence Framework in Rust logo

Xberg — Polyglot Document Intelligence Framework in Rust

A cross-language document extraction framework with a Rust core that parses PDFs, Office files, images, and 97+ formats into structured text and metadata.

AI Open Source 274Configs
Umi-OCR — Free Offline OCR Tool for Screenshots, Images & PDFs logo

Umi-OCR — Free Offline OCR Tool for Screenshots, Images & PDFs

Open-source, privacy-first OCR software that runs entirely offline. Supports batch image import, PDF recognition, QR code scanning, and multi-language text extraction without sending data to external servers.

Script Depot 245Scripts
Papermerge — Self-Hosted Document Management for Digital Archives logo

Papermerge — Self-Hosted Document Management for Digital Archives

Papermerge is a self-hosted, open-source document management system with OCR, full-text search, and hierarchical folder organization for scanned documents and PDFs.

AI Open Source 238Configs
DeepSeek OCR — Context-Aware Document Optical Compression logo

DeepSeek OCR — Context-Aware Document Optical Compression

High-accuracy document OCR system by DeepSeek that converts scanned documents and PDFs into structured text with layout-aware compression.

AI Open Source 143Configs
Surya — Document OCR for 90+ Languages logo

Surya — Document OCR for 90+ Languages

Surya is a document OCR toolkit with 19.5K+ GitHub stars. Text recognition in 90+ languages, layout analysis, table detection, reading order, and LaTeX OCR. Benchmarks favorably against cloud OCR serv

Script Depot 1,036Skills
Paperless-ngx — Self-Hosted Document Management with OCR logo

Paperless-ngx — Self-Hosted Document Management with OCR

Paperless-ngx is an open-source document management system that scans, OCRs, indexes, and archives all your physical and digital documents for full-text search.

Script Depot 729Skills
Stirling PDF — Self-Hosted PDF Editor & Toolkit logo

Stirling PDF — Self-Hosted PDF Editor & Toolkit

Stirling PDF is the #1 open-source PDF tool on GitHub. Merge, split, convert, compress, OCR, sign, and edit PDFs — all self-hosted with no data leaving your server.

Script Depot 689Skills
Gotenberg — API-Driven Document Conversion and PDF Generation Server logo

Gotenberg — API-Driven Document Conversion and PDF Generation Server

Docker-powered API server for converting HTML, Markdown, Office documents, and URLs into PDFs using Chromium and LibreOffice.

Script Depot 546Skills
BentoPDF — Privacy-First Self-Hosted PDF Toolkit logo

BentoPDF — Privacy-First Self-Hosted PDF Toolkit

BentoPDF is a self-hosted web application that provides a comprehensive set of PDF tools including merging, splitting, converting, and OCR without sending files to external services.

AI Open Source 541Skills
Tesseract OCR — Open Source Text Recognition Engine for 100+ Languages logo

Tesseract OCR — Open Source Text Recognition Engine for 100+ Languages

Tesseract is an open-source OCR engine maintained by Google, supporting over 100 languages. It converts images and scanned documents into machine-readable text with high accuracy across multiple output formats.

Script Depot 532Skills
Claude Office Skills — Docs/PDF/Sheets Skill Set logo

Claude Office Skills — Docs/PDF/Sheets Skill Set

A curated repo of office-focused skills (docs, PDF, spreadsheets) and an Office MCP server; copy skills into Claude Code to standardize document workflows.

Skill Factory 519Skills
PaddleOCR — AI-Powered OCR Toolkit for 100+ Languages logo

PaddleOCR — AI-Powered OCR Toolkit for 100+ Languages

A lightweight, production-ready OCR system supporting 100+ languages. Bridges documents and images to structured data for LLM pipelines.

Script Depot 510Skills
Grimmory — Self-Hosted eBook and Comics Library Server logo

Grimmory — Self-Hosted eBook and Comics Library Server

Grimmory is a self-hosted digital library server for managing and reading eBooks, comics, and documents. It supports EPUB, PDF, CBR, CBZ, and MOBI formats with metadata management, OPDS feeds, and a responsive web reader.

AI Open Source 330Configs
jsPDF — Generate PDF Documents in JavaScript logo

jsPDF — Generate PDF Documents in JavaScript

A client-side JavaScript library for generating PDF documents programmatically in the browser and Node.js.

AI Open Source 327Configs
React PDF — Display PDF Documents in React Applications logo

React PDF — Display PDF Documents in React Applications

A React component library for rendering PDF files in the browser using Mozilla pdf.js, with support for pagination, zoom, text selection, and annotations.

AI Open Source 318Configs
Chandra — OCR Model for Complex Tables, Forms, and Handwriting logo

Chandra — OCR Model for Complex Tables, Forms, and Handwriting

High-accuracy OCR model that handles structured documents with complex tables, nested forms, and handwritten annotations while preserving full layout fidelity.

Script Depot 288Scripts
pdfmake — Client-Server PDF Generation for JavaScript logo

pdfmake — Client-Server PDF Generation for JavaScript

Create complex PDF documents in the browser or Node.js using a declarative document-definition object.

Script Depot 278Scripts
Teedy — Lightweight Self-Hosted Document Management System logo

Teedy — Lightweight Self-Hosted Document Management System

Teedy is a lightweight, open-source document management system with full-text search, OCR, workflow automation, and a clean web interface for organizing files and metadata.

Script Depot 242Scripts
PhpSpreadsheet — Read and Write Spreadsheet Files in Pure PHP logo

PhpSpreadsheet — Read and Write Spreadsheet Files in Pure PHP

PhpSpreadsheet is a library for reading, creating, and writing spreadsheet documents in PHP, supporting formats including Excel (XLSX/XLS), LibreOffice Calc (ODS), CSV, and PDF export.

Script Depot 231Scripts
Unlimited-OCR — One-Shot Long-Horizon Document Parsing by Baidu logo

Unlimited-OCR — One-Shot Long-Horizon Document Parsing by Baidu

Open-source OCR system from Baidu that parses complex documents in a single pass with high accuracy across diverse layouts.

AI Open Source 210Configs
Quarkdown — Markdown with Superpowers for Papers, Slides & Websites logo

Quarkdown — Markdown with Superpowers for Papers, Slides & Websites

Quarkdown extends Markdown with a scripting layer, enabling complex documents like academic papers, presentations, and books from a single source. Built in Kotlin, it compiles to HTML and PDF.

AI Open Source 180Configs
Read the Docs — Documentation Hosting Platform for Open Source logo

Read the Docs — Documentation Hosting Platform for Open Source

Free documentation hosting service that builds and publishes Sphinx and MkDocs projects automatically from version control, with versioning, search, and PDF export built in.

Script Depot 178Scripts
react-pdf — Display PDFs in React as Easily as Images logo

react-pdf — Display PDFs in React as Easily as Images

A React component library that renders PDF documents in the browser using Mozilla's pdf.js under the hood. Supports page navigation, text selection, annotations, and responsive scaling out of the box.

AI Open Source 172Configs

L'intelligence documentaire par l'IA

AI Document Intelligence

AI document processing has leapfrogged traditional OCR. Modern tools don't just recognize characters — they understand document layout, hierarchy, tables, and semantic structure. OCR & Text Extraction — Surya delivers state-of-the-art multilingual OCR with layout detection. Marker converts PDFs to clean Markdown preserving structure. MinerU handles complex scientific papers with equations and diagrams.

Document ETL — DocETL and Unstructured build production pipelines that ingest PDFs, Word docs, scanned images, and HTML into normalized, chunked output ready for RAG or database storage. Translation & Accessibility — PDFMathTranslate preserves mathematical notation while translating academic papers across 100+ languages.

Knowledge Extraction — RAGFlow and Kotaemon combine document parsing with retrieval, letting you ask natural language questions over your document collection with source citations. MarkItDown converts any Office format to Markdown for AI processing.

The world's knowledge is trapped in PDFs — AI document tools are the key that unlocks it.

Questions fréquentes

Quel est le meilleur outil IA pour extraire du texte de PDF ?+

Pour les PDF généraux : Marker convertit en Markdown propre avec une excellente préservation de la mise en page. Pour les documents scannés : Surya OCR gère 90+ langues avec une précision supérieure sur les mises en page complexes. Pour les articles scientifiques : MinerU se spécialise dans les équations, tableaux et figures. Pour les pipelines de production : Unstructured et DocETL fournissent un traitement de documents de bout en bout avec chunking et extraction de métadonnées.

L'IA peut-elle extraire des tableaux de PDF avec précision ?+

Oui. Les outils modernes comme Surya, Marker et MinerU utilisent des modèles de vision qui comprennent la structure des tableaux — en-têtes, cellules fusionnées, lignes étendues — pas seulement les lignes de grille. La précision dépasse 95 % sur des tableaux bien formatés. Pour les tableaux complexes ou inconsistants, combiner plusieurs outils (OCR + détection de mise en page + post-traitement LLM) donne les meilleurs résultats.

Comment traiter des milliers de documents avec l'IA ?+

Utilisez des outils de pipeline comme DocETL ou Unstructured qui gèrent le batching, le traitement parallèle et la récupération d'erreurs. Ils normalisent différents formats (PDF, DOCX, images, HTML) en un format de sortie unique, extraient les métadonnées, chunkent le contenu pour le RAG et stockent les résultats dans votre base ou vector store. TokRepo héberge des configurations de pipeline préconfigurées pour les workflows de traitement de documents courants.

Explorer les catégories associées