PyMuPDF

A high-performance Python library for extracting, rendering, converting, and manipulating PDF and other document formats.

Library
PyPI
v1.28.2
10,871 stars
GNU AGPLv3

Repository Health

Pre-computed score based on development activity, maintenance, community, maturity, and trend momentum. How we score it →
91 /100 Excellent
Development Activity 100
Maintenance 96
Community 68
Maturity 60
Momentum 40

Technical Analysis

AI-assessed by reading the actual repository — architecture, code quality, innovation, and documentation. How we score it →
78 /100 Good
Architecture 78
Code Quality 82
Innovation 70
Learning Curve 80

PyMuPDF is a Python binding for MuPDF, a lightweight and fast C library for parsing and rendering PDF, XPS, EPUB, and other document formats. It gives Python code low-level, pixel-accurate control over document content — text with full font/position/color metadata, embedded images, vector graphics, and page structure — alongside high-level convenience methods for common tasks like text extraction, table detection, annotation, redaction, and page rendering to images.

With no mandatory external dependencies and prebuilt wheels for Windows, macOS, and Linux, it installs with a single pip install pymupdf and is used heavily in document-processing pipelines, from traditional text extraction to LLM/RAG ingestion workflows via its companion package PyMuPDF4LLM. The project ships as open source under AGPL-3.0, with an optional commercial license (PyMuPDF Pro) unlocking Office-format conversion (DOCX, XLSX, PPTX) for teams that can’t comply with AGPL’s copyleft terms.

What You Get

  • Text extraction as plain text, structured dicts with font/size/color/bbox metadata, HTML, XML, or raw blocks
  • Table detection via find_tables(), exportable directly to Markdown or a Pandas DataFrame
  • Page rendering to high-resolution Pixmap images at any DPI, plus SVG vector output
  • Annotation, highlighting, and redaction APIs, including add_redact_annot() / apply_redactions() for permanently removing sensitive content
  • Document assembly operations — merge, split, insert, and reorder pages across PDFs
  • OCR integration via Tesseract for extracting text from scanned pages

Common Use Cases

  • Building document-ingestion pipelines that extract text and layout for search indexing or RAG/LLM retrieval
  • Rendering PDF pages to images for thumbnails, previews, or downstream computer-vision processing
  • Redacting sensitive fields (SSNs, account numbers) from contracts and forms before sharing externally
  • Merging, splitting, and reassembling PDFs as part of an automated document-generation workflow
  • Extracting tabular data from PDF reports and invoices into structured Pandas DataFrames

Under The Hood

Architecture PyMuPDF is a thin, ergonomic Python layer over MuPDF, a C rendering engine vendored and built from source at install time (see setup.py’s MuPDF download/build step and the pipcl-based pyproject.toml build backend). The public API lives in src/ (pymupdf.py, utils.py, table.py, plus SWIG-generated bindings), with an older compatibility surface preserved separately in src_classic/ for backward compatibility with pre-rename fitz imports (fitz___init__.py, fitz_utils.py). Document, Page, and Pixmap objects wrap MuPDF’s C structs directly, so most operations are thin dispatches into the underlying engine rather than reimplemented logic — the core abstraction to understand is this C-to-Python boundary: changing it means touching the SWIG interface files (extra.i) and rebuilding the native extension, not just editing Python.

Tech Stack The project builds a bundled MuPDF C/C++ library via a custom pipcl-based PEP-517 backend (no setuptools/hatchling), producing platform wheels for Windows, macOS, and Linux across Python 3.10-3.14. Optional companion packages extend it without adding mandatory dependencies: pymupdf-fonts for extended font coverage, pymupdf4llm for Markdown/RAG-oriented extraction, and pymupdfpro for Office-format (DOCX/XLSX/PPTX) conversion gated behind a commercial license key. Table extraction integrates optionally with Pandas; OCR integrates optionally with a system Tesseract install.

Code Quality The repo carries an extensive test suite (76+ files under tests/, driven by pytest.ini with --tb=native) and a correspondingly heavy CI setup — seven separate GitHub Actions workflows covering standard tests, a matrix across Python versions (test_multiple.yml), Pyodide/WASM builds, system-installed MuPDF builds, Valgrind memory checks, and wheel builds for every supported platform. This level of CI investment is unusual for a Python package and reflects the project’s need to validate a compiled C extension across many platform/Python combinations, not just pure-Python logic.

What Makes It Unique PyMuPDF’s differentiator is that it exposes MuPDF’s C-level performance and accuracy (pixel-perfect text positioning, font/color metadata, fast rasterization) through a Pythonic API, rather than shelling out to a CLI tool or relying on a pure-Python PDF parser. Its redaction API removes content from the underlying stream rather than just overlaying it visually, and its ecosystem of companion packages (pymupdf4llm for RAG, pymupdfpro for Office formats) lets it scale from simple text extraction up to full document-conversion pipelines without changing the core library.

Used by 12 apps in this directory

Python
59%
Other

Arkon

AI Assistants · Knowledge Management · Mcp

1,463

Self-hosted enterprise AI knowledge hub that compiles internal docs into a scoped, reviewable wiki and serves it to Claude and other LLMs through an MCP server.

View details
46
Repo Health
74
Technical
70
Dependency
Built with
Python 59%
TypeScript 41%
Updated 4 months ago
Python
97%
MIT

auto-news

AI Assistants · Productivity

908

An AI-powered personal news aggregator that filters multi-source feeds through LLMs and delivers curated, noise-free summaries to your Notion workspace.

View details
43
Repo Health
53
Technical
66
Dependency
Built with
Python 97%
Updated 1 years ago
Rust
52%
Apache 2.0

cocoindex

AI Development · Data Engineering

11,607

An incremental data indexing engine that keeps AI agent context perpetually fresh by reprocessing only what changed.

View details
87
Repo Health
85
Technical
65
Dependency
Built with
Rust 52%
Python 48%
Updated 1 weeks ago
Python
70%
Apache 2.0

GPT Researcher

AI Assistants · Productivity

29,650

The pioneering open-source autonomous AI agent that conducts deep, multi-source research and produces citation-backed reports exceeding 2,000 words — faster and more reliably than any human researcher.

View details
91
Repo Health
91
Technical
63
Dependency
Built with
Python 70%
TypeScript 18%
Updated 2 weeks ago
Python
100%
MIT

hyperresearch

AI Agents · Developer Tools · Knowledge Management

3,691

A tier-adaptive 16-step deep research pipeline for Claude Code that produces adversarially-audited reports with full source provenance, backed by a persistent, searchable vault.

View details
79
Repo Health
85
Technical
72
Dependency
Built with
Python 100%
Updated 2 weeks ago
Python
51%
AGPL 3.0

Khoj

AI Assistants · Knowledge Management · Productivity

37,526

A self-hostable AI second brain that chats with your documents, searches the web, builds custom agents, and runs entirely on your own LLM.

View details
63
Repo Health
82
Technical
66
Dependency
Built with
Python 51%
TypeScript 36%
Updated 2 months ago
Python
86%
Apache 2.0

knowhere

AI Development · AI Memory · Developer Tools

3,541

Transform messy, unstructured documents into persistent, navigable memory that AI agents can actually use.

View details
82
Repo Health
75
Technical
66
Dependency
Built with
Python 86%
HTML 14%
Updated 2 weeks ago
Rust
84%
Apache 2.0

liteparse

Developer Tools

12,682

A fast, lightweight, open-source document parser that extracts spatial text, bounding boxes, and Markdown from PDFs and Office files — entirely on your machine.

View details
82
Repo Health
80
Technical
74
Dependency
Built with
Rust 84%
Updated 2 weeks ago
Python
66%
Other

Morphik

AI Development · Databases · Search

3,716

Morphik is an AI-native ingestion and retrieval engine that lets developers store, search, and reason over visually rich documents — scanned PDFs, manuals, slides, and video — without duct-taping together OCR, an embedding model, and a vector database.

View details
60
Repo Health
71
Technical
67
Dependency
Built with
Python 66%
TypeScript 23%
Updated 2 weeks ago

Join founders buildingwith open source

Opinionated takes, migration guides, cost-saving tips, and insights from the open source ecosystem.

Subscribe on Substack
Join 750+ subscribers