Document Processing

  1. pypdf: Pure Python PDF Toolkit

    When you need to manipulate PDFs in Python without heavy external dependencies, pypdf is the go-to solution. This pure Python library provides...

    Open Source
  2. PyMuPDF: High-Performance PDF Processing for Python

    When you need raw speed for PDF processing, PyMuPDF is the performance leader among Python PDF libraries. Built as a Python binding to the C-based...

    Open Source
  3. OmniParse: Open-Source Universal Data Parsing for GenAI Pipelines

    Modern GenAI applications consume data in many forms -- PDFs, spreadsheets, images, audio recordings, and video files. Building a RAG pipeline that...

    AI
  4. Marker: Open-Source PDF to Markdown Conversion with Deep Learning

    PDF documents remain one of the most common formats for knowledge distribution, yet they are among the most difficult to process programmatically...

    AI
  5. GPT-PDF: Parse PDFs into Markdown Using Vision LLMs with Just 293 Lines of Code

    PDF documents are the universal format for sharing information, but they are notoriously difficult for software to parse. Traditional PDF parsers...

    AI