PDF Parsing

  1. Open Parse: Visually-Driven Document Parser for LLM-Ready RAG Pipelines

    The RAG (Retrieval-Augmented Generation) ecosystem has matured rapidly, but one bottleneck persists: garbage in, garbage out. Most document parsing...

    AI
  2. olmOCR: AI2's Open-Source PDF-to-Markdown Toolkit for LLM Training Data

    Converting PDFs to clean, machine-readable text at scale is one of the foundational challenges in LLM dataset preparation. Traditional PDF parsers...

    AI
  3. MinerU: Open-Source PDF Document Parsing and Data Extraction

    PDF is the universal format for document distribution, but it is arguably the worst format for data extraction. PDFs store visual layouts —...

    AI
  4. GPT-PDF: Parse PDFs into Markdown Using Vision LLMs with Just 293 Lines of Code

    PDF documents are the universal format for sharing information, but they are notoriously difficult for software to parse. Traditional PDF parsers...

    AI