LLM Training

  1. The NanoGPT Speedrun Frontier: AI Agents Closed 81.7% of the Human Gap — But Couldn't Invent Anything New

    Prime Intellect ran the largest open experiment in autonomous AI research: 18 frontier models (Fable 5, Opus 5, Kimi K3, GPT-5.6 Sol, DeepSeek V4 Pro...) executed ~10,000 …

    AI
  2. olmOCR: AI2's Open-Source PDF-to-Markdown Toolkit for LLM Training Data

    Converting PDFs to clean, machine-readable text at scale is one of the foundational challenges in LLM dataset preparation. Traditional PDF parsers...

    AI
  3. llm.c: Karpathy's Minimal C Implementation of LLM Training

    Most developers and researchers who work with large language models interact with them through high-level frameworks like PyTorch or Hugging Face...

    AI
  4. LlamaFactory: Open-Source LLM Fine-Tuning Framework

    Fine-tuning large language models was once a complex, resource-intensive process reserved for organizations with large GPU clusters. LlamaFactory has...

    Open Source