pdfmux Self-healing PDF extraction with per-page confidence scoring. Open-source LlamaParse alternative for RAG pipelines, MCP server for Claude Desktop, LangChain + LlamaIndex loaders. Ranked #2 on opendataloader-bench (0.900). The only PDF extractor that audits its own output. Catches blank pages, scrambled columns, broken tables — re-extracts them with a stronger backend. So your LLM gets clean data, not silent garbage. Routes each page to the best of 5 rule-based backends + BYOK LLM fallback (Gemini / Claude / GPT-4o / Ollama). One CLI. One API. Zero config.
| Stars | 67 |
| Forks | 9 |
| Language | Python |
| Category | MCP 服务器 |
| License | MIT |
| Quality Score | 36.25/100 |
| Open Issues | 3 |
| Last Updated | 2026-06-10 |
| Created | 2026-03-03 |
| Platforms | cli, mcp, python |
| Est. Tokens | ~52k |
Explore other popular mcp 服务器 tools:
pdfmux is PDF extraction that checks its own work. #2 reading order accuracy — zero AI, zero GPU, zero cost.. It is categorized as a MCP 服务器 with 67 GitHub stars.
pdfmux is primarily written in Python. It covers topics such as ai-agent, docling, document-parsing.
You can find installation instructions and usage details in the pdfmux GitHub repository at github.com/NameetP/pdfmux. The project has 67 stars and 9 forks, indicating an active community.
pdfmux is released under the MIT license, making it free to use and modify according to the license terms.