by NameetP · MCP Server · ★ 79
pdfmux Self-healing PDF extraction with per-page confidence scoring. Open-source LlamaParse alternative for RAG pipelines, MCP server for Claude Desktop, LangChain + LlamaIndex loaders. Ranked #2 on opendataloader-bench (0.900). The only PDF extractor that audits its own output. Catches blank pages, scrambled columns, broken tables — re-extracts them with a stronger backend. So your LLM gets clean data, not silent garbage. Routes each page to the best of 5 rule-based backends + BYOK LLM fallback (Gemini / Claude / GPT-4o / Ollama). One CLI. One API. Zero config.
| Stars | 79 |
| Forks | 12 |
| Language | Python |
| Category | MCP Server |
| License | MIT |
| Quality Score | 68.870137129299/100 |
| Open Issues | 4 |
| Last Updated | 2026-08-13 |
| Created | 2026-03-03 |
| Platforms | cli, mcp, python |
| Est. Tokens | ~19k |
Explore other popular mcp server tools:
pdfmux is PDF extraction that audits its own output — and certifies any other extractor's, catching pages they silently dropped. Verify signed manifests offline: free, MIT, no account. 0.903 on opendataloader-b. It is categorized as a MCP Server with 79 GitHub stars.
pdfmux is primarily written in Python. It covers topics such as ai-agent, docling, document-ai.
You can find installation instructions and usage details in the pdfmux GitHub repository at github.com/NameetP/pdfmux. The project has 79 stars and 12 forks, indicating an active community.
pdfmux is released under the MIT license, making it free to use and modify according to the license terms.