FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

document-processing · GitHub Topics · GitHub

#

document-processing

Here are 1,683 public repositories matching this topic...

Turn any technical book PDF into a Claude Code skill — ready to study, reference, and use while you work.

  • Updated Aug 24, 2026
  • Python

A fast, helpful, and open-source document parser

  • Updated Aug 24, 2026
  • Rust

Modular SenseNova skills for building AI-powered office assistants and productivity workflows

  • Updated Aug 20, 2026
  • JavaScript

A system for agentic LLM-powered data processing and ETL

  • Updated Aug 10, 2026
  • Python

在保留版面、公式与结构的前提下进行 PDF 翻译,适用于科研与技术文档

  • Updated Jul 22, 2026
  • Python

ExtractThinker is a Document Intelligence library for LLMs, offering ORM-style interaction for flexible and powerful document workflows.

  • Updated Aug 27, 2025
  • Python

OpenOCR: An Open-Source Toolkit for General-OCR Research and Applications, integrates a unified training and evaluation benchmark, commercial-grade OCR and Document Parsing systems, and faithful reproductions of the core implementations from a wide range of academic papers.

  • Updated Aug 4, 2026
  • Python

The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.

  • Updated Aug 24, 2026
  • Rust

Give your AI agent eyes for PDFs — structured text, tables, OCR, visual evidence, and page-level citations via MCP. Native Rust, local-first.

  • Updated Aug 23, 2026
  • TypeScript

Local-first, open-source AI assistant for your data. Unify tasks, notes, docs, photos, and bookmarks. Private, self-hosted, and extensible via APIs.

  • Updated May 14, 2026
  • TypeScript

Transform unstructured documents into validated, rich and queryable knowledge graphs.

  • Updated Aug 19, 2026
  • Python

有情绪的 AI 桌面伴侣 | 隐私优先 · 语音唤醒 · 协助办公 · MCP 无限扩展 | An AI desktop companion with emotions — privacy-first, voice wake, office assistance, and MCP extensibility

  • Updated Aug 12, 2026
  • TypeScript

Java PDF table extraction & OCR library. Extract structured tables from text-based and scanned PDFs using stream, lattice (OpenCV-style grid detection), and hybrid parsing.

  • Updated Jul 25, 2026
  • Java

Open-source batch OCR workbench — a free, local alternative to ABBYY FineReader. Powered by Ollama + GLM-OCR + PP-DocLayoutV3, ~0.5s/page on RTX 4090. Three-panel editor, layout-aware, PDF/image batch processing, Markdown/Word export. 批量OCR工作台,纯本地运行,免费平替ABBYY,适合书籍文档数字化。

  • Updated Jun 18, 2026
  • Python

PowerPoint .NET library for reading, modifying, and generating PPTX presentations without Microsoft Office

  • Updated Aug 16, 2026
  • C#

Generic framework for historical document processing

  • Updated Jul 9, 2021
  • Python

TWIX is an open-source data extraction tool that reconstructs structured data from documents at scale, accurately and at low cost, by inferring the shared underlying visual template across documents

  • Updated May 10, 2026
  • Python

⚡ Cloud-native, AI-powered, document processing pipelines on AWS.

  • Updated Jan 22, 2026
  • TypeScript

Improve this page

Add a description, image, and links to the document-processing topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the document-processing topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL