FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

data-extraction · GitHub Topics · GitHub

#

data-extraction

Here are 3,855 public repositories matching this topic...

The context API to search, scrape, and interact with the web at scale. 🔥

  • Updated Aug 28, 2026
  • TypeScript

🕷️ An adaptive Web Scraping framework that handles everything from a single request to a full-scale crawl!

  • Updated Aug 25, 2026
  • Python

🔥 The open-source no-code platform for web scraping, crawling, search and AI data extraction • Turn websites into structured APIs in minutes 🔥

  • Updated Aug 27, 2026
  • TypeScript

Declarative data automation language and Go runtime for structured extraction workflows.

  • Updated Aug 25, 2026
  • Go

Extract Keywords from sentence or Replace keywords in sentences.

  • Updated Apr 13, 2025
  • Python

Browser automation CLI built for AI agents. Break through anti-bot walls, hand off to humans across platforms when stuck. Parallel multi-task execution, independent multi-session operation, isolated multi-account browsing.

  • Updated Aug 24, 2026
  • Python

A powerful Model Context Protocol (MCP) server that provides an all-in-one solution for public web access.

  • Updated Aug 12, 2026
  • JavaScript

The undetected self-hosted browser automation platform. Powered by Camoufox (Firefox) for 0% detection rates. Built for speed, privacy, and scalability.

  • Updated Aug 26, 2026
  • TypeScript

To extract article from given URL

  • Updated Aug 20, 2026
  • TypeScript

Converts a pdf file into a text file while keeping the layout of the original pdf. Useful to extract the content from a table in a pdf file for instance. This is a subclass of PDFTextStripper class (from the Apache PDFBox library).

  • Updated Dec 17, 2023
  • Java

A beginner-friendly yet powerful Python toolkit for financial analysis and automation — built to make modern investing accessible to everyone

  • Updated Jul 24, 2026
  • Python

Lightweight library for scraping web-sites with LLMs

  • Updated Dec 17, 2025
  • Python

⛏️ The extraction engine behind Maigret: turn any profile URL into a structured OSINT record across 150+ sites

  • Updated Aug 18, 2026
  • Python

The fastest PDF library for Python and Rust. Text extraction, image extraction, markdown conversion, PDF creation & editing. 0.8ms mean, 5× faster than industry leaders, 100% pass rate on 3,830 PDFs. MIT/Apache-2.0.

  • Updated Aug 27, 2026
  • Rust

Local-first, open-source AI assistant for your data. Unify tasks, notes, docs, photos, and bookmarks. Private, self-hosted, and extensible via APIs.

  • Updated May 14, 2026
  • TypeScript

Fast, lightweight Firecrawl/Tavily alternative in Rust. Web scraper, crawler & search API with MCP server for AI agents. Drop-in Firecrawl-compatible API (/scrape, /crawl, /search). 2.3x faster than Tavily, 1.5x faster than Firecrawl in 1K-URL benchmarks. 6 MB RAM, single binary. Self-host or use managed cloud.

  • Updated Aug 27, 2026
  • Rust

Improve this page

Add a description, image, and links to the data-extraction topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the data-extraction topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL