FazBrowse GitHub Viewer | Trending |
URL:
| Home
Tools: [Download Repo ZIP]   [Original HTTPS Page]

html-extractor · GitHub Topics · GitHub

#

html-extractor

Here are 16 public repositories matching this topic...

Module for automatic summarization of text documents and HTML pages.

  • Updated Aug 14, 2026
  • Python

Reworked https://www.readability.com/ parsing library (now https://mercury.postlight.com/ is living alternative)

  • Updated Aug 12, 2026
  • HTML

Automatically extract the main text content (and more) from an HTML document

  • Updated Sep 1, 2022
  • Kotlin

基于行块分布函数的通用网页正文抽取算法优化,Python实现

  • Updated Feb 17, 2020
  • Python

从html中提取正文,用于新闻类网页

  • Updated Feb 24, 2023
  • Go

PHP library which determines which css is used from html snippets.

  • Updated Nov 7, 2019
  • PHP

Xtract-htmlV2 is a tool for getting the HTML code from the website you want and is the successor to the previous version

  • Updated Oct 16, 2025
  • Python

Xtract-html is a tool for extracting HTML display code from a website, which you can also use for your website.

  • Updated Oct 16, 2025
  • Python

Go package that cleans a HTML page for better readability.

  • Updated Aug 1, 2023
  • HTML

Media Graper is a open source tool for Linux which is developed to extract all the Images, links, Videos from a Webpage.

  • Updated Mar 17, 2023
  • Shell

ℂ𝕀ℕ𝔼𝔽𝕪 ℙℝ𝕆 - The ultimate website code lookup+API finder (OPEN SOURCE)

  • Updated Jul 19, 2026
  • JavaScript

Python CLI: Question bank vignettes → AI-enhanced Anki flashcards. Clerkship, shelf exam, and board exam prep for medical students and residents.

  • Updated May 6, 2026
  • Python

A simple extractor based on BeatufulSoup, You can use it to iterate through all the HTML files in the website root directory and get the text, placeholders and other text.

  • Updated Dec 16, 2019
  • Python

A Java-based server leveraging Apache Tika to extract content and metadata from files (PDF, DOCX, TXT, etc.) in a local files-to-extract directory. Supports HTML (with CSS styling) and text extraction, file listing, and metadata retrieval via MCP-compliant tools and REST APIs. Built with Spring Boot, Jetty, and MCP SDK.

  • Updated Aug 30, 2025
  • Java

Public repository for the SavedPixel HTML CSS JS Extractor Chrome extension: pick webpage elements and export clean HTML, CSS, fonts, images, optional scripts, and Markdown locally.

  • Updated May 21, 2026
  • JavaScript

Improve this page

Add a description, image, and links to the html-extractor topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the html-extractor topic, visit your repo's landing page and select "manage topics."

Learn more


Back | FazBrowse Home | New Git URL