Ross ROSS = Recommend OSS · open-source software intelligence for agents

ageitgey/node-unfluff

Automatically extract body content (and other cool stuff) from an html document observed · 2026-08-28

github.com/ageitgey/node-unfluff · HTML · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

32/100

  • Activity 0
  • Release rhythm 35
  • Longevity 100

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 4447
  • days_rel: n/a
  • days_push: 1195
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

2158 stars · 210 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

A Node.js library and CLI tool that automatically extracts the main body content and metadata (title, author, date, images, tags, links) from HTML web pages, based on python-goose. It turns cluttered webpages into clean plain text or JSON data.

Use cases

  • extract main article text from a webpage html
  • build machine learning datasets from web pages
  • get article metadata like author and publish date from html
  • strip ads and navigation from web pages to get clean text
  • build an Instapaper-style read-it-later app
  • extract embedded videos and images from articles from the command line

When to choose

  • you need to extract readable article content from raw HTML in Node.js
  • you want a CLI to quickly pull main text and metadata from web pages
  • you're building datasets or text-processing pipelines from web content

When to avoid

  • you need a Python or JVM content extractor (use python-goose or goose instead)
  • you need actively maintained scraping with modern site support
  • you want to scrape pages requiring JavaScript rendering

Facets

library · maturity maintenance

parser web-scraping nlp web-development crawlers developer-tools cli cross-platform content-extraction readability html-parsing article-extraction metadata-extraction natural-language-processing nodejs

1 source

Member repositories

RepositoryRoleHealth v2
ageitgey/node-unfluffmain32

For agents

markdown · JSON · MCP: product_card(name="ageitgey/node-unfluff")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem