GravityLabs/goose
Html Content / Article Extractor in Scala - open sourced from Gravity Labs observed · 2026-08-28
Health v2 · maintenance only
10/100
- Activity 0
- Release rhythm 35
- Longevity 100
Flags: no_releases archived
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: n/a
- age_days: 5739
- days_rel: n/a
- days_push: 3424
- n_releases_24m: 0
Adoption not part of the score
1526 stars · 311 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded
Goose is a Scala library (originally Java) that extracts the main body text, metadata, publish date, embedded videos, and top image from news article web pages. It was open sourced by Gravity.com in 2011 for powering article-snippet apps like Flipboard/Pulse.
Use cases
- extract main article text from a news page url
- get the main image from an article for a preview card
- pull meta description and publish date from web pages
- build a read-it-later or feed reader app that shows article snippets
- extract embedded youtube/vimeo videos from articles
When to avoid
- you need an actively maintained extractor - the last release was 2017
- you are not on the JVM ecosystem
- you need modern web page handling for heavily JavaScript-rendered sites
Facets
library · maturity abandoned
parser web-scraping nlp web-development crawlers jvm cli article-extraction html-parsing content-extraction metadata-extraction readability natural-language-processing
2 sources
- readme: https://github.com/GravityLabs/goose · fetched 2026-08-28 · 9b6e1701ae33
- homepage: http://gravity.com · fetched 2026-08-29 · 05f39b2c0595
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| GravityLabs/goose | main | 10 |
For agents
markdown · JSON · MCP: product_card(name="GravityLabs/goose")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem