# tinyfish-io/bigset-oss

Open-source BigSet — self-hostable live datasets populated by TinyFish web agents

Repository: https://github.com/tinyfish-io/bigset-oss
Canonical: https://ross.abutalabs.com/products/bigset-oss
Language: TypeScript
License: AGPL-3.0
License Family: copyleft
Topics: open-source, tinyfish, bigset
Last push: 2026-08-24T09:05:53+00:00

## Health v2 (maintenance only)
Score: 76/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 99, release rhythm 87, longevity 7
- inputs: {"age_days": 110, "days_push": 9, "days_rel": 85, "gap_med": 0.0, "n_releases_24m": 3}
- flags: young
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1684, forks 201 (observed 2026-08-28T04:05:21.958667+00:00)

## What it is
BigSet is a self-hostable application that turns a natural-language sentence into a structured, regularly refreshed dataset by dispatching autonomous web agents to research, verify, deduplicate, and export the data. It infers schemas automatically and supports refresh cadences from 30 minutes to weekly, exporting to CSV or XLSX.

## Use cases
- build a dataset of YC companies hiring engineers with funding stage and location
- keep a price list refreshed daily from the live web
- generate verified lead lists from company websites
- turn a one-sentence data request into a CSV table
- schedule agents to re-crawl and update a dataset every 6 hours
- extract structured job postings across many sites without writing scrapers

## When to choose
- you need fresh structured data from the web on a recurring schedule
- you want schema inference and verification handled automatically instead of stitching scrapers, search APIs, and cron jobs
- you prefer self-hosting an open-source data pipeline

## When to avoid
- you need a battle-tested production tool with stable APIs
- you only need simple scraping of known URLs
- your project requires a permissive license since BigSet is AGPL-3.0

## Facets
- artifact type: application
- maturity: experimental
- function: web-scraping, data-generation, agent-framework, etl, search-engine
- domain: data-science, crawlers, artificial-intelligence
- platform: self-hosted
- tags: live-datasets, autonomous-agents, schema-inference, scheduled-refresh, csv-export, xlsx-export, tinyfish, ai-agents, automation, web-server, nodejs, docker

## Member repositories
- tinyfish-io/bigset-oss (main) score 76

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:21.958667+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:41:17.659498+00:00, confidence not recorded.
  - readme: https://github.com/tinyfish-io/bigset-oss (fetched 2026-08-28T04:05:21.958667+00:00, sha 755c4d95e349)
- Data as of 2026-08-30T08:39:29.467469+00:00.
