# bellingcat/auto-archiver

Automatically archive links to videos, images, and social media content from Google Sheets (and more).

Repository: https://github.com/bellingcat/auto-archiver
Canonical: https://ross.abutalabs.com/products/auto-archiver
Homepage: https://pypi.org/project/auto-archiver/
Language: Python
License: MIT
License Family: permissive
Topics: archive, docker, open-source-research, python, service, scraping, web-archiving
Last push: 2026-08-18T10:40:29+00:00

## Health v2 (maintenance only)
Score: 92/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 98, release rhythm 81, longevity 100
- inputs: {"age_days": 2056, "days_push": 15, "days_rel": 128, "gap_med": 5, "n_releases_24m": 20}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1109, forks 107 (observed 2026-08-28T04:03:37.243730+00:00)

## What it is
A Python tool by Bellingcat that automatically archives web content such as videos, images, social media posts, and webpages from URLs supplied via Google Sheets, CSV files, or the command line. Archived content can be enriched and stored locally or remotely (S3, Google Drive), with status reports written back to the source.

## Use cases
- archive social media posts before they get deleted
- bulk archive URLs from a Google Sheet
- preserve videos and images for open-source investigations
- automatically snapshot webpages for evidence preservation
- download and store online content for OSINT research

## When to choose
- you need verifiable, automated archiving of online content for research or journalism
- you want to feed URLs from Google Sheets or CSV and get status back
- you want a Docker-deployable archiving pipeline with remote storage options

## When to avoid
- you need a browser-based one-click webpage archiver like the Wayback Machine
- you only need simple link bookmarking without content capture
- your sources require heavy JavaScript interaction beyond the supported extractors

## Facets
- artifact type: cli-tool
- maturity: active
- function: web-scraping, cli, etl, file-upload
- domain: osint, crawlers, developer-tools
- platform: python, cli, cross-platform
- tags: web-archiving, social-media-archiving, google-sheets, osint, bellingcat, evidence-preservation, automation, docker

## Member repositories
- bellingcat/auto-archiver (main) score 92

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:03:37.243730+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T06:43:33.867035+00:00, confidence not recorded.
  - readme: https://github.com/bellingcat/auto-archiver (fetched 2026-08-28T04:03:37.243730+00:00, sha a9dc3bddcb82)
  - homepage: https://pypi.org/project/auto-archiver/ (fetched 2026-08-29T12:47:42.738398+00:00, sha 4b4e8fead74a)
  - registry_pypi: https://pypi.org/pypi/auto-archiver/json (fetched 2026-08-29T12:47:42.747990+00:00, sha 1cc9da65614b)
- Data as of 2026-08-30T08:39:29.467469+00:00.
