# DocNow/twarc

A command line tool (and Python library) for archiving Twitter JSON

Repository: https://github.com/DocNow/twarc
Canonical: https://ross.abutalabs.com/products/twarc
Homepage: https://twarc-project.readthedocs.io
Language: Python
License: MIT
License Family: permissive
Last push: 2025-10-31T18:03:47+00:00

## Health v2 (maintenance only)
Score: 50/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 49, release rhythm 22, longevity 100
- inputs: {"age_days": 4979, "days_push": 306, "days_rel": 306, "gap_med": null, "n_releases_24m": 1}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1394, forks 250 (observed 2026-08-28T04:04:36.360817+00:00)

## What it is
twarc is a command line tool and Python library for collecting and archiving Twitter JSON data via the Twitter v1.1 and v2 APIs. It was developed for researchers and archivists to harvest and preserve Twitter data.

## Use cases
- archive tweets as  for research
- collect twitter data via the api from the command line
- harvest tweets matching a search query
- download a user's tweets for analysis
- build a twitter dataset for academic study
- use a python library to page through twitter api results

## When to choose
- you need to collect and archive Twitter JSON data via the v1.1 or v2 API
- you want a scriptable command line tool or Python library for Twitter data harvesting
- you are doing social media research or web archiving of Twitter content

## When to avoid
- you need actively maintained Twitter collection tooling, since API quota changes made twarc unusable and it is no longer supported
- you want to collect data from platforms other than Twitter
- you need a GUI-based social media analytics tool

## Facets
- artifact type: cli-tool
- maturity: abandoned
- function: cli, web-scraping, etl
- domain: social-media, data-science, developer-tools
- platform: python, cli, cross-platform
- tags: twitter-api, twitter-archive, social-media-data, data-collection, academic-research, -archiving, command-line

## Member repositories
- DocNow/twarc (main) score 50

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:04:36.360817+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T04:39:25.243852+00:00, confidence not recorded.
  - readme: https://github.com/DocNow/twarc (fetched 2026-08-28T04:04:36.360817+00:00, sha fb6029de223f)
- Data as of 2026-08-30T08:39:29.467469+00:00.
