# collabH/bigdata-growth

大数据知识仓库涉及到数据仓库建模、实时计算、大数据、数据中台、系统设计、Java、算法等。

Repository: https://github.com/collabH/bigdata-growth
Canonical: https://ross.abutalabs.com/products/bigdata-growth
Language: Python
License: MIT
License Family: permissive
Topics: flink, kafka, hive, mapreduce, spark, olap, kudu, hadoop, hbase, debezium, hdfs, bigdata, hudi, bigdatalearning
Last push: 2026-04-18T13:12:36+00:00

## Health v2 (maintenance only)
Score: 58/100 (v2, computed 2026-09-02T17:46:02.011165+00:00)
- activity 78, release rhythm 8, longevity 100
- inputs: {"age_days": 2275, "days_push": 137, "days_rel": null, "gap_med": null, "n_releases_24m": 0}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1806, forks 395 (observed 2026-08-28T04:05:39.048505+00:00)

## What it is
A curated Chinese-language knowledge repository covering big data topics such as data warehouse modeling, stream processing, data lakes, and distributed systems, with notes on Flink, Spark, Hadoop, Hudi, Paimon, and Iceberg. It also includes an AI skill assistant for interview prep and study guidance.

## Use cases
- learn big data from scratch
- prepare for big data engineer interviews
- understand flink checkpoint mechanism
- study data warehouse modeling
- learn hudi iceberg paimon data lakehouse
- get a spark streaming learning path
- review hive tuning strategies

## When to choose
- you want structured study notes on the Hadoop/Flink/Spark ecosystem
- you are preparing for big data interviews
- you need data lake (Hudi/Paimon/Iceberg) learning material

## When to avoid
- you need production-ready code or tooling rather than documentation
- you need English-language materials

## Facets
- artifact type: learning-resource
- maturity: active
- function: documentation
- domain: big-data, tutorials
- platform: python
- tags: big-data, data-warehouse, flink, spark, hadoop, interview-preparation, knowledge-base, chinese, data-engineering

## Member repositories
- collabH/bigdata-growth (main) score 58

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:39.048505+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:21:29.274057+00:00, confidence not recorded.
  - readme: https://github.com/collabH/bigdata-growth (fetched 2026-08-28T04:05:39.048505+00:00, sha 6eeb94dce2fe)
- Data as of 2026-08-30T08:39:29.467469+00:00.
