Ross ROSS = Recommend OSS · open-source software intelligence for agents

DTStack/Taier

Taier is a big data development platform for submission, scheduling, operation and maintenance, and indicator information display observed · 2026-08-28

github.com/DTStack/Taier · homepage · Java · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

84/100

  • Activity 95
  • Release rhythm 62
  • Longevity 100
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 2010
  • days_rel: 41
  • days_push: 33
  • n_releases_24m: 1

Full methodology

Adoption not part of the score

1284 stars · 349 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Taier is a self-hosted distributed dispatching platform for big data that handles task submission, DAG-based scheduling, and operations/maintenance with indicator dashboards. It provides an IDE-style web development environment for batch and streaming jobs across Hadoop, Flink, and Spark clusters with multi-tenant, kerberos-aware isolation.

Use cases

  • schedule and orchestrate ETL jobs with upstream/downstream dependencies
  • self-hosted alternative to Azkaban or DolphinScheduler for big data workflows
  • web IDE for developing FlinkSQL, SparkSQL, and data synchronization tasks
  • manage batch and stream data pipelines in one platform with versioned tasks
  • monitor scheduled data tasks, view error logs, and run an operations center
  • submit Hadoop/Flink/Spark jobs with tenant and cluster isolation

When to choose

  • You need a web-based platform to author, schedule, and monitor ETL workflows with complex DAG dependencies
  • You run heterogeneous Hadoop/Flink/Spark clusters and want one UI with multi-tenant isolation and Kerberos support
  • You want an IDE-like environment for FlinkSQL/SparkSQL development plus task versioning and custom task plugins

When to avoid

  • You only need simple cron-style scheduling without big data cluster integration
  • Your workloads live outside the Hadoop/Flink/Spark ecosystem
  • You need a lightweight embeddable library rather than a full self-hosted platform

Facets

application · maturity active

scheduling workflow-automation etl monitoring developer-tools big-data self-hosted jvm self-hosted workflow-scheduler dag-orchestration job-scheduler big-data-platform hadoop flink spark hive task-scheduling azkaban-alternative etl-orchestration multi-tenant data-engineering automation docker web-server linux

2 sources

Member repositories

RepositoryRoleHealth v2
DTStack/Taiermain84

For agents

markdown · JSON · MCP: product_card(name="DTStack/Taier")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem