Ross ROSS = Recommend OSS · open-source software intelligence for agents

spotify/scio

A Scala API for Apache Beam and Google Cloud Dataflow. observed · 2026-08-28

github.com/spotify/scio · homepage · Scala · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

96/100

  • Activity 97
  • Release rhythm 91
  • Longevity 100
How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.

  • gap_med: 20.5
  • age_days: 4178
  • days_rel: 63
  • days_push: 22
  • n_releases_24m: 23

Full methodology

Adoption not part of the score

2628 stars · 533 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-30, confidence not recorded

Scio is a Scala API for Apache Beam and Google Cloud Dataflow, inspired by Apache Spark and Scalding. It provides a unified batch and streaming programming model with type-safe integrations for Google Cloud products and various data formats.

Use cases

  • write batch data pipelines in Scala on Apache Beam
  • run streaming data processing jobs on Google Cloud Dataflow
  • read and write BigQuery tables in a type-safe way
  • process Avro, Parquet, and Protobuf data at scale
  • integrate pipelines with Pub/Sub, Bigtable, and Cloud Storage
  • run machine learning pipelines with TensorFlow IO
  • develop and test data pipelines interactively with a Scala REPL

When to choose

  • your team works in Scala and needs a Spark-like API on Beam
  • you run data pipelines on Google Cloud Dataflow
  • you need unified batch and streaming processing with strong typing
  • you want deep integration with Google Cloud data products like BigQuery

When to avoid

  • you prefer Python or Java directly over Scala
  • you need a runner or cloud other than Beam-supported ones without extra work
  • your pipelines are small enough that a simpler tool suffices

Facets

framework · maturity active

etl streaming data-science machine-learning serialization big-data data-science cloud-computing microservices jvm cloud apache-beam google-cloud-dataflow bigquery batch-processing scala-api data-pipelines data-engineering

2 sources

Member repositories

RepositoryRoleHealth v2
spotify/sciomain96

For agents

markdown · JSON · MCP: product_card(name="spotify/scio")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem