Ross ROSS = Recommend OSS · open-source software intelligence for agents

h2oai/h2o-3

H2O is an Open Source, Distributed, Fast & Scalable Machine Learning Platform: Deep Learning, Gradient Boosting (GBM) & XGBoost, Random Forest, Generalized Linear Modeling (GLM with Elastic Net), K-Means, PCA, Generalized Additive Models (GAM), RuleFit, Support Vector Machine (SVM), Stacked Ensembles, Automatic Machine Learning (AutoML), etc. observed · 2026-08-28

github.com/h2oai/h2o-3 · homepage · Jupyter Notebook · Apache-2.0 (permissive) observed · 2026-08-28

Health v2 · maintenance only

77/100

  • Activity 99
  • Release rhythm 35
  • Longevity 100

Flags: no_releases

How is this computed?

round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-02. Adoption (stars, forks) is never an input.

  • gap_med: n/a
  • age_days: 4566
  • days_rel: n/a
  • days_push: 7
  • n_releases_24m: 0

Full methodology

Adoption not part of the score

7494 stars · 2025 forks observed · 2026-08-28

What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded

H2O-3 is an open-source, distributed, in-memory machine learning platform implementing algorithms such as GLM, GBM/XGBoost, Random Forest, Deep Learning, Stacked Ensembles, and AutoML. It is accessible from Python, R, Scala, Java, and a Flow web UI, integrates with Hadoop and Spark, and exports models as POJO/MOJO for fast production scoring.

Use cases

  • train gradient boosting models on large datasets
  • run automatic machine learning to find the best model
  • train deep learning models from Python or R
  • score models in production with MOJO export
  • run machine learning on Hadoop or Spark clusters
  • build stacked ensembles of multiple models
  • do PCA, K-Means, and GLM on big data

When to choose

  • you need distributed, scalable ML beyond a single machine's memory
  • you want AutoML to automatically train and tune many models
  • you need fast production scoring via POJO/MOJO export
  • you work in Python or R but need big-data performance
  • you want an Apache-licensed ML platform with Hadoop/Spark integration

When to avoid

  • you need lightweight scikit-learn-style ML on small datasets
  • you need deep learning with GPU training on neural architectures like CNNs or transformers
  • you want a pure-Python stack without a JVM dependency
  • you need LLM fine-tuning or generative AI features

Facets

library · maturity stable

machine-learning deep-learning data-science benchmarking machine-learning data-science big-data artificial-intelligence python jvm cross-platform cloud automl gbm gradient-boosting random-forest glm stacked-ensembles distributed-ml hadoop spark mojo-scoring r scala docker

8 sources

Member repositories

RepositoryRoleHealth v2
h2oai/h2o-3main77

For agents

markdown · JSON · MCP: product_card(name="h2oai/h2o-3")

Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem