# Qihoo360/hbox

AI on Hadoop

Repository: https://github.com/Qihoo360/hbox
Canonical: https://ross.abutalabs.com/products/hbox
Language: Java
License: Apache-2.0
License Family: permissive
Topics: hadoop, tensorflow, caffe, mxnet, ai, deeplearning, machinelearning, yarn
Last push: 2025-07-01T08:13:28+00:00

## Health v2 (maintenance only)
Score: 47/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 29, release rhythm 40, longevity 100
- inputs: {"age_days": 3197, "days_push": 428, "days_rel": 428, "gap_med": 5, "n_releases_24m": 6}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1727, forks 383 (observed 2026-08-28T04:05:28.274818+00:00)

## What it is
Hbox (formerly XLearning) is a scheduling platform that runs machine learning and deep learning frameworks like TensorFlow, MXNet, Caffe, and PyTorch on Hadoop YARN clusters. It provides GPU resource scheduling, HDFS-based unified data management, Docker container support, and a RESTful API management interface.

## Use cases
- run distributed TensorFlow training on a Hadoop YARN cluster
- schedule GPU resources for deep learning jobs alongside big data workloads
- train models using HDFS-stored data with multiple input strategies
- run PyTorch or MXNet training jobs in Docker containers on YARN
- manage deep learning training jobs via a RESTful API
- share a cluster between Spark/MapReduce and deep learning workloads

## When to choose
- you already operate a Hadoop YARN cluster and want to run deep learning training on it
- you need GPU-aware scheduling mixed with big data jobs
- you want unified HDFS-based data management for training data and model outputs

## When to avoid
- you run a Kubernetes-based ML platform (use Kubeflow or similar instead)
- you need a modern cloud-native scheduler without Hadoop dependencies
- your training is single-node and needs no cluster scheduling

## Facets
- artifact type: application
- maturity: maintenance
- function: scheduling, machine-learning, deep-learning, workflow-automation, gpu-computing
- domain: machine-learning, deep-learning, big-data, microservices, artificial-intelligence
- platform: jvm, self-hosted
- tags: hadoop-yarn, tensorflow, mxnet, caffe, pytorch, gpu-scheduling, hdfs, distributed-training, resource-management, linux, docker

## Member repositories
- Qihoo360/hbox (main) score 47

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:05:28.274818+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T03:32:15.152960+00:00, confidence not recorded.
  - readme: https://github.com/Qihoo360/hbox (fetched 2026-08-28T04:05:28.274818+00:00, sha a1c4079ff4a6)
- Data as of 2026-08-30T08:39:29.467469+00:00.
