# zinggAI/zingg

Scalable master data management, identity resolution, entity resolution, and deduplication using ML

Repository: https://github.com/zinggAI/zingg
Canonical: https://ross.abutalabs.com/products/zingg
Language: Java
License: AGPL-3.0
License Family: copyleft
Topics: fuzzymatch, fuzzy-matching, deduplication, dedupe, masterdata, dataengineering, entity-resolution, identity-resolution, data-science, spark, dataquality, datalake, master-data-management, customer-data-platform, databricks, snowflake, cdp, mdm, master-data
Last push: 2026-09-02T14:53:35+00:00

## Health v2 (maintenance only)
Score: 88/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 100, release rhythm 65, longevity 100
- inputs: {"age_days": 1834, "days_push": 0, "days_rel": 23, "gap_med": 201.5, "n_releases_24m": 3}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 1243, forks 175 (observed 2026-09-03T02:15:20.907436+00:00)

## What it is
Zingg is an ML-based tool for scalable master data management, entity resolution, identity resolution, and record deduplication. It runs on Apache Spark and integrates with platforms like Databricks and Snowflake to build unified views of business entities from messy, duplicated data.

## Use cases
- deduplicate customer records across systems
- resolve duplicate identities in a data warehouse
- build a single customer view from multiple data silos
- fuzzy match records with field variations at scale
- clean up dimension tables before analytics
- master data management for a customer data platform

## When to choose
- you need scalable entity resolution on large datasets using Spark
- you want ML-driven matching with an interactive training data builder
- you are consolidating customer data across Databricks, Snowflake, or a data lake

## When to avoid
- you need a lightweight in-memory dedupe for small datasets
- you require a permissive license (AGPL-3.0)
- your stack does not run JVM/Spark workloads

## Facets
- artifact type: application
- maturity: active
- function: machine-learning, etl, data-science, search-engine
- domain: data-science, big-data, analytics
- platform: jvm, windows
- tags: entity-resolution, identity-resolution, master-data-management, deduplication, fuzzy-matching, data-quality, customer-data-platform, apache-spark, data-engineering, spark, databricks, snowflake, linux, macos

## Member repositories
- zinggAI/zingg (main) score 88

## Provenance
- Observed fields: from GitHub, fetched 2026-09-03T02:15:20.907436+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T08:21:55.254096+00:00, confidence not recorded.
  - readme: https://github.com/zinggAI/zingg (fetched 2026-09-03T02:15:20.907436+00:00, sha 8289ba5e6a1c)
- Data as of 2026-08-30T08:39:29.467469+00:00.
