# whylabs/whylogs

An open-source data logging library for machine learning models and data pipelines. 📚 Provides visibility into data quality & model performance over time. 🛡️ Supports privacy-preserving data collection, ensuring safety & robustness. 📈

Repository: https://github.com/whylabs/whylogs
Canonical: https://ross.abutalabs.com/products/whylogs
Homepage: https://whylogs.readthedocs.io/
Language: Jupyter Notebook
License: Apache-2.0
License Family: permissive
Topics: ai-pipelines, approximate-statistics, statistical-properties, data-quality, calculate-statistics, python, logging, mlops, dataops, ml-pipelines, data-pipeline, dataset, machine-learning, data-science, analytics, constraints, data-constraints, model-performance
Last push: 2025-01-10T20:14:49+00:00

## Health v2 (maintenance only)
Score: 34/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 0, release rhythm 40, longevity 100
- inputs: {"age_days": 2210, "days_push": 600, "days_rel": 638, "gap_med": 7, "n_releases_24m": 10}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2831, forks 144 (observed 2026-08-28T04:07:24.377117+00:00)

## What it is
whylogs is an open-source Python data logging library for machine learning models and data pipelines. It profiles datasets with privacy-preserving approximate statistics to track data quality and model performance over time.

## Use cases
- profile datasets to monitor data quality in ML pipelines
- detect data drift between training and production data
- log model performance metrics over time
- collect privacy-preserving data statistics without storing raw data
- validate data against quality constraints before training
- integrate data observability into MLOps workflows

## When to choose
- you need lightweight, privacy-preserving data profiling for ML pipelines
- you want to track data drift and model performance over time
- you're building MLOps observability and want an open-source profiling standard

## When to avoid
- you need full-fidelity storage of raw data rather than statistical profiles
- you need a complete commercial monitoring dashboard out of the box
- your stack is not Python-based

## Facets
- artifact type: library
- maturity: active
- function: logging, monitoring, analytics, data-science, machine-learning
- domain: machine-learning, data-science, monitoring, analytics
- platform: python, cross-platform
- tags: data-logging, data-quality, mlops, data-profiling, drift-detection, privacy-preserving, approximate-statistics, model-monitoring, data-engineering

## Member repositories
- whylabs/whylogs (main) score 34

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:24.377117+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T07:38:09.479876+00:00, confidence not recorded.
  - readme: https://github.com/whylabs/whylogs (fetched 2026-08-28T04:07:24.377117+00:00, sha 217ed82f87b7)
- Data as of 2026-08-30T08:39:29.467469+00:00.
