# lakesoul-io/LakeSoul

LakeSoul is an end-to-end, realtime cloud-native Lakehouse framework for fast data ingestion, concurrent updates, incremental analytics, multimodal data processing and vector search — powering next-generation BI and AI workloads.

Repository: https://github.com/lakesoul-io/LakeSoul
Canonical: https://ross.abutalabs.com/products/lakesoul
Homepage: https://lakesoul-io.github.io/
Language: Java
License: Apache-2.0
License Family: permissive
Topics: lakehouse, spark, flink, streaming, postgresql, rust, sql, huggingface, python, pytorch, arrow, datafusion, vectorized, velox, gluten, daft, ray, vector-search
Last push: 2026-08-25T10:47:58+00:00

## Health v2 (maintenance only)
Score: 82/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 99, release rhythm 49, longevity 100
- inputs: {"age_days": 1709, "days_push": 8, "days_rel": 342, "gap_med": 1, "n_releases_24m": 4}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 3247, forks 423 (observed 2026-08-28T04:07:50.698295+00:00)

## What it is
LakeSoul is a cloud-native, real-time lakehouse framework with a Rust-native core providing ACID table format, concurrent upserts, incremental reads, and vector search. It integrates with Spark, Flink, Presto, Ray, Daft, and DuckDB, with PostgreSQL-based metadata management and built-in compaction and RBAC.

## Use cases
- build a real-time lakehouse with streaming ingestion from Kafka and Flink CDC
- run concurrent upserts and incremental reads on data lake tables
- query lakehouse tables with SQL from Spark, Flink, Presto, or DuckDB
- prepare tabular training datasets for AI and PyTorch workloads
- perform vector search over lakehouse data
- unify batch and stream processing on one table format
- avoid stitching together separate catalogs, compaction services, and auth layers

## When to choose
- you need a batteries-included lakehouse platform rather than just a table format
- you require high-concurrency writes with ACID guarantees and auto conflict resolution
- you want one Rust-native core shared consistently across Java, Python, and C++ engines
- you need real-time incremental pipelines on Hadoop or Kubernetes clusters
- you want built-in compaction, RBAC, and vector retrieval out of the box

## When to avoid
- you only need a widely adopted table format with the broadest ecosystem support, such as Apache Iceberg
- your stack relies on engines not in LakeSoul's compatibility matrix
- you prefer a fully serverless managed warehouse over self-managed lakehouse infrastructure
- your workloads are small-scale and don't need lakehouse complexity

## Facets
- artifact type: framework
- maturity: active
- function: database, streaming, etl, vector-database, search-engine, serialization, data-science
- domain: big-data, databases, analytics, machine-learning
- platform: jvm, python, rust, cloud, self-hosted
- tags: lakehouse, table-format, spark, flink, apache-arrow, datafusion, upsert, acid-transactions, cdc, incremental-processing, olap, ray, daft, rust-core, vector-search, data-engineering, real-time, docker, kubernetes, linux, macos

## Member repositories
- lakesoul-io/LakeSoul (main) score 82

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:07:50.698295+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-29T18:43:54.403546+00:00, confidence not recorded.
  - readme: https://github.com/lakesoul-io/LakeSoul (fetched 2026-08-28T04:07:50.698295+00:00, sha b5c5d8495aa7)
  - homepage: https://lakesoul-io.github.io/ (fetched 2026-08-29T09:37:04.849255+00:00, sha 69a022d69e2b)
- Data as of 2026-08-30T08:39:29.467469+00:00.
