lakesoul-io/LakeSoul
LakeSoul is an end-to-end, realtime cloud-native Lakehouse framework for fast data ingestion, concurrent updates, incremental analytics, multimodal data processing and vector search — powering next-generation BI and AI workloads. observed · 2026-08-28
Health v2 · maintenance only
82/100
- Activity 99
- Release rhythm 49
- Longevity 100
How is this computed?
round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10) — computed 2026-09-03. Adoption (stars, forks) is never an input.
- gap_med: 1
- age_days: 1709
- days_rel: 342
- days_push: 8
- n_releases_24m: 4
Adoption not part of the score
3247 stars · 423 forks observed · 2026-08-28
What it is AI-extracted, prompt v1, taxonomy v1, 2026-08-29, confidence not recorded
LakeSoul is a cloud-native, real-time lakehouse framework with a Rust-native core providing ACID table format, concurrent upserts, incremental reads, and vector search. It integrates with Spark, Flink, Presto, Ray, Daft, and DuckDB, with PostgreSQL-based metadata management and built-in compaction and RBAC.
Use cases
- build a real-time lakehouse with streaming ingestion from Kafka and Flink CDC
- run concurrent upserts and incremental reads on data lake tables
- query lakehouse tables with SQL from Spark, Flink, Presto, or DuckDB
- prepare tabular training datasets for AI and PyTorch workloads
- perform vector search over lakehouse data
- unify batch and stream processing on one table format
- avoid stitching together separate catalogs, compaction services, and auth layers
When to choose
- you need a batteries-included lakehouse platform rather than just a table format
- you require high-concurrency writes with ACID guarantees and auto conflict resolution
- you want one Rust-native core shared consistently across Java, Python, and C++ engines
- you need real-time incremental pipelines on Hadoop or Kubernetes clusters
- you want built-in compaction, RBAC, and vector retrieval out of the box
When to avoid
- you only need a widely adopted table format with the broadest ecosystem support, such as Apache Iceberg
- your stack relies on engines not in LakeSoul's compatibility matrix
- you prefer a fully serverless managed warehouse over self-managed lakehouse infrastructure
- your workloads are small-scale and don't need lakehouse complexity
Facets
framework · maturity active
database streaming etl vector-database search-engine serialization data-science big-data databases analytics machine-learning jvm python rust cloud self-hosted lakehouse table-format spark flink apache-arrow datafusion upsert acid-transactions cdc incremental-processing olap ray daft rust-core vector-search data-engineering real-time docker kubernetes linux macos
2 sources
- readme: https://github.com/lakesoul-io/LakeSoul · fetched 2026-08-28 · b5c5d8495aa7
- homepage: https://lakesoul-io.github.io/ · fetched 2026-08-29 · 69a022d69e2b
Member repositories
| Repository | Role | Health v2 |
|---|---|---|
| lakesoul-io/LakeSoul | main | 82 |
For agents
markdown · JSON · MCP: product_card(name="lakesoul-io/LakeSoul")
Data as of 2026-08-30T08:39:29.467469+00:00 · Report a problem