# dotnet/spark

.NET for Apache® Spark™ makes Apache Spark™ easily accessible to .NET developers.

Repository: https://github.com/dotnet/spark
Canonical: https://ross.abutalabs.com/products/dotnet-spark
Homepage: https://dot.net/spark
Language: C#
License: MIT
License Family: permissive
Topics: spark, csharp, dotnet, analytics, bigdata, spark-streaming, spark-sql, machine-learning, fsharp, dotnet-core, dotnet-standard, streaming, apache-spark, tpcds, tpch, azure, hdinsight, databricks, emr, microsoft
Last push: 2026-08-19T09:49:03+00:00

## Health v2 (maintenance only)
Score: 77/100 (v2, computed 2026-09-03T02:20:16.233290+00:00)
- activity 98, release rhythm 38, longevity 100
- inputs: {"age_days": 2690, "days_push": 14, "days_rel": 201, "gap_med": 275, "n_releases_24m": 2}
- flags: none
- formula: round(0.45*activity + 0.35*rhythm + 0.20*longevity); archived -> min(score, 10)

## Adoption (not part of the score)
Stars 2098, forks 334 (observed 2026-08-28T04:06:13.430225+00:00)

## What it is
.NET for Apache Spark provides high-performance C# and F# bindings for Apache Spark, exposing DataFrames, SparkSQL, and Structured Streaming to .NET developers. It is a .NET Standard library that runs on Windows, Linux, and macOS and on major cloud Spark services like Azure HDInsight, Databricks, and Amazon EMR.

## Use cases
- process big data with spark from c#
- run spark sql queries using .net
- build etl pipelines in c# on spark
- stream processing with spark structured streaming in .net
- run .net spark jobs on databricks or emr
- analyze large datasets with f# and spark

## When to choose
- your team is .NET-based and needs distributed big data processing
- you want to reuse existing .NET skills and libraries on Spark
- you deploy Spark workloads on Azure HDInsight, Databricks, or EMR

## When to avoid
- you primarily use Python or Scala, where native Spark APIs are richer
- you need the newest Spark features immediately, as .NET bindings lag behind
- your project needs a very active community; the project is in maintenance mode

## Facets
- artifact type: library
- maturity: maintenance
- function: etl, streaming, data-science, machine-learning, analytics
- domain: big-data, data-science, microservices, analytics
- platform: cross-platform, windows, dotnet, cloud
- tags: apache-spark, spark-sql, structured-streaming, dataframe, csharp, fsharp, databricks, azure-hdinsight, amazon-emr, net-standard, data-engineering, linux, macos, docker

## Member repositories
- dotnet/spark (main) score 77

## Provenance
- Observed fields: from GitHub, fetched 2026-08-28T04:06:13.430225+00:00.
- Health v2: computed from the inputs above; adoption is never an input.
- Inferred fields (summary, facets, guidance): AI-extracted, prompt v1, taxonomy v1, on 2026-08-30T02:54:29.380848+00:00, confidence not recorded.
  - readme: https://github.com/dotnet/spark (fetched 2026-08-28T04:06:13.430225+00:00, sha b4361a4cc73d)
  - homepage: https://dot.net/spark (fetched 2026-08-29T10:34:44.432747+00:00, sha 02e07dd27390)
  - site_page: https://learn.microsoft.com/en-us/lifecycle/faq/internet-explorer-microsoft-edge (fetched 2026-08-29T10:34:44.445827+00:00, sha c363b2f43d4b)
- Data as of 2026-08-30T08:39:29.467469+00:00.
