PySpark

Python API for Apache Spark, the unified engine for large-scale distributed data processing

Framework
PyPI
v4.2.0
44,074 stars
Apache License 2.0

Repository Health

Pre-computed score based on development activity, maintenance, community, maturity, and trend momentum. How we score it →
87 /100 Excellent
Development Activity 100
Maintenance 52
Community 96
Maturity 60
Momentum 40

Technical Analysis

AI-assessed by reading the actual repository — architecture, code quality, innovation, and documentation. How we score it →
73 /100 Good
Architecture 88
Code Quality 82
Innovation 78
Learning Curve 45

PySpark is the official Python API for Apache Spark, a unified analytics engine for distributed data processing across clusters. It exposes Spark SQL and DataFrames, a pandas-compatible API on Spark, MLlib for distributed machine learning, and Structured Streaming, all backed by Spark’s JVM-based execution engine via Py4J (classic mode) or the gRPC-based Spark Connect protocol (client mode). PySpark scaffolds an entire distributed-computation lifecycle — a SparkSession entry point, cluster resource management, and a lazily-evaluated DataFrame query plan optimized by Catalyst and executed by Tungsten — rather than being a library you simply call into.

What You Get

  • pyspark.sql — DataFrame and SQL APIs with a Catalyst-optimized, lazily-evaluated query plan
  • pyspark.pandas — a pandas-API-compatible layer that runs on Spark’s distributed engine (pandas API on Spark)
  • pyspark.ml/pyspark.mllib — distributed machine learning pipelines, feature transformers, and models
  • pyspark.streaming/Structured Streaming — micro-batch and continuous stream processing on the same DataFrame API
  • Two connection modes: classic Py4J-based local/cluster execution, or Spark Connect (gRPC) for a thin, decoupled client (pyspark-client) against a remote Spark cluster

Common Use Cases

  • Running distributed ETL and batch transformations across a Spark cluster (EMR, Databricks, on-prem YARN/Kubernetes) too large for a single machine
  • Training and scoring machine learning pipelines at scale with MLlib across partitioned data
  • Processing continuous data streams (Kafka, files, sockets) with Structured Streaming using the same DataFrame API as batch jobs
  • Migrating pandas-based analysis to a distributed engine using the pandas API on Spark with minimal code changes

Under The Hood

Architecture: PySpark lives under python/pyspark/ inside the Apache Spark monorepo, structured into core (RDD/SparkContext primitives), sql (DataFrame/Catalyst-facing API), ml/mllib (machine learning), streaming (Structured Streaming), pandas (pandas API on Spark), and resource (cluster resource management), with two separate packaging trees under python/packaging/ — classic (the traditional Py4J-bridged pyspark package driving a JVM Spark session) and client (the Spark Connect gRPC-based pyspark-client, a thin client that talks to a remote Spark cluster without embedding a JVM). Tech Stack: the core engine is Scala/JVM (Spark’s execution runtime, Catalyst optimizer, and Tungsten execution engine), bridged to Python via Py4J in classic mode or gRPC/protobuf in Spark Connect mode; the Python codebase itself targets modern CPython with type stubs, checked via mypy.ini. Code Quality: 34+ dedicated test modules under python/pyspark/tests/ plus per-subpackage test suites (sql/tests, ml/tests, streaming/tests), run through run-tests.py with coverage tracked via run-tests-with-coverage and reported to Codecov; the project enforces strict review and CI gating befitting an Apache top-level project with tens of thousands of commits. API Design: SparkSession.builder centralizes configuration, and the DataFrame API mirrors SQL/pandas semantics closely enough to be approachable, but real usage requires understanding lazy evaluation, partitioning, and cluster resource tuning — meaningful boilerplate and conceptual overhead compared to single-machine libraries, which is inherent to any framework that scaffolds a distributed execution model rather than just exposing a function call.

Used by 5 apps in this directory

Python
89%
Apache 2.0

Apache Airflow

Data Engineering

46,995

Define, schedule, and monitor complex data workflows as Python code — with a powerful UI, 80+ provider integrations, and battle-tested scalability across thousands of production deployments.

View details
96
Repo Health
89
Technical
64
Dependency
Built with
Python 89%
Updated 4 days ago
Python
89%
Apache 2.0

Apache Airflow

Data Engineering

46,995

Define, schedule, and monitor complex data workflows as Python code — with a powerful UI, 80+ provider integrations, and battle-tested scalability across thousands of production deployments.

View details
96
Repo Health
89
Technical
64
Dependency
Built with
Python 89%
Updated 4 days ago
C++
68%
Apache 2.0

ClickHouse

Analytics · Data Engineering · Databases

50,116

Open-source column-oriented database that delivers real-time analytical queries on petabyte-scale data with millisecond latency.

View details
95
Repo Health
90
Technical
64
Dependency
Built with
C++ 68%
Python 14%
Updated 4 days ago
Rust
95%
Other

Databend

Data Engineering · Databases

9,452

Open-source enterprise data warehouse unifying analytics, vector search, full-text search, and AI agent orchestration in a single Rust-built engine on S3.

View details
93
Repo Health
84
Technical
77
Dependency
Built with
Rust 95%
Updated 4 days ago
Python
65%
Apache 2.0

WrenAI

AI Agents · Analytics · Data Engineering

17,763

Open-source GenBI engine that lets AI agents turn natural-language questions into governed SQL, charts, and shareable dashboards across 20+ data sources — no vendor lock-in, no black-box prompts.

View details
90
Repo Health
91
Technical
69
Dependency
Built with
Python 65%
Rust 32%
Updated 1 weeks ago

Join founders buildingwith open source

Opinionated takes, migration guides, cost-saving tips, and insights from the open source ecosystem.

Subscribe on Substack
Join 750+ subscribers