Open Source Data Engineering Apps

Discover open source data engineering tools for building reliable data pipelines, ETL processes & scalable data storage. Unlock the power of your data!

36 apps available

Apps in Data Engineering

C++
68%
Apache 2.0

ClickHouse

Analytics · Data Engineering · Databases

50,116

Open-source column-oriented database that delivers real-time analytical queries on petabyte-scale data with millisecond latency.

View details
95
Repo Health
90
Technical
64
Dependency
Built with
C++ 68%
Python 14%
Updated 5 days ago
Java
70%
GPL 3.0

DataEase

AI Assistants · Analytics · Data Engineering

24,558

Open-source BI tool with drag-and-drop dashboards, 20+ data source connectors, and AI-powered natural language queries — a self-hosted alternative to Tableau.

View details
94
Repo Health
71
Technical
65
Dependency
Built with
Java 70%
Vue 29%
Updated 5 days ago
Rust
95%
Other

Databend

Data Engineering · Databases

9,452

Open-source enterprise data warehouse unifying analytics, vector search, full-text search, and AI agent orchestration in a single Rust-built engine on S3.

View details
93
Repo Health
84
Technical
77
Dependency
Built with
Rust 95%
Updated 5 days ago
Java
58%
Apache 2.0

Kestra

Automation · Data Engineering · Devops

28,388

Event-driven orchestration platform for data, AI, and infrastructure workflows — define everything in YAML, run anywhere at scale.

View details
93
Repo Health
81
Technical
72
Dependency
Built with
Java 58%
TypeScript 26%
Vue 15%
Updated 1 weeks ago
C++
75%
Apache 2.0

Timeplus Proton

Analytics · Data Engineering

2,262

Single C++ binary SQL engine for real-time stream processing, ETL, and analytics on Kafka, Redpanda, and ClickHouse with sub-millisecond latency.

View details
88
Repo Health
82
Technical
68
Dependency
Built with
C++ 75%
Python 11%
Updated 1 weeks ago
Rust
52%
Apache 2.0

cocoindex

AI Development · Data Engineering

11,607

An incremental data indexing engine that keeps AI agent context perpetually fresh by reprocessing only what changed.

View details
87
Repo Health
85
Technical
65
Dependency
Built with
Rust 52%
Python 48%
Updated 5 days ago
Rust
86%
Other

Arroyo

Analytics · Data Engineering

5,039

A distributed stream processing engine written in Rust that lets you write SQL to run stateful, real-time computations over data streams with subsecond results.

View details
86
Repo Health
74
Technical
62
Dependency
Built with
Rust 86%
Updated 1 weeks ago
Rust
97%
Apache 2.0

Volga

Data Engineering

161

A Rust-based real-time data processing engine for AI/ML feature computation, built on Apache DataFusion and Arrow — positioned as an alternative to Flink, Spark, Chronon, and OpenMLDB with unified streaming, batch, and request-time execution.

View details
67
Repo Health
67
Technical
72
Dependency
Built with
Rust 97%
Updated 6 days ago
Java
34%
Apache 2.0

Enso

Analytics · Data Engineering · Low Code Platforms

7,441

A visual and textual programming platform for data prep and analysis where the node graph and the underlying Enso code are always perfectly in sync, built by an Alteryx co-founder on a GraalVM engine.

View details
59
Repo Health
90
Technical
61
Dependency
Built with
Java 34%
TypeScript 27%
Scala 26%
Updated 1 months ago
C++
43%
MIT

openduck

Data Engineering · Databases

570

OpenDuck brings MotherDuck-style cloud capabilities to self-hosted DuckDB — attach remote databases, run hybrid queries across local and remote nodes, and own your data with an open gRPC and Arrow IPC protocol.

View details
26
Repo Health
76
Technical
76
Dependency
Built with
C++ 43%
Rust 38%
HTML 13%
Updated 5 months ago

About Data Engineering

Data engineering focuses on building and maintaining robust data pipelines that enable organizations to make data-driven decisions. These tools are essential for turning raw data into actionable insights, automating data workflows, and ensuring data quality.

Typical features within this category include:

  • Data Integration: Connecting to various data sources (databases, APIs, cloud storage) and ingesting data.
  • Data Transformation: Cleaning, validating, enriching, and transforming data into usable formats using techniques like ETL (Extract, Transform, Load).
  • Data Storage: Managing and organizing data in efficient and scalable storage systems (data warehouses, data lakes).
  • Data Pipeline Automation: Scheduling and monitoring data workflows to ensure reliability and consistency.
  • Data Quality & Governance: Implementing checks and balances to maintain data accuracy, completeness, and security.

Data engineering solves critical problems such as siloed data, inefficient workflows, and a lack of reliable data for analytics. By streamlining the data process, organizations can unlock business value faster, improve decision-making accuracy, and gain a competitive edge. Furthermore, robust data pipelines are foundational for machine learning initiatives, enabling teams to build and deploy predictive models with confidence.

Join founders buildingwith open source

Opinionated takes, migration guides, cost-saving tips, and insights from the open source ecosystem.

Subscribe on Substack
Join 750+ subscribers