Open Source Data Engineering Apps

Discover open source data engineering tools for building reliable data pipelines, ETL processes & scalable data storage. Unlock the power of your data!

36 apps available

Apps in Data Engineering

Python
89%
Apache 2.0

Apache Airflow

Data Engineering

46,995

Define, schedule, and monitor complex data workflows as Python code — with a powerful UI, 80+ provider integrations, and battle-tested scalability across thousands of production deployments.

View details
96
Repo Health
89
Technical
64
Dependency
Built with
Python 89%
Updated 5 days ago
Python
47%
Other

Airbyte

Data Engineering · Developer Tools

22,143

Open-source ELT platform with 600+ connectors for moving data from any source to warehouses, lakes, and AI agents.

View details
95
Repo Health
80
Technical
67
Dependency
Built with
Python 47%
Kotlin 43%
Updated 5 days ago
C++
68%
Apache 2.0

ClickHouse

Analytics · Data Engineering · Databases

50,116

Open-source column-oriented database that delivers real-time analytical queries on petabyte-scale data with millisecond latency.

View details
95
Repo Health
90
Technical
64
Dependency
Built with
C++ 68%
Python 14%
Updated 5 days ago
Java
70%
GPL 3.0

DataEase

AI Assistants · Analytics · Data Engineering

24,558

Open-source BI tool with drag-and-drop dashboards, 20+ data source connectors, and AI-powered natural language queries — a self-hosted alternative to Tableau.

View details
94
Repo Health
71
Technical
65
Dependency
Built with
Java 70%
Vue 29%
Updated 5 days ago
Java
58%
Apache 2.0

Kestra

Automation · Data Engineering · Devops

28,388

Event-driven orchestration platform for data, AI, and infrastructure workflows — define everything in YAML, run anywhere at scale.

View details
93
Repo Health
81
Technical
72
Dependency
Built with
Java 58%
TypeScript 26%
Vue 15%
Updated 1 weeks ago
Python
46%
Other

Redash

Analytics · Data Engineering

28,817

Redash lets anyone connect to 35+ SQL and NoSQL data sources, write a query in the browser, and turn the result into a shared dashboard — no separate BI suite required.

View details
92
Repo Health
74
Technical
60
Dependency
Built with
Python 46%
JavaScript 30%
TypeScript 17%
Updated 5 days ago
Python
65%
Apache 2.0

WrenAI

AI Agents · Analytics · Data Engineering

17,763

Open-source GenBI engine that lets AI agents turn natural-language questions into governed SQL, charts, and shareable dashboards across 20+ data sources — no vendor lock-in, no black-box prompts.

View details
90
Repo Health
91
Technical
69
Dependency
Built with
Python 65%
Rust 32%
Updated 1 weeks ago
Python
62%
Apache 2.0

marimo

Data Engineering · Developer Tools

22,918

A reactive Python notebook that eliminates hidden state, runs reproducibly, and deploys as a web app or script — stored as pure Python, built for the AI era.

View details
89
Repo Health
91
Technical
65
Dependency
Built with
Python 62%
TypeScript 37%
Updated 6 days ago
C++
75%
Apache 2.0

Timeplus Proton

Analytics · Data Engineering

2,262

Single C++ binary SQL engine for real-time stream processing, ETL, and analytics on Kafka, Redpanda, and ClickHouse with sub-millisecond latency.

View details
88
Repo Health
82
Technical
68
Dependency
Built with
C++ 75%
Python 11%
Updated 1 weeks ago
Python
88%
Apache 2.0

sirchmunk

AI Development · Data Engineering

1,351

Drop your files and search them instantly — no vector DB, no indexing pipeline, just raw data queried by a self-evolving intelligence layer.

View details
84
Repo Health
70
Technical
72
Dependency
Built with
Python 88%
TypeScript 11%
Updated 1 weeks ago
Python
65%
MIT

Flowfile

Data Engineering

363

Visual ETL that compiles to Polars — build pipelines on a canvas, export as standalone Python, and run anywhere without platform lock-in.

View details
83
Repo Health
81
Technical
66
Dependency
Built with
Python 65%
Vue 17%
TypeScript 17%
Updated 6 days ago
Python
58%
MIT

Docglow

Data Engineering

147

A next-generation documentation site generator for dbt Core projects — lineage explorer, health scoring, and full-text search for teams without access to dbt Cloud's built-in docs features.

View details
67
Repo Health
65
Technical
82
Dependency
Built with
Python 58%
TypeScript 42%
Updated 1 weeks ago
Python
59%
Apache 2.0

argilla

AI Development · Data Engineering

5,125

Collaborate on high-quality AI training data with a self-hosted annotation platform built for LLMs, NLP, and multimodal models.

View details
65
Repo Health
81
Technical
61
Dependency
Built with
Python 59%
Jupyter Notebook 21%
Updated 1 weeks ago
Python
94%
Apache 2.0

SWIRL

Data Engineering · Databases · Search

3,047

Federated AI search and RAG across 100+ enterprise sources—no data extraction, no vector database required.

View details
62
Repo Health
83
Technical
65
Dependency
Built with
Python 94%
Updated 1 weeks ago
Java
34%
Apache 2.0

Enso

Analytics · Data Engineering · Low Code Platforms

7,441

A visual and textual programming platform for data prep and analysis where the node graph and the underlying Enso code are always perfectly in sync, built by an Alteryx co-founder on a GraalVM engine.

View details
59
Repo Health
90
Technical
61
Dependency
Built with
Java 34%
TypeScript 27%
Scala 26%
Updated 1 months ago

About Data Engineering

Data engineering focuses on building and maintaining robust data pipelines that enable organizations to make data-driven decisions. These tools are essential for turning raw data into actionable insights, automating data workflows, and ensuring data quality.

Typical features within this category include:

  • Data Integration: Connecting to various data sources (databases, APIs, cloud storage) and ingesting data.
  • Data Transformation: Cleaning, validating, enriching, and transforming data into usable formats using techniques like ETL (Extract, Transform, Load).
  • Data Storage: Managing and organizing data in efficient and scalable storage systems (data warehouses, data lakes).
  • Data Pipeline Automation: Scheduling and monitoring data workflows to ensure reliability and consistency.
  • Data Quality & Governance: Implementing checks and balances to maintain data accuracy, completeness, and security.

Data engineering solves critical problems such as siloed data, inefficient workflows, and a lack of reliable data for analytics. By streamlining the data process, organizations can unlock business value faster, improve decision-making accuracy, and gain a competitive edge. Furthermore, robust data pipelines are foundational for machine learning initiatives, enabling teams to build and deploy predictive models with confidence.

Join founders buildingwith open source

Opinionated takes, migration guides, cost-saving tips, and insights from the open source ecosystem.

Subscribe on Substack
Join 750+ subscribers