Open Source Data Engineering Apps

Discover open source data engineering tools for building reliable data pipelines, ETL processes & scalable data storage. Unlock the power of your data!

36 apps available

Apps in Data Engineering

Go
85%
Apache 2.0

Argo Workflows

Data Engineering · Devops

17,006

The most popular Kubernetes-native workflow engine for orchestrating containerized DAGs, ML pipelines, CI/CD, and parallel batch jobs at scale.

View details
96
Repo Health
90
Technical
68
Dependency
Built with
Go 85%
TypeScript 11%
Updated 1 weeks ago
Java
70%
GPL 3.0

DataEase

AI Assistants · Analytics · Data Engineering

24,558

Open-source BI tool with drag-and-drop dashboards, 20+ data source connectors, and AI-powered natural language queries — a self-hosted alternative to Tableau.

View details
94
Repo Health
71
Technical
65
Dependency
Built with
Java 70%
Vue 29%
Updated 1 weeks ago
Java
58%
Apache 2.0

Kestra

Automation · Data Engineering · Devops

28,388

Event-driven orchestration platform for data, AI, and infrastructure workflows — define everything in YAML, run anywhere at scale.

View details
93
Repo Health
81
Technical
72
Dependency
Built with
Java 58%
TypeScript 26%
Vue 15%
Updated 1 weeks ago
Go
80%
Apache 2.0

Dolt

Data Engineering · Databases · Developer Tools

24,532

The SQL database you can branch, merge, diff, and clone — Git for your data, MySQL-compatible and ready for multi-agent AI workflows.

View details
90
Repo Health
9
Technical
65
Dependency
Built with
Go 80%
Shell 19%
Updated 1 weeks ago
Go
81%
AGPL 3.0

PeerDB

Data Engineering · Databases

3,288

Postgres-native ETL that streams change data capture in real time to Snowflake, BigQuery, ClickHouse, S3, and Kafka — up to 10x faster than general-purpose pipelines, managed through a familiar Postgres SQL interface.

View details
88
Repo Health
76
Technical
66
Dependency
Built with
Go 81%
TypeScript 12%
Updated 1 weeks ago
Go
39%
Apache 2.0

Rill

Analytics · Data Engineering

2,914

The fastest BI tool for humans and agents — define metrics, models, and dashboards as code and query them instantly on ClickHouse or DuckDB.

View details
87
Repo Health
88
Technical
64
Dependency
Built with
Go 39%
TypeScript 38%
Svelte 21%
Updated 1 weeks ago
Go
84%
AGPL 3.0

Beta9

AI Development · Automation · Data Engineering

1,794

Run AI workloads at scale with a Pythonic serverless runtime that handles GPU inference, background jobs, and sandboxes with zero infrastructure overhead.

View details
85
Repo Health
78
Technical
66
Dependency
Built with
Go 84%
Python 15%
Updated 1 weeks ago
Go
54%
MPL 2.0

shaper

Analytics · Data Engineering

1,251

Build analytics dashboards, reports, and customer-facing analytics by writing pure SQL — powered by DuckDB and designed for self-hosted deployment.

View details
80
Repo Health
79
Technical
74
Dependency
Built with
Go 54%
TypeScript 45%
Updated 1 weeks ago
Java
34%
Apache 2.0

Enso

Analytics · Data Engineering · Low Code Platforms

7,441

A visual and textual programming platform for data prep and analysis where the node graph and the underlying Enso code are always perfectly in sync, built by an Alteryx co-founder on a GraalVM engine.

View details
59
Repo Health
90
Technical
61
Dependency
Built with
Java 34%
TypeScript 27%
Scala 26%
Updated 1 months ago

About Data Engineering

Data engineering focuses on building and maintaining robust data pipelines that enable organizations to make data-driven decisions. These tools are essential for turning raw data into actionable insights, automating data workflows, and ensuring data quality.

Typical features within this category include:

  • Data Integration: Connecting to various data sources (databases, APIs, cloud storage) and ingesting data.
  • Data Transformation: Cleaning, validating, enriching, and transforming data into usable formats using techniques like ETL (Extract, Transform, Load).
  • Data Storage: Managing and organizing data in efficient and scalable storage systems (data warehouses, data lakes).
  • Data Pipeline Automation: Scheduling and monitoring data workflows to ensure reliability and consistency.
  • Data Quality & Governance: Implementing checks and balances to maintain data accuracy, completeness, and security.

Data engineering solves critical problems such as siloed data, inefficient workflows, and a lack of reliable data for analytics. By streamlining the data process, organizations can unlock business value faster, improve decision-making accuracy, and gain a competitive edge. Furthermore, robust data pipelines are foundational for machine learning initiatives, enabling teams to build and deploy predictive models with confidence.

Join founders buildingwith open source

Opinionated takes, migration guides, cost-saving tips, and insights from the open source ecosystem.

Subscribe on Substack
Join 750+ subscribers