Open Source Data Engineering Apps
Discover open source data engineering tools for building reliable data pipelines, ETL processes & scalable data storage. Unlock the power of your data!
Apps in Data Engineering
Apache Airflow
Data Engineering
Define, schedule, and monitor complex data workflows as Python code — with a powerful UI, 80+ provider integrations, and battle-tested scalability across thousands of production deployments.
Argo Workflows
Devops · Data Engineering
The most popular Kubernetes-native workflow engine for orchestrating containerized DAGs, ML pipelines, CI/CD, and parallel batch jobs at scale.
Airbyte
Developer Tools · Data Engineering
Open-source ELT platform with 600+ connectors for moving data from any source to warehouses, lakes, and AI agents.
ClickHouse
Databases · Analytics · Data Engineering
Open-source column-oriented database that delivers real-time analytical queries on petabyte-scale data with millisecond latency.
Kestra
Devops · Data Engineering · Automation
Event-driven orchestration platform for data, AI, and infrastructure workflows — define everything in YAML, run anywhere at scale.
Databend
Databases · Data Engineering
Open-source enterprise data warehouse unifying analytics, vector search, full-text search, and AI agent orchestration in a single Rust-built engine on S3.
DataEase
Analytics · Data Engineering · AI Assistants
Open-source BI tool with drag-and-drop dashboards, 20+ data source connectors, and AI-powered natural language queries — a self-hosted alternative to Tableau.
Label Studio
AI Development · Data Engineering
Label Studio is an open-source, multi-type data labeling platform that lets teams annotate images, text, audio, video, and time series data with a configurable XML-based UI and export annotations in formats ready for any ML framework.
Lightdash
Analytics · Data Engineering
The open-source Looker alternative that turns your dbt project's metrics and dimensions into governed, self-serve charts and dashboards — no license key required.
Dolt
Databases · Data Engineering · Developer Tools
The SQL database you can branch, merge, diff, and clone — Git for your data, MySQL-compatible and ready for multi-agent AI workflows.
WrenAI
Analytics · AI Agents · Data Engineering
Open-source GenBI engine that lets AI agents turn natural-language questions into governed SQL, charts, and shareable dashboards across 20+ data sources — no vendor lock-in, no black-box prompts.
marimo
Developer Tools · Data Engineering
A reactive Python notebook that eliminates hidden state, runs reproducibly, and deploys as a web app or script — stored as pure Python, built for the AI era.
Jitsu
Data Engineering
Open-source, fully-scriptable data ingestion engine that streams events from web, apps, and APIs to any data warehouse in real time.
PeerDB
Data Engineering · Databases
Postgres-native ETL that streams change data capture in real time to Snowflake, BigQuery, ClickHouse, S3, and Kafka — up to 10x faster than general-purpose pipelines, managed through a familiar Postgres SQL interface.
Timeplus Proton
Data Engineering · Analytics
Single C++ binary SQL engine for real-time stream processing, ETL, and analytics on Kafka, Redpanda, and ClickHouse with sub-millisecond latency.
About Data Engineering
Data engineering focuses on building and maintaining robust data pipelines that enable organizations to make data-driven decisions. These tools are essential for turning raw data into actionable insights, automating data workflows, and ensuring data quality.
Typical features within this category include:
- Data Integration: Connecting to various data sources (databases, APIs, cloud storage) and ingesting data.
- Data Transformation: Cleaning, validating, enriching, and transforming data into usable formats using techniques like ETL (Extract, Transform, Load).
- Data Storage: Managing and organizing data in efficient and scalable storage systems (data warehouses, data lakes).
- Data Pipeline Automation: Scheduling and monitoring data workflows to ensure reliability and consistency.
- Data Quality & Governance: Implementing checks and balances to maintain data accuracy, completeness, and security.
Data engineering solves critical problems such as siloed data, inefficient workflows, and a lack of reliable data for analytics. By streamlining the data process, organizations can unlock business value faster, improve decision-making accuracy, and gain a competitive edge. Furthermore, robust data pipelines are foundational for machine learning initiatives, enabling teams to build and deploy predictive models with confidence.