Open Source Data Engineering Apps

Discover open source data engineering tools for building reliable data pipelines, ETL processes & scalable data storage. Unlock the power of your data!

36 apps available

Apps in Data Engineering

TypeScript
39%
Apache 2.0

Label Studio

AI Development · Data Engineering

28,358

Label Studio is an open-source, multi-type data labeling platform that lets teams annotate images, text, audio, video, and time series data with a configurable XML-based UI and export annotations in formats ready for any ML framework.

View details
93
Repo Health
87
Technical
67
Dependency
Built with
TypeScript 39%
JavaScript 27%
Python 25%
Updated 5 days ago
TypeScript
97%
Other

Lightdash

Analytics · Data Engineering

6,166

The open-source Looker alternative that turns your dbt project's metrics and dimensions into governed, self-serve charts and dashboards — no license key required.

View details
93
Repo Health
84
Technical
64
Dependency
Built with
TypeScript 97%
Updated 5 days ago
TypeScript
43%
Apache 2.0

OpenMetadata

AI Development · Analytics · Data Engineering

15,344

Open-source metadata platform that unifies data catalog, lineage, quality, and governance into a single searchable graph, with an MCP server that gives AI agents governed access to that context.

View details
93
Repo Health
85
Technical
74
Dependency
Built with
TypeScript 43%
Java 36%
Python 18%
Updated 5 days ago
TypeScript
82%
MIT

evidence

Analytics · Data Engineering

6,962

Turn SQL queries and markdown files into polished, interactive data apps and business intelligence reports — no drag-and-drop, no GUI, just code.

View details
90
Repo Health
79
Technical
64
Dependency
Built with
TypeScript 82%
Svelte 17%
Updated 1 weeks ago
TypeScript
62%
MIT

Jitsu

Data Engineering

5,094

Open-source, fully-scriptable data ingestion engine that streams events from web, apps, and APIs to any data warehouse in real time.

View details
88
Repo Health
79
Technical
66
Dependency
Built with
TypeScript 62%
Go 36%
Updated 1 weeks ago
TypeScript
84%
Apache 2.0

ktx

AI Development · Analytics · Data Engineering

1,603

ktx builds a self-improving context layer over your data warehouse so AI agents like Claude Code and Codex query it with approved metric definitions instead of reinventing SQL logic from scratch.

View details
66
Repo Health
85
Technical
72
Dependency
Built with
TypeScript 84%
Updated 3 weeks ago
TypeScript
94%
Apache 2.0

reader

Data Engineering · Developer Tools

561

Production-grade open source web scraping engine that turns any URL into clean markdown for AI agents — with built-in anti-bot bypass, proxy rotation, and browser session management.

View details
59
Repo Health
79
Technical
76
Dependency
Built with
TypeScript 94%
Updated 1 months ago
TypeScript
96%
Other

superglue

AI Agents · Data Engineering · Developer Tools

2,063

superglue is an AI-agent-driven integration engine that turns plain-English descriptions of enterprise systems into production-grade API tools, ERP/CRM connectors, and data pipelines — self-hosted or cloud, Y Combinator-backed (W25).

View details
49
Repo Health
79
Technical
67
Dependency
Built with
TypeScript 96%
Updated 1 months ago
TypeScript
79%
MIT

Trench

Analytics · Data Engineering · Monitoring

1,663

Open-source event tracking infrastructure built on Kafka and ClickHouse that handles thousands of events per second on a single node, with full Segment API compatibility and no cookies.

View details
44
Repo Health
81
Technical
73
Dependency
Built with
TypeScript 79%
MDX 15%
Updated 6 months ago

About Data Engineering

Data engineering focuses on building and maintaining robust data pipelines that enable organizations to make data-driven decisions. These tools are essential for turning raw data into actionable insights, automating data workflows, and ensuring data quality.

Typical features within this category include:

  • Data Integration: Connecting to various data sources (databases, APIs, cloud storage) and ingesting data.
  • Data Transformation: Cleaning, validating, enriching, and transforming data into usable formats using techniques like ETL (Extract, Transform, Load).
  • Data Storage: Managing and organizing data in efficient and scalable storage systems (data warehouses, data lakes).
  • Data Pipeline Automation: Scheduling and monitoring data workflows to ensure reliability and consistency.
  • Data Quality & Governance: Implementing checks and balances to maintain data accuracy, completeness, and security.

Data engineering solves critical problems such as siloed data, inefficient workflows, and a lack of reliable data for analytics. By streamlining the data process, organizations can unlock business value faster, improve decision-making accuracy, and gain a competitive edge. Furthermore, robust data pipelines are foundational for machine learning initiatives, enabling teams to build and deploy predictive models with confidence.

Join founders buildingwith open source

Opinionated takes, migration guides, cost-saving tips, and insights from the open source ecosystem.

Subscribe on Substack
Join 750+ subscribers