Open-Source ETL That Runs on Your Servers, Not a Vendor Cloud
You've probably felt it: your data pipelines live in someone else's cloud, you pay per row, and every job you write is tangled up in a platform you can't take with you. Duckle is an open-source ETL platform built for teams who'd rather keep their pipelines on their own infrastructure—author them however you like, then deploy them to a server or cloud account you actually control.
What It Does
Duckle is an ETL platform that lets you author pipelines on a canvas, in Python, or in SQL—and then ship that same file to your own server or cloud account. When you're ready to run it, duckle-runner serve executes it headless on a schedule, whether that's in Docker or on a box you own. Alongside the runner you get a web console, roles, and an audit trail, so it's not just a script you kick off and forget about.
Under the hood, Duckle compiles your pipelines to SQL on DuckDB. It's designed to use every core you give the box, which means a bigger instance translates directly into a faster pipeline. The README cites a concrete number: 96 million rows out of Postgres to Parquet in 39.9 seconds. Every pipeline is a single file in git, and the project connects to 190 sources and destinations—databases, warehouses, SaaS apps, and the DuckDB ecosystem—all running locally on DuckDB. It's an independent project by SlothFlowLabs, built on the DuckDB engine but not affiliated with or endorsed by DuckDB Labs or MotherDuck.
Why It's Cool
-
Your infrastructure, your rules. The whole premise is pipelines you own. You author once and ship the same file to your own server or cloud account—no vendor cloud, no per-row billing, no lock-in. If that sentence doesn't resonate, you probably haven't been burned by a per-row bill yet.
-
One file in git. Every pipeline is a single file checked into version control. That's a quietly powerful design decision: the pipeline outlives whoever wrote it. When the person who built the job moves on, the next engineer inherits a readable file, not a mystery buried in a proprietary UI.
-
Three ways to author, one way to ship. Some people think in a visual canvas, some reach for Python, some just want SQL. Duckle lets you pick your lane and still deploy the same artifact. That flexibility matters when a team has mixed skill sets.
-
It scales with the machine you give it. Because it compiles to SQL on DuckDB and uses every core available, throwing a bigger instance at a slow pipeline is a legitimate answer. That 96-million-row benchmark is the kind of thing you can reason about when sizing hardware.
-
Operational features that usually show up late. A web console, roles, and an audit trail aren't glamorous, but they're the difference between a toy and something a team can run in production. Having them from the start is a good sign.
-
Broad connectivity. 190 sources and destinations is a lot of surface area, covering databases, warehouses, SaaS apps, and the DuckDB ecosystem—and it all runs locally.
How to Try It
The README points to a 60-second quickstart, so getting a pipeline running is meant to be fast. Here's the shape of it:
-
Download or install Duckle for your platform. It supports Windows, macOS, and Linux. The README links to a Download / Install section, and you can also build from source if you prefer.
-
Run your first pipeline. Follow the "Run your first pipeline" section in the README to go from install to a working job.
-
Deploy it. Once you've authored a pipeline, use
duckle-runner serveto run it headless on a schedule—in Docker or on a box you own, with the web console, roles, and audit trail in place.
The project is at beta (currently v0.7.4), so expect things to move. You can grab the repo, releases, and docs here:
- Repository: https://github.com/slothflowlabs/duckle
- Website: https://duckle.org/
- Community: the README links a Discord invite
Final Thoughts
Duckle is a solid fit for data engineers and small teams who are tired of renting their pipelines and want them running on hardware they control. The combination of a git-tracked single-file pipeline, DuckDB's local execution, and real operational features makes it worth a look—especially if per-row billing has ever made you wince. It's beta software, so treat it accordingly, but the design choices are thoughtful and the direction is clear. If owning your pipelines sounds better than renting them, this is a project to keep an eye on.
Follow @githubprojects for more developer tools and open source projects.