opensourceprojects.dev

A broadsheet for software that doesn't ask for your email

A terminal agent that reads your project and checks its own work
GitHub RepoImpressions3

Project Description

View on GitHub

A Terminal Agent That Actually Checks Its Own Work

You've probably tried pointing an AI at your codebase and watching it confidently break things. The hard part isn't getting a model to write code—it's getting it to read your project, make changes, run the tests, and notice when something went wrong. Codewhale is a terminal agent built around exactly that loop: read, edit, run, verify.

What It Does

Codewhale is an open-source agent that lives in your terminal. You point it at a project folder, pick a model, and give it a task. From there it can read your repository, edit files, run commands, and inspect the results—keeping at it until the goal is met. The README's own example is refreshingly concrete: Fix the failing tests and explain what changed.

It works with either a hosted model or a local one, so you're not locked into a single provider. The first run drops you straight into the composer rather than walking you through setup; until you connect a model, the launch screen just tells you "no model connected." You add a hosted key or pick a local runtime with /provider (or F3), and if Ollama is already running with a chat model, Codewhale will switch to it on its own.

It's distributed as a Rust CLI (there's a crates.io package) with an npm package as a secondary route, plus Docker, Nix, Scoop, Android/Termux, and an optional CNB mirror. The README is translated into a long list of languages, which suggests the project is thinking beyond a single audience.

Why It's Cool

It's designed to verify, not just generate. The core pitch isn't "AI writes your code." It's that the agent runs commands and inspects results as part of the job. That feedback loop is what separates a useful agent from an autocomplete that hallucinates APIs.

You can hand off parts of a bigger job. For larger work, you can give different parts of the task to agents with different models and roles. That's a genuinely interesting architecture choice—instead of one model doing everything, you can match the model to the subtask.

There's a real safety valve for exploration. /mode plan lets the agent explore your codebase without making file changes or executing shell commands. When you're ready for it to actually do things, /mode work flips it on. That two-mode split is a small feature with a big effect on how comfortable you'll feel letting an agent loose.

You don't have to open the TUI at all. If you'd rather script it, codewhale exec "..." runs a task headlessly. That makes it usable in pipelines and one-off shell invocations, not just interactive sessions.

The install path is honest about friction. The installer prints the line you need for your shell if ~/.local/bin isn't on your PATH yet, and there's a dedicated doc for it. The README also notes that the changelog describes an unreleased candidate that isn't in published downloads—a small detail, but it tells you the maintainers care about not overpromising.

Updates are handled. For direct installs, codewhale update handles it, and codewhale update --check lets you inspect before committing. The updater prints the executable path and preserves newer builds.

How to Try It

On macOS or Linux, install the official GitHub release:

curl -fsSL https://codewhale.net/install.sh | sh
"$HOME/.local/bin/codewhale"

If plain codewhale says "command not found," ~/.local/bin isn't on your PATH yet—run the line the installer printed for your shell, or check the PATH section in the install docs.

On Windows, grab the matching installer or archive from GitHub Releases.

Then:

  1. Open a terminal in your project folder and run codewhale.
  2. Choose your provider with /provider and your model with /model.
  3. Describe a concrete task, like Fix the failing tests and explain what changed.

If you'd rather skip the TUI, run it inline:

codewhale exec "fix the failing tests and explain what changed"

Shell completions are one command per shell: codewhale completion bash|zsh|fish|powershell|elvish. Full details live in the repository.

Final Thoughts

Codewhale is aimed at developers who already work in a terminal and want an agent that stays there—no IDE plugin, no browser tab. The self-checking loop and the plan/work mode split are the parts that matter most; they're what make an agent trustworthy enough to leave running. If you're comfortable with a CLI and you've been waiting for something that verifies its own work rather than just generating plausible-looking code, this is worth a look. Start with /mode plan on a repo you know well, and see how it reasons before you let it touch anything.


Follow @githubprojects for more developer tools and open source projects.

Back to Projects
Project ID: fad0f076-2e36-4b19-b7d5-72fdb73ec70bLast updated: October 1, 2026 at 06:03 AM