Open-source developer tool

Dev Agent Autopilot

Codex plans. Claude Code builds. Codex reviews. You merge.

A command-line orchestrator for Claude Code, the Codex CLI and GitHub CI. Write a task in NEXT_TASK.md, run dev-autopilot run, and get back a reviewed, CI-checked pull request, without copying diffs and test failures between agents.

View on GitHub Quick start Download a release
77% less

agent-facing text on the LeanLoop benchmark, with every required-information check passing

0 auto-merges

merges, deploys, payments, secrets and DNS stay behind human gates

No API keys

it drives the claude, codex and gh CLIs you're already signed in to

MIT

licensed, tested on Linux, Windows and macOS

Stop being the glue between your AI agents

Without Autopilot

You ask Claude Code for a feature, copy the diff into Codex for a review, paste the review back, run the tests, paste the failures, push, wait for CI, paste those failures too, and open the pull request by hand. You are the message bus.

With Autopilot

You write the task and run one command. It returns immediately, and the work continues in a Claude Code background session until the pull request is ready for you. Close the terminal or reboot, run it again, and the same session picks up where it left off.

How one task runs

  1. Codex plans

    Approach, files to touch, risks and a test plan, in a read-only sandbox. Codex never edits code.

  2. Claude Code implements

    In a background session and its own worktree, running your real test, lint and build commands until they pass.

  3. Codex reviews the branch

    A fresh, independent codex review. The git diff sets the budget: docs-only changes skip it, risky ones get every round. A clean review ends the loop.

  4. Claude opens the pull request

    With a Codex review summary, then watches CI and fixes failures the task caused.

  5. You review and merge

    Autopilot never merges, deploys or touches secrets. It stops and reports a blocker instead.

Try it

You need Node.js 22.13+, Git, the GitHub CLI, Claude Code (2.1.139+) and the Codex CLI, each signed in. Autopilot isn't on the npm registry; install a tagged release straight from GitHub.

# Try a command without installing
npx --yes github:iammurtaza53/dev-agent-autopilot#v0.4.2 --help

# Install it (Autopilot sessions call dev-autopilot by name)
npm install -g github:iammurtaza53/dev-agent-autopilot#v0.4.2
dev-autopilot install-reviewer

# Onboard a project once, then commit the two files it writes
cd your-project
dev-autopilot init
dev-autopilot doctor

# Describe the work in NEXT_TASK.md, commit it, then
dev-autopilot run
dev-autopilot status

Want to see it work first? The demo project is a tiny Node app with a ready-made task and takes about ten minutes.

Why it's different

Real checks, not self-reports

Your own deterministic commands and GitHub CI decide when the work is done, not the agent's say-so.

Two vendors keep each other honest

One company's model writes the code; another's plans and reviews it.

LeanLoop: evidence, not history

Context Capsules, quiet checks and delta resumes keep what reaches the agents small, without an extra model and without hiding a failure.

No hard-coded model IDs

Claude and Codex use whatever models you've configured, so it doesn't go stale when new ones ship.

Optional trust-handoff gate

Switch on HostLatch and every check flags agent-written changes that could later run with your authority.

Small on purpose

A thin Node launcher over Claude Code's native background agents and Codex's native CLI. No custom agent runtime.

Who does what

PartOwns
Dev Agent AutopilotThe task lifecycle: launching and resuming the session, the workflow rule Claude follows, per-session permissions, the bounded review loop, status and the human gates
Claude CodeBackground sessions and worktrees, implementation, your checks, Git, the pull request and CI fixes
Codex CLIRead-only planning and the automated review loop
YouThe task, the review of the pull request, the merge and anything behind a human gate

Read more