Dev Agent Autopilot
Codex plans. Claude Code builds. Codex reviews. You merge.
A command-line orchestrator for Claude Code, the Codex CLI and GitHub CI. Write a task in NEXT_TASK.md, run dev-autopilot run, and get back a reviewed, CI-checked pull request, without copying diffs and test failures between agents.
agent-facing text on the LeanLoop benchmark, with every required-information check passing
merges, deploys, payments, secrets and DNS stay behind human gates
it drives the claude, codex and gh CLIs you're already signed in to
licensed, tested on Linux, Windows and macOS
Stop being the glue between your AI agents
Without Autopilot
You ask Claude Code for a feature, copy the diff into Codex for a review, paste the review back, run the tests, paste the failures, push, wait for CI, paste those failures too, and open the pull request by hand. You are the message bus.
With Autopilot
You write the task and run one command. It returns immediately, and the work continues in a Claude Code background session until the pull request is ready for you. Close the terminal or reboot, run it again, and the same session picks up where it left off.
How one task runs
Codex plans
Approach, files to touch, risks and a test plan, in a read-only sandbox. Codex never edits code.
Claude Code implements
In a background session and its own worktree, running your real test, lint and build commands until they pass.
Codex reviews the branch
A fresh, independent
codex review. The git diff sets the budget: docs-only changes skip it, risky ones get every round. A clean review ends the loop.Claude opens the pull request
With a Codex review summary, then watches CI and fixes failures the task caused.
You review and merge
Autopilot never merges, deploys or touches secrets. It stops and reports a blocker instead.
Try it
You need Node.js 22.13+, Git, the GitHub CLI, Claude Code (2.1.139+) and the Codex CLI, each signed in. Autopilot isn't on the npm registry; install a tagged release straight from GitHub.
# Try a command without installing
npx --yes github:iammurtaza53/dev-agent-autopilot#v0.4.2 --help
# Install it (Autopilot sessions call dev-autopilot by name)
npm install -g github:iammurtaza53/dev-agent-autopilot#v0.4.2
dev-autopilot install-reviewer
# Onboard a project once, then commit the two files it writes
cd your-project
dev-autopilot init
dev-autopilot doctor
# Describe the work in NEXT_TASK.md, commit it, then
dev-autopilot run
dev-autopilot status
Want to see it work first? The demo project is a tiny Node app with a ready-made task and takes about ten minutes.
Why it's different
Real checks, not self-reports
Your own deterministic commands and GitHub CI decide when the work is done, not the agent's say-so.
Two vendors keep each other honest
One company's model writes the code; another's plans and reviews it.
LeanLoop: evidence, not history
Context Capsules, quiet checks and delta resumes keep what reaches the agents small, without an extra model and without hiding a failure.
No hard-coded model IDs
Claude and Codex use whatever models you've configured, so it doesn't go stale when new ones ship.
Optional trust-handoff gate
Switch on HostLatch and every check flags agent-written changes that could later run with your authority.
Small on purpose
A thin Node launcher over Claude Code's native background agents and Codex's native CLI. No custom agent runtime.
Who does what
| Part | Owns |
|---|---|
| Dev Agent Autopilot | The task lifecycle: launching and resuming the session, the workflow rule Claude follows, per-session permissions, the bounded review loop, status and the human gates |
| Claude Code | Background sessions and worktrees, implementation, your checks, Git, the pull request and CI fixes |
| Codex CLI | Read-only planning and the automated review loop |
| You | The task, the review of the pull request, the merge and anything behind a human gate |
Read more
- README: requirements, configuration, commands and FAQ
- LeanLoop design: Context Capsules, Delta Resume, quiet checks and the adaptive review budget
- Benchmark: the method, the per-scenario results and the limits of the measurement
- The official Codex plugin: how Autopilot works alongside it
- Changelog and security policy
- HostLatch, the companion trust-handoff scanner for AI-written repositories