Flow lets AI coding agents take tasks off your backlog, build them, test them and open the pull request on their own. You say what you want, and you approve the merge. Everything in between is checked by CI, not by you staring at diffs.
The problem
The agent says it’s finished. Then you find the test it skipped, the file it shouldn’t have touched, the edge case it waved through.
Every diff lands on you, so the time you saved writing code goes straight into reading it. Two projects in, you’re the bottleneck.
Context lives in chat windows. Sessions end with a list of questions. Close the laptop and the work stops.
How it works
Flow is a protocol that lives inside your repo. Tasks are plain files, coordination happens in git, and the rules are enforced by GitHub Actions, where a prompt can’t talk its way past them.
Write down the outcome and why it matters, or just file an issue. Agents turn it into small tasks for you to approve. Each one has checkable acceptance criteria, a list of the files it’s allowed to change, and the goal from your vision that it serves.
A fresh agent picks up each task, builds it on a branch and opens a pull request. Then the gate takes over:
You get a PR where every criterion is ticked and linked to the test that proves it. Read it, merge it, move on. The merge always stays with a human.
It starts with a vision
You write a short VISION.md once: your goals and your non-goals, in your own words. Every task has to name the goal it serves, a task pointing at a goal that doesn’t exist is flagged, and a weekly audit reports when the work starts drifting from what you said you wanted. Agents work fast. The vision keeps them going somewhere.
What lands on your desk
By the time a PR reaches you, the agent has linked every acceptance criterion to the test that proves it, and three independent reviewers have posted their verdicts. These three examples are real pull requests from Flow’s own repo, shortened.
[flow-0100] A repo sets its review diff limit in config.yml #136
Task: .flow/tasks/flow-0100-review-diff-limit-in-config.md
Repos that legitimately open large PRs went red because the review limit was fixed at 300 KB. It’s now a per-repo setting, read from the base branch so a PR can’t raise its own limit.
criterion 1: … a 789 000-byte diff survives wholecriterion 2: no key and no env override — the limit is still 300 000criterion 4: … and 4 more65/65 pass in flow-review.test.mjs. 1,448 pass in the full suite.
All 8 acceptance criteria have a proving test that asserts the actual outcome, not just that nothing threw. I re-ran the suite with no ambient environment variables: 65/65 pass.
The precedence is implemented as specified, and the ceiling applies to all three sources. A bad value fails loudly, naming the key.
No High or Critical findings. I checked that a PR can’t raise its own limit: the reviewer reads the base branch’s config, not the PR’s.
Your turn. One read, one click. Merging closes the task automatically.
Merge pull request[flow-0101] code-review and security refuse to pass a truncated diff #134
| # | Criterion | Status |
|---|---|---|
| 1 | All three reviewer prompts carry the same truncation rule | PROVEN |
| 2 | The rule forbids PASS and makes the verdict name the truncation | PROVEN |
| 3 | The changelog entry exists and says no caller action is needed | UNPROVEN |
The changelog file does exist and is correct. But no test proves it, and a criterion without a proving test doesn’t pass.
Criterion 3 is now proved by changes/flow-0101.md exists and states that no caller action is needed. I verified it independently rather than taking the test on trust.
You never saw the red. The gap was caught and fixed before the PR needed you.
Merge pull request[flow-0107] One test proves every task’s changelog entry #139
This PR is a draft and the task is blocked. The build is complete, but one acceptance criterion can’t be met inside the task’s declared scope.
The new test found a real defect in an older task: a changelog file that was declared but never written. The fix is one file, but that file isn’t in this task’s scope. Widening the scope is the orchestrator’s call, not the worker’s, so the task is blocked with both ways out written on it.
| Gate | Result |
|---|---|
| build | green |
| lint | green |
| test | 1,482 pass, 1 fail: the defect the new test exists to find |
| coverage | 95.89% |
One decision: widen the scope, or split the fix into its own task.
Merge blockedShortened from pull requests #136, #134 and #139 in Flow’s repo. Wording is condensed; verdicts, criteria, counts and timings are as they happened.
Automation
Flow ships as a set of small GitHub Actions workflows. The gate and task tracking are the core. The rest are optional, and the unattended ones (queue runner, triage and review) stay off until you switch them on.
Build, lint, tests, coverage and scope on every PR. It fails loudly, and a failure can’t be talked around.
The task file follows the PR: in progress, in review, done. Your backlog is always accurate without anyone updating it.
Opens the pull request as a draft the moment an agent pushes, so a stalled session can’t strand finished work.
Picks the top ready task, starts a fresh agent, and lets it build and open the PR. If a run produces nothing it can verify, it fails loudly instead of claiming success.
Three independent reviewers on every PR: QA, code review, and security when the change touches sensitive paths. Each runs in a clean session that never saw the code being written.
Reads new issues and posts a proposed task spec on each one. Add an “approved” label, even from your phone, and it becomes a ready task. It only listens to people who can already direct the repo.
Finds tasks stuck mid-flight, such as a crashed session or a PR with no status, and puts them back where they belong.
A read-only audit of recent work against your stated goals. It flags drift before it becomes a direction.
When a new Flow version is out, it opens a PR with the update. You review it like any other change.
Agent-agnostic
Flow’s rules live in one plain Markdown file in your repo. CLAUDE.md and AGENTS.md both point at it, so any coding agent that reads them can pick up a task and follow the same loop.
A human, Claude Code or any other agent meets exactly the same checks. The trust lives in CI, not in which model you picked.
Every review runs in a fresh session that never saw the code being written. You choose the review models per repo, and a separate model can handle security.
Next, the reviewers can run on a different agent from the builder, such as Codex reviewing work Claude built. The builder and the reviewer then come from different vendors.
The runners that work unattended (queue runner, triage, review) run on Claude Code today. Any agent can work tasks in your own sessions.
Why Flow
A PR can’t go green on “looks good to me”. Each acceptance criterion maps to a named test, and the reviewers run outside the agent that wrote the code.
Every task declares the files it may touch. Stray outside them and the PR fails, so there are no surprise refactors.
An optional scheduled runner picks up ready tasks overnight. A watchdog tells you if any automation stops running.
Tasks are Markdown files in your repo. There’s no service, database or dashboard. Stop using Flow and you keep everything.
The same protocol runs a static site and a full app. Updates reach every repo as a pull request you review.
When a task is unclear or a real decision comes up, the agent stops and asks. It doesn’t invent an answer to keep going.
Proof
Every change to Flow goes through the same gate it gives you. Here’s what that looks like as of v2.2.
It was hardened by real failures. Each one is now a rule the system enforces for you.
A parse error in an unchecked folder silently dropped messages for a week.
CI refuses any source folder it can’t check.
An agent finished the work, then never opened the pull request.
A workflow opens the PR, so a stalled agent can’t strand the work.
On an oversized change, two AI reviewers passed code they hadn’t fully read.
If a reviewer couldn’t see the whole diff, the check fails.
Compared
| An agent on its own | An agent on Flow | |
|---|---|---|
| Who decides it’s done | The agent, in its own words | CI, criterion by criterion, with a test for each |
| Scope | Whatever the agent decides to touch | Declared up front and enforced on every PR |
| Review | You, reading every line | Three independent AI checks first, then you merge |
| Parallel work | Agents collide in the same files | Tasks claimed in git, scopes kept apart |
| When you step away | Work stops | Ready tasks keep moving on a schedule |
| Direction | Drifts a little with every session | Every task names a goal from your vision, and a weekly audit flags drift |
| Bug reports | Sit in a list until you get to them | Get a proposed task spec every weekday morning, and one label from you makes it ready |
| Where it lives | Chat history | Plain files in your repo that you own |
Roadmap
Flow ships weekly. This is the current plan, and it changes as early users tell us what matters.
Is it for you?
In Claude Code, Codex, Cursor or any coding agent that can read a web page, inside your repo. It fetches the setup steps, tells you what it’s about to do, proposes your config and stops when it needs you. That’s usually two GitHub settings.
Set up Flow: getflow.now/start
The steps are plain text you can read first, pinned to v2.2.0: getflow.now/start.