v2.2 Independent AI review on every pull request

Delegate the work, not the watching.

Flow lets AI coding agents take tasks off your backlog, build them, test them and open the pull request on their own. You say what you want, and you approve the merge. Everything in between is checked by CI, not by you staring at diffs.

Open source, Apache-2.0Runs on GitHub ActionsWorks with Claude Code and other agentsNothing to host

The problem

AI can write the code. You still end up babysitting it.

“Done” isn’t done

The agent says it’s finished. Then you find the test it skipped, the file it shouldn’t have touched, the edge case it waved through.

You become the reviewer

Every diff lands on you, so the time you saved writing code goes straight into reading it. Two projects in, you’re the bottleneck.

Nothing runs without you

Context lives in chat windows. Sessions end with a list of questions. Close the laptop and the work stops.

How it works

Two touchpoints. The rest runs itself.

Flow is a protocol that lives inside your repo. Tasks are plain files, coordination happens in git, and the rules are enforced by GitHub Actions, where a prompt can’t talk its way past them.

You

Say what you want, or what’s broken

Write down the outcome and why it matters, or just file an issue. Agents turn it into small tasks for you to approve. Each one has checkable acceptance criteria, a list of the files it’s allowed to change, and the goal from your vision that it serves.

Agents + CI

Agents build it, CI checks it

A fresh agent picks up each task, builds it on a branch and opens a pull request. Then the gate takes over:

  • Every criterion needs a passing test
  • Changes outside the task’s scope fail
  • Build, lint, tests and coverage must pass
  • Three independent AI reviewers check QA, code quality and security
You

Approve the merge

You get a PR where every criterion is ticked and linked to the test that proves it. Read it, merge it, move on. The merge always stays with a human.

It starts with a vision

Every change traces back to why you wanted it.

You write a short VISION.md once: your goals and your non-goals, in your own words. Every task has to name the goal it serves, a task pointing at a goal that doesn’t exist is flagged, and a weekly audit reports when the work starts drifting from what you said you wanted. Agents work fast. The vision keeps them going somewhere.

  1. VISION.mdG10: The gate tells the truthWhen CI says green, the work is actually right.
  2. Taskflow-0100 · serves G10A repo sets its own review size limit.
  3. Pull request#136 · 8 of 8 criteria provedThree reviewers passed. Merged by a human.

What lands on your desk

This is what you review.

By the time a PR reaches you, the agent has linked every acceptance criterion to the test that proves it, and three independent reviewers have posted their verdicts. These three examples are real pull requests from Flow’s own repo, shortened.

[flow-0100] A repo sets its review diff limit in config.yml #136

● OpenWorker session wants to merge 1 commit intomainfromflow/flow-0100-review-diff-limit-in-config
flow-open-pr opened this as a draft when the agent pushed its branch
AIworker wrote the description

Task: .flow/tasks/flow-0100-review-diff-limit-in-config.md

Repos that legitimately open large PRs went red because the review limit was fixed at 300 KB. It’s now a per-repo setting, read from the base branch so a PR can’t raise its own limit.

Acceptance criteria · 8 of 8
  • A 789 000-byte diff is not cut when the limit is set to 900 000proved by criterion 1: … a 789 000-byte diff survives whole
  • With no setting, the limit stays at 300 000proved by criterion 2: no key and no env override — the limit is still 300 000
  • An invalid or too-large value fails loudly, naming the settingproved by criterion 4: … and 4 more

65/65 pass in flow-review.test.mjs. 1,448 pass in the full suite.

Worker marked this ready for review
QAqabot reviewed in a separate session
QA verdict: PASS

All 8 acceptance criteria have a proving test that asserts the actual outcome, not just that nothing threw. I re-ran the suite with no ambient environment variables: 65/65 pass.

CRcode-reviewbot
code-review: PASS

The precedence is implemented as specified, and the ceiling applies to all three sources. A bad value fails loudly, naming the key.

SECsecuritybot
Security review: PASS

No High or Critical findings. I checked that a PR can’t raise its own limit: the reviewer reads the base branch’s config, not the PR’s.

✓
All checks have passed9 successful, 1 skipped
  • ✓flow-gates / gate build, lint, test, coverage
  • ✓flow-gates / touches diff stays inside the task’s declared files
  • ✓flow-gates / flow-tooling
  • ✓flow-gates / source-roots-plan
  • –flow-gates / source-root skipped, already covered by the gate
  • ✓flow-review / plan
  • ✓flow-review / qa
  • ✓flow-review / code-review
  • ✓flow-review / security
  • ✓flow-status / sync-status task moved to in_review

Your turn. One read, one click. Merging closes the task automatically.

Merge pull request

Shortened from pull requests #136, #134 and #139 in Flow’s repo. Wording is condensed; verdicts, criteria, counts and timings are as they happened.

Automation

Runners that keep the work moving without you.

Flow ships as a set of small GitHub Actions workflows. The gate and task tracking are the core. The rest are optional, and the unattended ones (queue runner, triage and review) stay off until you switch them on.

every pull request

Gate

Build, lint, tests, coverage and scope on every PR. It fails loudly, and a failure can’t be talked around.

PR opened, ready, merged

Status and done

The task file follows the PR: in progress, in review, done. Your backlog is always accurate without anyone updating it.

every push to a task branch

Open PR

Opens the pull request as a draft the moment an agent pushes, so a stalled session can’t strand finished work.

weekday mornings

Queue runner

Picks the top ready task, starts a fresh agent, and lets it build and open the PR. If a run produces nothing it can verify, it fails loudly instead of claiming success.

PR ready or updated

Review

Three independent reviewers on every PR: QA, code review, and security when the change touches sensitive paths. Each runs in a clean session that never saw the code being written.

weekday mornings

Triage

Reads new issues and posts a proposed task spec on each one. Add an “approved” label, even from your phone, and it becomes a ready task. It only listens to people who can already direct the repo.

every 3 hours

Recover

Finds tasks stuck mid-flight, such as a crashed session or a PR with no status, and puts them back where they belong.

weekly

Compass

A read-only audit of recent work against your stated goals. It flags drift before it becomes a direction.

weekly

Sync

When a new Flow version is out, it opens a PR with the update. You review it like any other change.

Core, installed by defaultOptional, add when you want itA watchdog also opens an issue if any scheduled runner stops running.

Agent-agnostic

Bring your own agent. Let a different one check it.

A protocol, not a plugin

Flow’s rules live in one plain Markdown file in your repo. CLAUDE.md and AGENTS.md both point at it, so any coding agent that reads them can pick up a task and follow the same loop.

The gate doesn’t care who wrote it

A human, Claude Code or any other agent meets exactly the same checks. The trust lives in CI, not in which model you picked.

The builder never grades its own work

Every review runs in a fresh session that never saw the code being written. You choose the review models per repo, and a separate model can handle security.

Review by a different vendor Coming

Next, the reviewers can run on a different agent from the builder, such as Codex reviewing work Claude built. The builder and the reviewer then come from different vendors.

Builds
Claude CodeCodexCursorYouany AGENTS.md agent

Gate
GitHub Actions, the same for everyone

Reviews
Claude, fresh sessionCodex · coming

Merges
You

The runners that work unattended (queue runner, triage, review) run on Claude Code today. Any agent can work tasks in your own sessions.

Why Flow

Built so you can trust a green tick.

TRUSTWORTHY

Done means done

A PR can’t go green on “looks good to me”. Each acceptance criterion maps to a named test, and the reviewers run outside the agent that wrote the code.

CONTAINED

Agents stay in their lane

Every task declares the files it may touch. Stray outside them and the PR fails, so there are no surprise refactors.

UNATTENDED

Work while you sleep

An optional scheduled runner picks up ready tasks overnight. A watchdog tells you if any automation stops running.

NO LOCK-IN

Nothing new to log into

Tasks are Markdown files in your repo. There’s no service, database or dashboard. Stop using Flow and you keep everything.

PORTFOLIO

One process, every project

The same protocol runs a static site and a full app. Updates reach every repo as a pull request you review.

HONEST

Agents ask instead of guessing

When a task is unclear or a real decision comes up, the agent stops and asks. It doesn’t invent an answer to keep going.

Proof

Flow is built with Flow.

Every change to Flow goes through the same gate it gives you. Here’s what that looks like as of v2.2.

82tasks shipped through its own gate
1,400+automated tests on Flow itself
96%line coverage, with a floor CI enforces

It was hardened by real failures. Each one is now a rule the system enforces for you.

A parse error in an unchecked folder silently dropped messages for a week.

CI refuses any source folder it can’t check.

An agent finished the work, then never opened the pull request.

A workflow opens the PR, so a stalled agent can’t strand the work.

On an oversized change, two AI reviewers passed code they hadn’t fully read.

If a reviewer couldn’t see the whole diff, the check fails.

Compared

Flow vs. prompting harder

An agent on its ownAn agent on Flow
Who decides it’s doneThe agent, in its own wordsCI, criterion by criterion, with a test for each
ScopeWhatever the agent decides to touchDeclared up front and enforced on every PR
ReviewYou, reading every lineThree independent AI checks first, then you merge
Parallel workAgents collide in the same filesTasks claimed in git, scopes kept apart
When you step awayWork stopsReady tasks keep moving on a schedule
DirectionDrifts a little with every sessionEvery task names a goal from your vision, and a weekly audit flags drift
Bug reportsSit in a list until you get to themGet a proposed task spec every weekday morning, and one label from you makes it ready
Where it livesChat historyPlain files in your repo that you own

Roadmap

What’s coming

Flow ships weekly. This is the current plan, and it changes as early users tell us what matters.

Now
  • A reading guide on every PROne comment telling you where to look: summary, riskiest spots, assumptions made and a quick smoke test.
  • Intents, end to endAn interview that turns “here’s what I want” into a written intent every task traces back to.
  • Runner controlsPause the overnight runner without pausing review, and schedule it in your own timezone.
Later
  • Mission controlOne view of what’s in flight across all your repos.
  • Real services in the gateDatabases and containers stood up in CI, so the gate can verify what your work needs.
  • Cost per taskSee what each task spent in AI tokens and CI minutes.

Is it for you?

Made for builders running more than they can watch.

A great fit if you

  • Are a solo builder or small team with several projects on the go
  • Already use GitHub and GitHub Actions
  • Want to hand real work to AI agents, not just autocomplete
  • Would rather trust a gate than read every diff

Good to know

  • Flow is early. It runs across its creator’s own products, and you’d be among the first outside teams
  • It needs GitHub Actions and has no hosted UI
  • It works best when work splits into small, well-scoped tasks

Tell your agent one line.

In Claude Code, Codex, Cursor or any coding agent that can read a web page, inside your repo. It fetches the setup steps, tells you what it’s about to do, proposes your config and stops when it needs you. That’s usually two GitHub settings.

Say this to your agent
Set up Flow: getflow.now/start
getflow.now/start
New repoIt lays down the task store, the gate and the workflows, proposes your config, then stops for the two GitHub settings only you can change.
Existing repoSame line. It merges Flow in beside your code, CI and CLAUDE.md without overwriting them, moves open tracker items into tasks, and asks before changing anything you wrote. Your application code isn’t touched.

The steps are plain text you can read first, pinned to v2.2.0: getflow.now/start.