How it works

Full cycle — from task to provably closed goal. Each step is objective, each result is verifiable.

/goal — agent work
Task Verifiable criteria Deeplink to agent Work Evidence Judge Result
harness: proof of done
1 You work with your agent — the spec lands in Planner

You talk to your agent the way you always do. The agent puts the task artifact (spec) into Planner and attaches it to the right place in the project hierarchy.

Your process doesn't change. Planner is the agent's tool, not yours.
2 Planner helps make the task verifiable

On goal creation, Planner checks acceptance criteria. A linter catches vague wording. An LLM judge rejects subjective criteria — ones that depend on someone's opinion rather than an observable artifact. A red-team attacker looks for a scenario where all criteria formally pass but the goal isn't met. Weak criteria? Planner asks the agent to reformulate.

The task is verifiable before work begins. You set the intent — linter, judge, and red-team harden the criteria.
3 Hand the task to an agent via deeplink

A well-formed task can be assigned as a goal to any agent — via a /goals deeplink. The agent receives full context: criteria, dependencies, position in the hierarchy.

One link — the agent knows what to do and how it will be checked.
4 The agent does the work

The agent works natively via MCP — same Claude Code, same tools. Planner doesn't get in the way.

Zero overhead. The agent writes code, runs tests, deploys — business as usual.
5 The agent attaches evidence

For each acceptance criterion — a file: screenshot, log, test output. Not a report saying "I did it", but an artifact you can see with your own eyes.

Artifacts, not narratives. The agent proves, not tells.
6 The judge verifies — as many times as needed

An independent judge (a separate model with vision) examines each artifact against the criterion. Verdict: matches, mismatch, or weak. Mismatch? The agent gets the reason, fixes, and resubmits. The loop repeats until proven closure.

You don't babysit the agent — the system does it for you.
7 Better results — without botsitting

The goal is provably closed. The judge synthesizes an outcome: what was achieved and what it unblocks. Artifacts are available at any time. You got the result without spending time on manual verification.

Higher quality, less time spent.
What you get

Verifiable and provably closed

Tasks become verifiable and close with proof. Frees up your time for what actually matters.

Transparency. Proof of done.

Task artifacts are available at any time. Every closure is backed by evidence — not by the agent's word.

Connect to Claude