Buyer manual · version 1.2 · September 2026

The Bench Book. Every word the judges listen for.

What to type, what happens, when to use it and what it costs. Keep this page open beside Claude Code, Codex or ChatGPT. Everything here works the same in all three editions, and the pass mark is always computed in code, so nothing talks its way through.

Chapter 01 · New in version 1.2
NEW

The low-cost first look

Speed Scan. Every expert, one fast read.

Type speed scan or quick scan

A full panel is the gold standard, but it is thorough, and thorough uses a lot of your AI plan. Speed Scan is built for everyday checks and for anyone on a smaller plan.

ONE judge takes every expert's seat for your kind of work, reads it once, scores it seat by seat, and hands you the top five fixes in order of importance. No chairman, no re-judging rounds. The same pass-mark maths still runs on every seat.

Use it on drafts, on small changes, on anything you want a fast honest read on. Then send in the full panel before anything important goes live.

Usage, measured

about 123k

tokens for a 5-seat sales page

A full panel run

1.75M to 2.3M

tokens across 2 to 3 rounds

Time

about 3 min

one worker instead of 12 to 18

Tested on the same sales page the full panel had judged: every seat's scan score landed within half a point of the full panel's, with the same top fixes. Roughly a fifteenth to a twentieth of the usage.

What you get

  • Every expert on your panel weighs in, not just one
  • A score out of 10 for each seat
  • The five fixes that matter most, ranked, each with the exact change to make
  • A plain-English summary of where the work stands

What you give up

  • The judges can see each other's thinking. In the full panel they are sealed off
  • No second round, so no proof your fixes actually worked
  • No chairman striking weak findings, so a little more noise gets through
The one rule

A Speed Scan can never mark your work CLEARED. That is locked in the code, not left to the AI. The best a scan can say is "looks ready for the full panel". Only the full panel can clear work.

Chapter 02 · The commands

Six things you can say, and exactly what each one does.

You never need more than a sentence. Finish a piece of work, then type one of these. Run modes and panel sizes combine: "relentless, full panel" or "speed scan this email" both work.

01NEW

Speed Scan

speed scan
What happens
One judge sits in every seat of your panel, one pass, top five fixes.
Use it for
Drafts, small changes, daily checks, smaller AI plans.
Cost
About 120k tokens. The cheapest way to hear every expert.
Ends with
A Speed Scan card. Never a CLEARED verdict. Full details in chapter 01.
02DEFAULT

The standard run

/perfection send in the judges
What happens
The right panel for your work judges it, rejects it with a fix list, the fixes get made, and they judge it again. Up to 3 rounds.
Use it for
Most finished work: a page, an email, an app screen, a script.
Cost
Roughly 95k tokens per judge per round. It tells you its estimate before it starts.
Ends with
CLEARED, or an honest scorecard of what is still short of the bar after round 3. You decide what happens next.
03HARDER

Relentless

relentless
What happens
Up to 5 rounds. Rounds 4 and 5 bring in a more powerful judging model, and a stronger worker makes the fixes.
Use it for
Work that has to clear the bar: a launch page, a big email, a client deliverable.
Cost
More than the standard run. It warns you again before the stronger rounds begin.
Ends with
CLEARED, or a scorecard plus a ruling on every finding still open.
04NO CAP

Until it passes

until it passes
What happens
It keeps going until the work clears, with two safety brakes: a hard ceiling of 15 rounds, and a stall detector. If the score stops genuinely improving, it pauses and asks you.
At a pause
You choose: accept as is, override one specific finding (it is recorded as a risk you accepted), or keep going.
Use it for
The one piece that must be as good as it can possibly get.
Cost
The highest. You get a cost update every round, with the running total.
05LEAN

One judge

one judge
What happens
Only the single most important expert for your kind of work judges it, with no chairman. Still the full judge, fix, re-judge loop.
Use it for
Real rounds on a budget. Unlike Speed Scan, one judge CAN clear work.
Cost
About a quarter of a four-seat panel.
Good to know
Small single pieces (one email, one banner) already get one judge automatically.
06MAXIMUM

Full panel

full panel
What happens
Every expert for your kind of work, plus a chairman who merges their findings and strikes anything a judge cannot back up.
Use it for
Anything important, and always after a Speed Scan when you want the real verdict.
Cost
The most per round, and the most independent opinions.
Good to know
Sales pages, funnels and checkouts get the full panel automatically.
Chapter 03 · Which one to use

Pick by the situation, not by the feature.

Your situationType thisWhy
I just want a quick honest readspeed scanEvery expert, one pass, a fraction of the usage.
I'm on a smaller plan and worried about usagespeed scan, then one judgeScan to find the big fixes, then one judge for real rounds that can clear it.
I finished something and want it judged properly/perfectionThe right panel, fixes made for you, up to 3 rounds.
It's a sales page or checkout/perfectionMoney pages get the full panel automatically.
It's going live to my whole list tomorrowrelentlessMore rounds and stronger judges when it counts.
This one has to be the best it can beuntil it passesNo cap, with safety brakes and your say at every pause.
A scan said "looks ready"send in the full panelOnly the full panel can clear work.
Chapter 04 · Your first run

Four questions, once. Then every judge knows your business.

The very first time you run it, before judging anything, it asks you these in one message. Two minutes. Your answers are saved in a file called profile.md inside the skill folder.

  1. Who is your customer? Age, situation, what they have tried before, how skeptical they are.
  2. What do you sell, and at roughly what prices?
  3. How do you write? A few words, plus anything banned: emojis, certain words, dashes.
  4. Do you have brand colours or fonts? Hex codes if you know them, or "none yet".

From then on the skeptical buyer judges become your customer, the voice judges enforce your voice, and the design judges check your colours. Edit profile.md any time. Delete it and the next run asks the four questions again.

Got a DESIGN.md or a project brief in your project folder? The judges use it as the bar and it takes priority over the built-in standard.

Chapter 05 · Who sits on the panel

It works out what you built and calls the right experts.

You never pick judges by hand. It reads the work, names the type in one line, and sizes the panel. Say "one judge" or "full panel" to change it.

Size of the workPanelChairman
One small piece (one email, one banner, one short script)1 elite judgeNo
A standard build (landing page, app feature, deck, long script)3 judgesYes
A money page or a full build (sales page, funnel, checkout)The full panelYes
Kind of workThe experts it calls
Sales page / funnelDirect-Response MasterONE JUDGE PICK, Scarred Buyer, Funnel Mechanic, Design Taste Judge, Compliance Skeptic
Website / landing pageConversion CriticONE JUDGE PICK, Design Taste Judge, Technical QA Judge, Target Buyer
App / product screenProduct Design JudgeONE JUDGE PICK, UX Flow Judge, Technical QA Judge, Microcopy Judge
Marketing emailEmail Conversion JudgeONE JUDGE PICK, Brand Compliance Judge, Skeptical Subscriber
Video scriptRetention JudgeONE JUDGE PICK, Voice Judge, Conversion Judge
Deck / presentationStory JudgeONE JUDGE PICK, Slide Design Judge, Room Judge
Anything elseDomain Expert (the best person alive at that craft)ONE JUDGE PICK, Ruthless Skeptic, End-User Proxy

A Speed Scan always uses the whole list for your kind of work. Hearing every expert is the point of it.

Chapter 06 · Reading the verdict

Three marks decide it. Code does the maths, not the AI.

To be CLEARED, all three must be trueCleared

8.0+

Every judge's overall score.

7.0+

Every single thing each judge scores, with no weak spot hiding under a good average.

0

Hard fails. An automatic reject like invented proof, fake scarcity or a dead buy button.

You seeIt means
ClearedEvery judge passed it. You still get three watch-outs per judge, the things closest to failing, so a pass is never a blank rubber stamp.
Not clearedThe round cap was reached. You get a table of what is still wrong, how bad it is, what a 10 looks like, and the effort to fix it.
Speed scanA first look with a score and the top five fixes. Never a verdict.
StalledScores stopped genuinely improving, so it stopped wasting your usage and shows you the stuck findings.
DisputedIt checks every finding against your actual work before fixing. If a judge got a fact wrong, it disputes it with proof instead of "fixing" something that was fine.

It never touches your offer. Prices, guarantees, bonuses and access rules are yours alone. If a judge has a concern about them, it is passed to you word for word, never changed.

Chapter 07 · The three editions

Same judges, same pass mark, wherever you work.

All three are in your zip. Install help for each one is on your install page.

Flagship

Claude Code

Type /perfection or say send in the judges. Every judge is its own sealed-off worker and fixes are made for you between rounds. speed scan works here.

Full power

Codex

Say send in the judges. Each judge runs as a separate Codex process that cannot see the others. If Codex looks unsure, say: "read my perfection skill at ~/.agents/skills/perfection/SKILL.md and judge this." speed scan runs as one process.

Lite

ChatGPT

Paste the Lite document into a fresh chat, then paste your work. Judges go one after another and ChatGPT runs the pass mark in Python. You apply the fixes and paste the new version back. Type speed scan for the lean version.

Chapter 08 · When things go sideways

Quick fixes for the usual bumps.

It's using more of my plan than I expected

Start with speed scan to find the big fixes for a fraction of the usage. When you want real rounds on a budget, add one judge. Save the full panel for work that is about to go live.

A Speed Scan said "looks ready". Is it done?

Not yet. A scan can never clear work by design. Say send in the full panel for the real verdict.

It asked me the four questions again

That happens when profile.md is missing from the skill folder, usually after a fresh reinstall. Answer them again, or copy your old profile.md back in.

I want to change my answers about my customer or brand

Open profile.md in the skill folder and edit it, or delete it and the next run asks the questions fresh.

A judge flagged my price or guarantee

That is on purpose. The skill never changes your offer terms. Concerns about them come to you word for word, and the decision is yours.

Codex says Node.js is missing

Just reply "install Node for me" and Codex will do it. The Codex edition needs it to run the judges and the pass mark.

How do I get the latest version?

Paste your original install prompt again. It downloads the fresh zip and replaces the old copy. Your current version is listed on the install page.

Something else went wrong

Reply to your welcome email with what your AI said and we will get you sorted. Nothing you do here can break anything.