A rubric for judging what a skill produced
Some qualities cannot be checked by a script: does the screen look designed, is the copy specific. A five-line rubric makes that judgement repeatable.
Scripts can check that a token was used. They cannot tell you whether a page looks designed. For that you need a person, and a person needs a short list so two reviews mean the same thing.
A rubric for a design skill
- Would a stranger say this page had a designer? Yes or no.
- Is the accent used only on the primary action?
- Is every number in a mono, tabular face?
- Is the hero showing the product rather than an illustration?
- Is there anything from the skill's banned list?
Each question maps to a rule in the skill, so a failing answer points straight at what to fix. The anti-generic list behind several of these is in why AI-generated UI all looks the same. Pair the rubric with automated property checks from golden outputs.
Questions
Why use a rubric instead of just looking?
Because looking drifts. A short rubric makes two reviews comparable, and makes it obvious which rule the output broke.
How long should a rubric be?
Five or six yes-or-no questions drawn from the skill's own rules. Longer rubrics stop being used.
Can a model grade against the rubric?
It can help as a first pass, but check its judgements against your own on a few cases before trusting it on many.