How scoring works
What an Updraft score is, what produces it, and what it does not claim.
The test
Every skill is run against real briefs for one kind of marketing work. Its output is checked against a written standard — hard rules a machine can verify, and judgement calls a model referee decides — and compared with what the same model produces from the same brief with no skill attached.
That bare-model control is the point. A skill that only reaches where the plain model already reaches is a cost, not a capability, and the “vs bare model” figure on every report says which side of that line it landed.
The six measures
- Quality
How much of your written standard the output actually met, checked automatically and by a second model. - Efficiency
What it costs to run: tokens per piece of work, and how many attempts it took to get something acceptable. - Bloat
How much weight the skill adds — its length and the number of files someone has to read and maintain. Higher is leaner. - Conflicts
Whether it collides with something you already run or already believe. Higher means fewer collisions. - Your preference
How often you picked this output in the blind comparison, without knowing which one it was. - Craft
How close to shippable it was, read from what you said you would still change. The softest number here.
Four are measured by machine when the assessment runs. Two — which output a person actually prefers, and how close it got to shippable — need a human blind review, so a score marked “machine score” is the four alone and says so.
Benchmark v0
Scores tagged were measured against briefs and standards drafted by Updraft itself, not yet against real published work. They are honest about their method and provisional about their basis; as real benchmarks replace the drafts, evaluations re-run and the tag goes.
What runs it
Generation, refereeing and the control all run on Claude, and each report states its exact models and effort settings at the top — including whether the call went to the API directly. The referee judging output quality is itself a model, so treat judgement checks as a strong opinion with the working shown, not ground truth. The blind human comparison exists precisely because that is not enough.
No score is ever adopted automatically anywhere: a person makes every decision that matters.