How scoring works

What an Updraft score is, what produces it, and what it does not claim.

The test

Every skill is run against real briefs for one kind of marketing work. Its output is checked against a written standard — hard rules a machine can verify, and judgement calls a model referee decides — and compared with what the same model produces from the same brief with no skill attached.

That bare-model control is the point. A skill that only reaches where the plain model already reaches is a cost, not a capability, and the “vs bare model” figure on every report says which side of that line it landed.

The six measures

Four are measured by machine when the assessment runs. Two — which output a person actually prefers, and how close it got to shippable — need a human blind review, so a score marked “machine score” is the four alone and says so.

Benchmark v0

Scores tagged Benchmark v0 were measured against briefs and standards drafted by Updraft itself, not yet against real published work. They are honest about their method and provisional about their basis; as real benchmarks replace the drafts, evaluations re-run and the tag goes.

What runs it

Generation, refereeing and the control all run on Claude, and each report states its exact models and effort settings at the top — including whether the call went to the API directly. The referee judging output quality is itself a model, so treat judgement checks as a strong opinion with the working shown, not ground truth. The blind human comparison exists precisely because that is not enough.

No score is ever adopted automatically anywhere: a person makes every decision that matters.

← All skills

Methodology · Updraft