Scoring a prompt library without a beauty contest
A prompt library that is judged by whether it “sounds right” will be rewritten every time a new person joins the channel. Prompt Systems therefore scores the library as a set: each template has a brief, a version stamp, an owner, and a twenty-case test. The score on the rubric is for the set, not for the prettiest instruction.
In the lab we freeze the template, run the twenty cases, and mark factual control and reproducibility in the same sitting. If the group wants to change the instruction, they change one variable and rerun. Cases that were added because they made the template look good are taken out. Cases that look ugly — truncated input, mixed languages, missing fields — stay in, because that is the work.
The artefact we accept at Band 3 is a folder a colleague can open on Monday: templates, cases, scores, and a one-line note on what failed. A screenshot of a clever chat is not a library. We say this in the first hour because people often arrive with screenshots.