Starting out
A rubric is a production tool before it is a quality claim
The reason to have defined axes is that it halves your time per job; the reason buyers like it is secondary.
A rubric is usually sold as a quality signal, and that is the smaller half of what it does. Its first job is to remove the decision about what to say, which is where most of the time in an unstructured assessment disappears.
Earners who work to fixed axes routinely report a large drop in minutes per job once the structure settles, and the drop comes almost entirely from not restarting from a blank page every time. The buyer-facing benefit is real and it arrives afterwards.
What the blank page actually costs
Time an unstructured written assessment and the minutes do not go where you expect. The writing is fast. Deciding what is worth mentioning, in what order, and how much to say about each thing is slow, and it is slow every single time because none of it carries over.
A rubric converts that decision into a lookup. You are no longer asking what to write about; you are asking what the answer is on axis three. That is a different cognitive task, and it is roughly the difference between composing and filling in.
The second cost of the blank page is variance. Two jobs done a fortnight apart end up covering different ground at different depths, and the buyer who books you twice notices. Consistency across jobs is worth more to a reputation than any single strong delivery, and a rubric is the cheapest way to get it.
Choosing the axes
Three to six. Below three the structure is not doing anything a paragraph could not; above six the buyer stops reading and you have built yourself a chore.
Each axis has to pass three tests.
It is observable. You can point at what in the submission produced the score. An axis you cannot evidence turns into a vibe with a number attached, and it is the one a buyer will query.
It is independent. If two axes always move together, they are one axis wearing two hats, and you are charging yourself twice for the same judgement. The usual offender is a pair where one is a restatement of the other in more flattering words.
It survives a bad case. Imagine the submission you least want to receive and check that the axis still produces a defensible answer. Axes that only work on good material collapse exactly when you need them most.
What the axes should not do is stray from description into anything that sounds like a claim about the person. That line matters more here than in most trades, and the piece on staying inside what you actually know is the longer treatment. The theory of what a scoring axis can and cannot legitimately measure belongs to the property that covers scoring as a subject, and Penis Rater's scores hub sets out the reasoning better than a pricing site should attempt to.
Designing the scale
Most people reach for ten because scores out of ten are the convention. Ten is a bad working scale and a fine display scale, which is a distinction worth keeping.
A ten-point scale asks you to distinguish a 6 from a 7 on every axis, on every job, forever. You cannot, nobody can, and the effort of pretending is a real fraction of your per-job time. Worse, the distinction is not stable across months: your 7 in March is somebody else's 6 and possibly your own 6 in September.
Score on five points internally. Anchor each point in a sentence you write once.
| Point | Anchor |
|---|---|
| 1 | Clearly weak on this axis, with an obvious specific reason |
| 2 | Below the middle, one identifiable shortcoming |
| 3 | Unremarkable, nothing to flag either way |
| 4 | Above the middle, one identifiable strength |
| 5 | Clearly strong on this axis, with an obvious specific reason |
The anchors do the work. Once each point has a written definition, scoring is matching rather than judging, and matching is fast and repeatable.
If the buyer expects a number out of ten, convert on display. Two of your points map to one of theirs, and you have kept a scale you can actually apply. What a resulting score means to the person receiving it is a different question and belongs to the platform side; Rate Cock's explainer on reading a score covers what buyers are told, which is useful context for how yours will be read.
Writing the justification
An axis with a number and no sentence is worth almost nothing. The sentence is the product.
Write one per axis, and give each one the same shape: what you observed, then what it implies. Two clauses. The observation has to be specific enough that it could only have been written about this submission, and the implication has to be general enough to be useful.
The reason to fix the shape is that fixed shapes are fast. You are not deciding how to phrase a judgement, you are filling two slots. This is the same principle as templates that speed you up without flattening the work, applied at sentence level rather than document level.
The trap is the reusable clause. A phrase you have written forty times reads as a phrase you have written forty times, and buyers detect it faster than you would think. The rule that holds: the structure repeats, the content never does. If a sentence could be pasted into another job unchanged, it is padding and it should be cut rather than reworded.
Scored work with written justifications prices above prose covering identical ground, which is the argument in written versus scored. The rubric is what makes the scored version cheap for you to produce.
Version it, and date the version
A rubric changes. You will drop an axis nobody engages with, split one that turned out to be two, and rewrite an anchor that kept producing arguments.
Give it a version number and a date, and keep the old versions. This costs one line in a text file and buys three things.
You can answer a returning buyer who asks why this assessment looks different from the one last spring. You can compare a job scored under v2 against a job scored under v2 rather than silently comparing across a definition change. And when you review your own numbers, you can see whether the format change coincided with the point where bookings moved.
Change one thing at a time and leave it for a stretch of jobs before changing another. A rubric that changes every fortnight is a rubric you are never getting faster at, and speed was the point.
The rubric is also a scoping document
This is the part people find by accident, usually after a job has gone sideways.
Published axes tell a buyer exactly what they are getting, which means they also tell a buyer what they are not getting. "Not covered by my axes" is a complete and unarguable answer to a request that arrives after the price was agreed. It is a far better answer than a negotiation, because it was visible before they booked.
That makes the rubric the backbone of your scope box. The axes, the scale, the length of each justification and the turnaround are four concrete facts a buyer can check, and four facts are what turns a scope box buyers actually read into something you can cite when the job drifts.
It also prices your extras cleanly. An additional axis on request is a defined unit of extra work with an obvious price. An unstructured "more detail please" is not.
Drift, and the calibration habit
Rubrics drift. Not by decision, by accumulation: you have seen four hundred submissions, your middle has moved, and a 3 now means something different from what it meant in your first month.
Drift is not a fault, it is what expertise looks like from the inside. It becomes a problem only when it happens invisibly, because then a repeat buyer gets two incompatible assessments and you cannot explain why.
The cheap correction is a calibration file. Keep three anonymised past submissions with their scores, one at each end and one in the middle, on generic material you own. Every couple of months, re-score them without looking at the old numbers, then compare.
If they match, you are stable. If they have moved a point in one direction, either the anchors need rewriting to say what you now mean, or you need to correct back. Either answer is fine. Not knowing which is happening is the bad outcome.
The formal name for this check is reliability, and the research literature has numbers for it. McHugh (2012) explains why raw percent agreement flatters raters: some agreement happens by chance, which is what Cohen's kappa was developed to correct for. Koo and Li (2016) suggest reading an intraclass correlation below 0.5 as poor and above 0.9 as excellent. Three calibration files will not produce a statistic, and they do not need to; the point is that "roughly the same as last time" is a claim that can be checked rather than felt.
Fifteen minutes, four times a year. It is the least glamorous hour in this list and it is the reason your third year looks like your first from the buyer's side.
Where the rubric does not help
Not everything is axis-shaped, and forcing it is how a structure starts flattening the work.
Comparison jobs need a different apparatus: the same axes applied across submissions produce a table, and the table needs a paragraph on top of it that the rubric cannot generate. Recorded formats resist heavy structure because a spoken assessment that follows five numbered headings sounds like someone reading a form. Use the axes as an internal running order there and let the delivery be looser.
And a rubric never rescues a thin brief. If you do not know what the buyer wanted assessed, structured wrongness arrives faster than unstructured wrongness, which is not an improvement.
Building yours this week
Do it in one sitting, from your own past work rather than from first principles.
Take your last five assessments and highlight every distinct thing you commented on. Group the highlights. The groups that appear in four or five of them are your axes; the ones that appeared once were responses to a specific submission and belong in the free-text section, not the structure.
Write the five anchors. Write one example justification per axis, taken from real past work, so that you have a shape to match rather than a rule to remember. Then run the whole thing once, end to end, on practice material and time it - the practice run is what converts the rubric from a document into a number of minutes, and the number of minutes is what your price is built from.
Expect the first rubric-scored job to be slower than an unstructured one. The saving arrives around the fifth, and it compounds from there.
Two things are worth reading from outside this site while the structure is still forming. Automated scoring already produces axis-based output instantly and free, so any axis of yours that a model does equally well is not what you are being paid for; the account of what automated scoring reliably does is the clearest guide to which of your axes carry the fee. And buyers' expectations about what a structured assessment should contain are set by what they have received before, which Rate Penis's material on reviews describes from the commissioning side.
The rubric earns its keep in a way that is easy to miss because it shows up as absence: no blank page, no re-litigated scope, no fortnight where your standard quietly slipped. It is the closest thing this trade has to fixed capital, and it costs an afternoon.