Starting out
Two buyers comparing notes should recognise the same person
A wildly variable standard is worse for reputation than a consistently modest one.
Guides on Starting out: A rubric is a production tool before it is a quality claim, The brief is rarely the request, Getting the first booking, in order
A buyer is not comparing your work to perfect. They are comparing it to what they were told to expect, and to what someone else told them they received, which means variance costs you more than an ordinary standard does.
The earner who is excellent twice and thin once has a worse reputation than the earner who is solidly good three times, because the thin one is the review that gets written.
Why variance is punished harder than level
Two mechanisms, both boring and both reliable.
The first is that reviews are written disproportionately by people whose expectations were violated, in either direction, and downward violations produce more text than upward ones. A consistently good deliverable generates a mild positive review. A deliverable that is worse than the buyer's last one generates a specific negative one, and specific negatives are the reviews future buyers actually read.
The second is that inconsistency is unforecastable, and buyers are managing risk more than they are seeking quality. A known modest standard can be planned around. A standard that might be excellent and might not is, from a buyer's position, indistinguishable from a coin flip with your fee attached.
This is the same logic behind everything in what buyers are really asking for: the purchase is mostly a bet on predictability.
Where variance actually comes from
Almost never from ability. Four causes account for most of it.
| Cause | What it looks like | Fix |
|---|---|---|
| Time of day | The 11pm job is 40% shorter | Batch work into fixed sessions |
| Fee level | Cheap jobs get less care | Price so every tier is worth doing properly |
| Interest | A dull submission gets a dull assessment | Fixed skeleton forces the sections anyway |
| Fatigue | Job five of the day is the weak one | Cap jobs per session, not per week |
The time-of-day one dominates and is the easiest to fix, because it is a scheduling decision rather than a discipline problem. Working in defined sessions rather than whenever a booking lands removes most of it at a stroke - batching work into sessions covers how that actually gets arranged around a normal week.
The fee one is worth being honest about. If your lowest tier is priced such that doing it properly is irritating, you will not do it properly, and the review will come from that tier. That is a pricing error being paid for by your reputation.
The three things that hold a standard
A fixed skeleton. The same sections in the same order every time, so the floor of the deliverable is set by structure rather than by how you feel. A tired writer working to a five-part skeleton still produces five parts. A tired writer working from nothing produces two paragraphs and a sign-off, and the shape of that skeleton is in the structure of a written assessment.
Templates that carry the frame and not the content. Headings, boilerplate scope lines, the handover message, the score table. Never observations, never phrases about the material. The boundary is the whole subject of templates that speed you up without flattening the work, and crossing it turns consistency into the generic writing buyers punish hardest.
A check before sending. Two minutes, against a fixed list, every job without exception. The exception is where the bad one comes from.
The pre-send check
Five items, in this order, and it works because it is short enough to actually do.
Does it have every section the skeleton has. Is there a specific referent in each section, or is one of them general. Does the summary say the same thing as the body. Is the buyer's name, the brief's own vocabulary, and anything they specifically asked about, present. Is there anything in it that could not be defended if quoted back at you.
That last one catches the sentences written at speed that read differently in daylight. The rest of the routine, including where it sits relative to delivery, is in the final check before you send.
Consistency across formats and platforms
If you offer more than one format, the recognisable thing has to be the judgement rather than the layout, because a recording cannot have headings.
The practical version is a fixed running order that you follow in speech as well as in writing: overall read first, observations in a set sequence, what would change it, close. A buyer who has bought both formats should hear the same person, and the same order is most of what produces that.
Across platforms it is a stronger requirement, because listings get compared. A rate card and a described format that differ between two sites read as either carelessness or as two different offers, and cross-platform consistency covers what has to match and what can differ.
What consistency is not
It is not saying the same things. An assessment that reaches the same conclusion regardless of the material is consistent in the useless sense, and buyers spot it immediately - which is what an automated score reliably delivers at zero cost, and why matching a machine's uniformity is the wrong target.
The thing to hold steady is the process: the same axes, the same depth of attention, the same structure, applied to material that should produce genuinely different outputs. Rater consistency is measurable rather than a feeling: the reliability guideline by Koo and Li (2016) treats an intraclass correlation below 0.5 as poor and above 0.90 as excellent, and a process you repeat is what moves you up that scale. Where the axes involve figures, holding to a published convention rather than your own drifting one is the cheapest form of consistency available - the site that documents method is the reference, and using it means your third job scores the same way as your thirtieth.
Buyers who have already run their material through the automated tools that produce a quick score will notice if your read swings wildly from job to job in a way the material does not justify, because they have a stable comparison point in hand.
The platform-side version - what a judge's listed format commits you to, and what happens when deliverables diverge from it - is on Rate Cock's judges page, and the commitment there is stricter than most new earners assume.