Pay
A template is a scaffold, not a script
The parts of an assessment that repeat should be automatic so that the parts that do not can have your attention.
Guides on Pay: Pricing, from first principles to the annual review, The ceiling is hours, and it arrives sooner than people expect, What rating work actually pays
Template the container, never the contents: headings, axis definitions, caveats and delivery messages should cost you no time, while every observation is written for the job. Templates are the largest lever on your effective hourly rate, and also the fastest way to make your work worthless.
The difference is entirely a question of what you template. Anything that will be identical across every job should cost you no time at all; anything that varies is the reason someone paid a person.
What is safe to template
Some of a deliverable is genuinely the same every time, and reproducing it by hand is a small tax you pay on every job for no return.
The heading structure. The definitions of your axes, stated once so the buyer knows what each score means. The standard caveats - what you did not assess, what the material did not let you judge, what a score on your scale does and does not imply. The delivery message, the file naming convention, the closing note.
Those are the container. A buyer reading two of your reports back to back would expect them to match on all of it, and would be mildly alarmed if they did not, because inconsistent structure reads as improvisation.
Saving here is real and it compounds. Ten minutes per job of retyped boilerplate, across four jobs a week, is most of a working day a month recovered - which is a larger effect on your income than a price rise you could plausibly defend, and it is available immediately. That interaction between time per job and what an hour is actually worth to you is the arithmetic in the piece on what rating work pays.
The rubric underneath the axes is its own subject and is treated separately in the post on building a rubric you can reuse. What matters here is only that a rubric and a template are different objects: the rubric decides what you look at, the template decides how the result is laid out.
What must never be templated
The observation. The justification for each score. The comparison to whatever reference point you use. Anything the buyer will read as "this person looked at mine specifically".
The tell that a line has crossed over is simple: if you could paste it into any other job without changing a word, and it is not a caveat or a definition, it is padding. It is occupying space where the specific thing should have been, and buyers detect this reliably even when they cannot articulate what is wrong.
This is why structured scoring outperforms prose commercially without becoming generic - the structure carries the professionalism and the reasoning carries the specificity, an argument worked through in the comparison of written and scored deliverables.
The failure mode, in detail
Generic output does not usually arrive by a deliberate decision. It arrives by accretion, and it goes like this.
You write a good sentence for one job. It fits the next job too, so you keep it. Three months later that sentence is in your template, and eight sentences like it are in your template, and each of them was once a genuine observation.
Your deliverable is now sixty percent constant by volume and you cannot see it, because you never wrote a generic sentence - you wrote twenty specific ones and let them harden.
The consequences show up in order. First the reviews get shorter and less specific, because there is less specific work to describe. Then repeat bookings decline, because the second purchase is visibly the first one. Then the price becomes indefensible, because what you are selling is now something a buyer can approximate for nothing from the automated tools that already produce structured generic output instantly.
That last point is the whole commercial stake. Automation owns generic assessment completely and permanently. The only defensible position left is the part of the work that could not have been produced without looking at this particular submission, so a template that erodes that part is eroding the entire basis of the price.
Keeping it honest
Three practices are enough, and none of them takes long.
Read your last three deliverables side by side, monthly. Highlight every sentence that appears in more than one. Anything highlighted that is not a definition or a caveat gets deleted from the template. This takes fifteen minutes and it is the only reliable detection method, because you cannot notice drift from inside a single job.
Set a floor on specific content. A minimum number of observations per axis that must be about this submission. Not a word count - a count of things that could not be said about anything else.
Leave the variable parts visibly empty. A template with [what stood out] in it prompts you; a template with a plausible generic sentence there invites you to leave it. This is the single highest-value habit in the list, because it converts a passive default into an active decision every time.
Consistency is a feature, up to a point
None of this argues for making every deliverable different. Consistency is genuinely valuable - it is what makes a repeat buyer's second purchase legible, and it is why standardised structure is worth having at all. The professions that assess things for money almost all work to fixed forms, and the measurement side's account of why a stated method matters is the clearest statement of the case.
The distinction is that a form is a set of questions, and what you are selling is the answers. A radiologist's report has the same headings every time and nobody would call it generic, because the content under each heading is about one person.
The same logic applies to how buyers read your work. They are comparing risk rather than prose, and a consistent structure lowers perceived risk, which is why the format itself converts - the commissioning-side view of what makes a report feel trustworthy is on Rate Penis's reviews pages, and a stated method is most of it.
One caution on automating further than templates. Drafting tools will happily produce the observational content too, and the output is fluent, plausible and not about the submission in any way that survives a careful reading. The failure is structural rather than a bug of this year's models: Kalai and colleagues (2025) describe language models that "guess when uncertain, producing plausible yet incorrect statements instead of admitting uncertainty". If you use them at all, use them on the container, where being generic is the point.
For how structured deliverables are actually presented to buyers inside a marketplace - what the platform formats for you and what you have to bring - Rate Cock's judge listings show the current shape. The useful exercise is reading several from the same person and seeing how quickly you can spot which sentences they wrote once.