Starting out

Volume matters more than the average

Nobody reads twelve reviews; they read that there are twelve.

By Updated 5 min readStarting out

Guides on Starting out: A rubric is a production tool before it is a quality claim, The brief is rarely the request, Getting the first booking, in order

The number a buyer acts on is the count, not the average. Averages cluster so tightly across every marketplace that they carry almost no information, and the count is the only part of the pair that varies enough to be worth reading.

This has an uncomfortable implication for anyone new, and a useful one: your problem is not that your score is unproven, it is that there is nothing next to it.

Why averages stop discriminating

Ratings in service marketplaces compress upward. Buyers who had a normal experience give the top score, buyers who had a bad one frequently give nothing at all rather than a low one, and the result is that almost everyone visible sits in a narrow band near the top.

The effect is documented: Filippas, Horton and Golden (2022, Marketing Science) show raters feel pressure to leave "above average" ratings because they do not want to harm the seller, which pushes the average up over time. The consequence is worth stating precisely. When every listing on the page is inside a few tenths of the same figure, the average has stopped being a way to choose, and the buyer's eye moves to the thing beside it.

That thing is the count, and the count varies by two orders of magnitude across the same page.

What a count is actually evidence of

Not quality. A count is evidence that other people took the risk, which is the question the buyer is really asking, since they cannot assess the work in advance and know it.

It is also evidence of three things they cannot see directly. That you exist and are not a dormant listing. That you have delivered repeatedly, which is a claim about reliability rather than skill. And that whatever went wrong for other buyers was survivable enough that nobody bothered to warn anyone.

An illustration of how the same average reads at different counts, not a claim about any specific platform:

Average Count What a buyer infers
5.0 1 Unknown. One person, possibly a friend.
4.6 30 Established. Small problems exist and are minor.
5.0 4 Promising, still a gamble.
4.9 120 Default choice on the page.

The row that surprises people is the third against the second. A perfect four is worse positioned than a 4.6 with thirty behind it, and every earner protecting an unblemished average from a slightly risky job has the trade the wrong way round.

Why the early ones weigh so much

Each new review moves the average by less than the one before, arithmetically, so the first few are the only ones with real leverage on the number. That much is obvious.

The less obvious part is that they also set the text. A listing's first three reviews are the ones a buyer reads in full, because they are frequently all there is, and they establish what the work is in the reader's mind before any later review gets a chance to. This is the practical reason the first ten are worth engineering rather than hoping for. Engineering means earning and asking, never buying: the FTC rule finalised in August 2024 bans fake reviews and paying for reviews conditioned on a particular sentiment.

Early reviews compound in a second way through ranking. Marketplace sorting generally rewards completed work and recent activity, so the reviews that arrive early do not just persuade, they increase how many people see the listing at all - the mechanics of which are covered in the piece on ranking inside a marketplace.

A low count is a harder obstacle than a mediocre score

If you have four reviews averaging 4.5, your problem is four, not 4.5.

Everything you would do to fix the average - delivering only easy jobs, avoiding difficult buyers, over-servicing to guarantee a five - slows the count, and the count is what is blocking you. It is one of the few situations in this trade where the aggressive move and the correct move are the same, and the case for volume over price early on rests on exactly this.

The exception is a genuine outlier. A one-star with a specific complaint sitting in a set of five reviews is doing real damage, because it is the specific one and specific reviews are the ones read. That is not an averages problem either, and it has its own answer.

What buyers infer beyond the numbers

Three inferences run alongside the count and each is cheap to influence.

Recency. A last review from eight months ago reads as a person who stopped, and buyers check this more carefully than they read the profile text. What a maintained profile signals is mostly this.

Distribution of length. Several substantial reviews suggest jobs substantial enough to write about. A column of four-word entries suggests small transactions, which prices you accordingly.

Whether anything is described. A count with no describable content behind it converts worse than a smaller count with detail, which is the argument in the piece on what makes a review persuasive.

The comparison being made is not the one you think

Earners read their own listing against the listing above them. Buyers read a page of listings against the option of not booking anyone, and against the free automated alternative they have usually already tried.

That is worth holding onto, because it changes what the count needs to beat. It does not need to beat the person with a hundred and twenty reviews; it needs to be enough that a person willing to pay does not retreat to a free score instead. The people who work with rating data are notably careful about how much any aggregate figure supports, and the notes on reading aggregates apply to your star average as much as to any measurement. How buyers choose between the available options in the first place, and what actually makes the shortlist, is set out from their side. And what a score communicates as a bare number, stripped of who produced it, is a subject of its own with less in it than people assume.

On platforms where the sort order and the displayed aggregate are documented rather than guessed at, read the documentation instead of theorising; Rate Cock's judge pages set out how listings are surfaced there.

Read next

Full archive