Skip to content

How we rank listings

The browse board is ordered by a model, and this page is that model — the same constants the code actually runs on, not a description of them.

Scoring version 2. Changing any figure below means publishing a new version, not quietly editing this page.

The short version

Two things decide where a listing sits: whether the person behind it delivers, and how much they've told you. The first is worth far more than the second, and the second can never make up for a failure on the first.

On day one nobody has a delivery record, so the board starts out ordered by what can be verified — identity and an aged handle — and shifts to actual conduct as soon as there is any.

1. The starting point, before any history

Every lister starts at a base rate set by what they've verified. Each rung has to cost something that the same person can't simply regenerate — that is the whole test for being on this list.

TierBase rateWhat it means
T00.30Signed up. Nothing verified yet.
Free, and re-creatable in a minute — which is exactly why it starts here.
T10.50Payout identity verified.
Stripe has checked a real name, date of birth and bank account. Free to an honest lister — you need it to be paid at all — and genuinely expensive to fake at scale.
T20.58A machine-verified handle we've known for 60 days.
Counted from when we first verified it, not from when the account was created — so a fresh start resets it however old the underlying account is. Time is the one thing on this list nobody can buy retroactively.

The prior is worth 3 observations. In practice that means it's half the picture by the third completed order and about a quarter by the ninth — so verification gets a lister started and then stops mattering. Nobody can buy a permanently high position with it.

T0's 0.30 is not a number we chose. It's derived, so that abandoning a damaged account and signing up again is never worth doing — a lister carrying the worst record we think should still be recoverable scores exactly what starting over would score, and everything better than that is strictly better than walking away.

The ladder stops at T2 on purpose. There was a fourth rung for choosing print & post, justified as the lister staking their own money — but that fee is charged to the sponsor and never comes out of the lister's payout, so it cost them nothing and the rung was really just a button in the listing form. A rung nobody can honestly earn is worse than no rung, so it's gone until a real one exists. The fee still counts as evidence where it genuinely is unrecoverable: per order, below.

2. What counts as delivery evidence

There are no star ratings here and nobody is asked for an opinion. The only thing recorded is what we observed: the sticker photo arrived inside the window, or it didn't.

The denominator is every brief accepted, including the quiet ones nobody complained about. That matters more than it sounds: on platforms that rely on buyers leaving feedback, most transactions produce none, and the displayed score ends up describing the loud minority. Silence counts here.

Some orders are never the lister's to answer for

A sponsor who never sends artwork, a logo our own review rejects, a card flagged by its issuer, a listing we take down, a bid our collusion checks flag — none of these touch a lister's record. Each one is recorded with the reason it was excluded, and a lister can ask to see it.

Evidence is priced by what it would cost to fake

An order is worth a full observation once the money that cannot come back to the person faking it — our cut, plus the £12 print-and-post fee — reaches £3. Below £1 it is worth nothing at all, and it scales in between. A wash trade between two accounts of the same person costs them almost nothing and buys almost nothing.

Old evidence fades, but only once there's enough of it

Once a record holds the equivalent of 10 full observations, evidence halves in weight every 365 days. Below that it doesn't decay at all — fading the record of someone with four lifetime orders would erase it rather than age it. Note that this counts the same weighted observations described just above, not raw order count: a dozen orders too cheap to be worth much evidence still won't start decaying.

A clean record that has been idle more than 180 days drifts back toward the starting base rate: a dormant lister is less certain, not worse. A record carrying a failure never does that. Waiting is not a way to clear it.

3. The ten disclosure points

The second half of the score is how much a buyer can actually check. Every point is something the lister did:

  1. Photo uploaded and passed review
  2. Email confirmed
  3. All seven placement answers given
  4. Context photo uploaded
  5. Sticker measured with the ruler
  6. Placement answers confirmed in the last 90 days
  7. At least one handle machine-verified
  8. Payout identity verified
  9. Committed to posting the proof photo publicly
  10. Placement claims corroborated against the photos

No point rewards a flattering answer. Saying your laptop never leaves the house scores exactly what “out seven days a week” scores. Saying it's already covered in other brands' stickers scores exactly what “nothing on it” scores. Declining to name your area costs nothing at all.

That is deliberate and it is the most important rule here. The moment a rank point depends on what an answer says, the form stops measuring honesty and starts pricing it — and the answers underneath it become worthless to you.

Placement answers go stale after 90 days and have to be re-confirmed. Re-confirming the same answers earns the point; so does changing them. The point is the act of saying “still true”, not what's said.

We spot-check 10% of listings at random against their own photos. That figure is published because we run it — a stated audit rate you don't actually perform is just a false statement about your own process.

4. How the two combine

Disclosure multiplies the delivery score rather than being added to it. A fully disclosed listing is worth up to 2.5× a bare one — but because it multiplies, it can never rescue a delivery record that has gone to zero. Context cannot buy forgiveness for a failure.

Past 2 weighted failures, a listing sorts below every clean lister regardless of anything else. A single failure costs position, not that. One exception: a dispute decided against the lister does it on its own, at any weight — that is the one case where somebody has already ruled on the facts, and waiting for a second would be ignoring a finding rather than being careful.

Listings are then grouped into bands 0.05 wide and shuffled within each band, reshuffling every hour. That is on purpose: on a small board, exact ordering would freeze the same few listings at the top forever, and a new lister would never get seen. The shuffle hands out impressions by construction.

5. Why there's no score on the page

You will never see a trust score, a star rating, a percentage or a percentile here — not because the number doesn't exist, but because showing it would be the misleading part.

A single blended figure lets a strong half hide a weak one: a lister with excellent disclosure and a missed delivery would show up as a respectable number, and the one fact you actually needed would be the one the number buried. So the listing page shows the counts separately instead — delivered, missed or late, disputes upheld, out of every brief accepted — and lets you draw your own conclusion.

The blend is used for one thing only: deciding what order to show you things in. That is the finding of the only randomised study on the question, and it is also what the transparency rules on ranking are getting at.

6. If you think this got you wrong

Every verdict is computed fresh from your orders rather than stored, and every one carries the machine-readable reason that produced it. Nothing here is a number somebody typed in. Email [email protected] and we'll show you the exact list of orders behind your record and how each one was classified.