In the working paper Reputation Inflation, Apostolos Filippas (Fordham University), John J. Horton (MIT Sloan) and Joseph M. Golden examine roughly six years of transaction and feedback records from a large online labour marketplace, covering 2007 to 2016. The question is simple: what happens to a five-star rating system when it is left running for a long time.
The distribution moved a great deal. The share of workers receiving a perfect five-star rating rose from 33 per cent to 85 per cent across the period. The median feedback score rose from approximately 3.74 stars in early 2007 to 4.85 stars by May 2016. By the end of the window, more than 80 per cent of all evaluations fell into the single bin between 4.75 and 5.00 stars.
The mechanism the authors propose
Filippas, Horton and Golden attribute the drift to what they call reflected costs. An employer leaving a low rating knows the rating will damage the worker's future prospects on the platform, and that knowledge makes the low rating costly to give. The more heavily the marketplace relies on feedback scores in matching and hiring, the higher that cost becomes, and the more reluctant raters are to impose it.
The loop closes on itself. What counts as a bad score is defined by the surrounding distribution, so as scores compress upward, a merely good rating becomes a damaging one, and the reluctance intensifies. The authors describe the threshold as endogenous to the distribution it is measured against.
The private feedback comparison
The strongest evidence in the paper comes from a period during which the platform collected private feedback alongside the public score. The private ratings were shown neither to the worker nor to other employers, which removed the reflected cost.
Over the same window, average private scores were falling while average public scores for the same transactions were rising. The two series, describing identical transactions, moved in opposite directions. When the platform later made the private feedback consequential by releasing it publicly, the private scores began to inflate immediately.
Using sentiment analysis of written feedback as an independent check, the authors estimate that more than half of the six-year increase in scores is attributable to inflation rather than to improvement in the underlying work. They describe this as a conservative figure, since written feedback is subject to the same pressure and can inflate alongside the numeric score.
How far it generalises, and how far it does not
The authors obtained longitudinal data from four further online marketplaces, spanning home-sharing, asset rental and services, and report the same upward trend in each. They also mark the boundary: the mechanism should be weakest where reflected costs are low, such as product reviews, where the rater has no relationship with anyone whose prospects the rating affects.
Several caveats sit alongside the result. The primary dataset comes from one marketplace, identified only generically, so platform-specific policy changes cannot be fully separated from the trend. The paper circulated as a working paper, and figures have varied across drafts. The decomposition of the increase depends on a sentiment classifier applied to written comments, which carries its own error. And the counterfactual — what the quality of the underlying work actually did over nine years — is not directly observed by any series in the paper.
The practical consequence the authors draw is narrow and worth stating in their terms. As scores compress into a single bin, they carry progressively less information for the purpose they were built to serve, and matching in the marketplace comes to rely on signals other than the visible rating.