Which listings are dragging your Etsy rating down
· 3 min read
The listing hurting your shop rating most is rarely the one with the lowest stars. Review volume changes the order, which means the right place to start is the listing creating the most drag.
When an Etsy shop rating drops, the natural response is to open the reviews, find the listing with the lowest stars, and start there. That feels objective. It is also often the wrong priority.
A listing average does not enter the shop average once. Its reviews enter one by one. A product with a slightly weak rating across dozens of reviews can pull the combined number down much more than a product with one or two terrible reviews. Raw stars tell you severity. They do not tell you impact.
Why the worst stars can be the wrong list
Take a simplified rating window. The rest of a shop has 100 reviews averaging 4.8 stars. One listing has 2 reviews averaging 3.0. Another has 50 reviews averaging 4.2.
Together, that is 696 stars across 152 reviews, or an average of about 4.58. Remove the 3.0-star listing and the average rises to 4.60, a gain of about 0.02. Remove the 4.2-star listing and it rises to about 4.76, a gain of about 0.19.
The 3.0 listing looks worse in a sorted table. The 4.2 listing creates almost nine times as much drag because far more reviews carry that lower-than-shop rating into the average. If the goal is to understand why the combined rating moved, the second listing belongs at the top of the work list.
This is why sales volume is only indirect context. More orders may produce more reviews, but orders without reviews do not enter the rating calculation. The useful weight is review count.
What a drag score measures
A drag score asks a practical counterfactual: what would the average be if this listing’s reviews were removed from the window?
Everlyst calculates the current average, calculates it again without the listing, and records the difference. A listing with a lower rating and meaningful review volume produces negative drag. A strong listing can produce positive lift. Listings without enough reviews stay unscored rather than receiving a confident number from thin evidence.
Sorting by drag changes the priority list from “lowest rating” to “largest effect on the shop.” Sometimes the same listing tops both lists. Often it does not. The distinction matters because the response is different too. Two bad reviews might need a careful read. Fifty consistently softer reviews point to a repeatable product or fulfillment problem.
Orders add the missing context
The rating identifies the listing, but the order can identify the cause. That requires joining each review to the transaction behind it.
Once that join exists, variation-level patterns become visible. A shirt might average 4.7 overall while one size receives repeated complaints about fit. A print might perform well in every option except one frame finish. That is not necessarily a product-wide failure. It is a specific option that can be corrected without rebuilding a listing that otherwise works.
The same transaction context makes shipping comparisons possible. Group reviewed orders by the number of days between purchase and shipment, then compare their average ratings. If slower orders consistently rate worse, the problem is not buried in product copy. It is in processing time or fulfillment consistency.
Etsy’s review screen presents the feedback, but it does not assemble these listing, variation, order, and shipping comparisons for you.
Edits need the same timeline
Reviews arrive after the order, sometimes weeks later. That delay makes memory unreliable. A rating dip in July might reflect a description change made in June, a new supplier introduced in May, or orders placed before either change.
The honest way to judge an edit is to place listing changes on the same timeline as the rating trend. Find the edit marker, allow for purchase and review lag, then compare the reviews before and after it. One new five-star review proves very little. A sustained shift across enough reviews is evidence that the change worked.
This also prevents false credit. If ratings began improving before the edit, the chart makes that visible. The goal is not to defend the decision. It is to see whether the data changed after it.
Put impact ahead of severity
Review Analytics in Everlyst brings this method into one view: reviews joined to transactions, listing drag scores, variation ratings, shipping-speed comparisons, and edit markers on the rating timeline.
Starter includes the review feed with filters and acknowledgements plus 90-day per-listing analytics. Growth and Pro add full history, the rating trend with edit markers, drag scores, variation ratings, and insights. The Review Analytics changelog entry has the release details.