Mosana
← Back to blog
analisis-de-resenassesgo-de-escalanlp-hotelesreputacion-online

High Score With a Complaint in the Text, Low Score With None: Why Review Sentiment Doesn't Predict the Score

Why a review's text and its score are different signals — and what that means for any automated analysis.

Share this article

Case A: a top score despite a complaint in the text

It's common to find a review scored 9 or 10 out of 10 whose text literally says check-in took 40 minutes or the air conditioning was noisy at night — and still, "everything was great." This happens because most guests score the overall experience, not each detail separately, and satisfaction scales carry a well-documented generosity bias: unless something completely ruins the stay, the score skews toward the top. The complaint in the text is real and specific; the score is an average weighted by factors that same text never mentions (price paid, prior expectations, comparison to the last trip).

Case B: a score that isn't perfect, with no complaint in the text

The opposite case is just as common and harder to detect: a guest leaves a 7 out of 10 with a text that only says "all good, great location," with no problem mentioned. There's no textual clue explaining why it wasn't a 10. This is silent dissatisfaction: something cost points — a room that didn't match the photos, impersonal treatment, a price that didn't feel justified — but the guest never put it into words, either because they didn't want to sound "harsh" in a public comment or because the reason wasn't even fully conscious.

The consequence: a very low noise ceiling

De Langhe, Fernbach and Lichtenstein (Journal of Consumer Research, 2016) analyzed hundreds of thousands of online reviews and found that a product's average rating barely predicts its objective quality as measured by independent consumer reports — the stars say less than they appear to. If the score is already a noisy summary of the real experience, trying to predict it from the text inherits that same ceiling: it's not that a better language model is missing, it's that the signal you're trying to predict doesn't carry all the information the text has, and vice versa.

What this means for any review-analysis tool

It's the same temptation any AI-driven review-analysis system faces, ours included: presenting a score prediction from the text as if it were a verified fact, instead of an inference with a real margin of error. It's easy to give in to that temptation because a clean number sells better than a nuanced answer — but if the system bases its recommendations on a prediction that can't hold up, the problem it detects, and the one it misses, can be wrong from the root.

How to spot this pattern in your own property's reviews

You don't need an algorithm to start seeing it: just manually pull a handful of reviews scored 8 or higher and read only the text, without looking at the score first. It's common to find at least one in ten with a specific complaint (noise, cleanliness, wait times) that, read in isolation, you'd expect to come with a lower score. The reverse exercise — reading complaint-free text from reviews scored 6 or 7 — is usually even more revealing, because there's no text left to investigate: only the guest themself could say what was missing.

If you manage several properties or a high volume of reviews, that manual sampling quickly stops being practical — and that's exactly the point at which it makes sense to automate the detection of discrepancies, not the prediction of the score itself.

Share this article

Frequently asked questions

So is there no point analyzing review text at all?

It's very useful — for what the text can actually answer: which topics repeat, in which category, how often, and with what tone. What it can't reliably do is substitute for or predict the numeric score. They're two different questions with two different answers.

Why don't guests just score more precisely?

Because most guests aren't scoring each detail separately — they're scoring a global impression shaped by the price paid, prior expectations, and comparisons with other stays, factors that almost never show up in the review text. It isn't a guest failing to be precise — it's simply what a general satisfaction scale does (and doesn't) measure.

How we designed Mosana knowing this

Mosana treats score and text as two complementary signals, not as one predicting the other. The text is analyzed to extract topics and sentiment specific to each category (cleanliness, service, location, noise...), and the score is used as an independent aggregate metric, never inferred from the text. When the two signals disagree — a high score with a detected complaint, or a low score with neutral text — the system flags it explicitly as a case worth attention, instead of forcing a consistency the data doesn't have.

We don't promise to predict your score from a review's text — no tool can do that reliably. What we do is make sure neither signal gets lost along the way.