WouldSwipe Data Study

How Many Votes Does a Dating Photo Need?

Five votes give an early signal. Ten are much more useful. Twenty make the average notably steadier.

Published July 28, 2026 · 8 min read
Fictional dating portrait, increasing groups of rating tokens, and a score line becoming stable

The short answer

Use five votes to spot a weak photo, ten to make a practical decision, and twenty when two photos are close. In this sample, the average difference from the eventual score fell from 0.46 after five votes to 0.19 after twenty.

What changed after 5, 10, and 20 votes

For every eligible photo, we ordered ratings by time, calculated the running score after a threshold, and compared it with that photo's final score in the July 28 data snapshot. The combined score is the equal-weight mean of attractiveness, trustworthiness, and fun.

After
5
votes
0.46
mean absolute change
58% within 0.5 points
0.74 rank correlation
After
10
votes
0.32
mean absolute change
81% within 0.5 points
0.84 rank correlation
After
20
votes
0.19
mean absolute change
92% within 0.5 points
0.90 rank correlation

Samples differ by threshold: 31 photos had later votes after vote 5, 21 had later votes after vote 10, and 13 had later votes after vote 20. Results are descriptive and should not be read as universal guarantees.

Why both "change" and "rank" matter

A score can be useful in two ways. The first is absolute stability: will 6.2 still look like roughly 6.2 after more people vote? The second is rank stability: will photo A remain ahead of photo B?

After five votes, the rank correlation with the final ordering was already 0.74, but 42% of photos still moved by more than half a point. This means five votes can often identify broad differences while remaining noisy for close calls.

After ten, rank correlation rose to 0.84 and four out of five photos sat within half a point of their eventual score. Twenty improved both measures again, but the gain from 10 to 20 was smaller than the gain from 5 to 10.

Attractiveness was the noisiest score

VotesAttractive changeTrustworthy changeFun change
50.750.560.65
100.420.360.45
200.250.160.26

At every threshold, attractiveness moved more than trustworthiness. That is consistent with the idea that attraction includes substantial personal taste, while some first-impression traits may draw more shared judgments.

What published research says about rating agreement

Scientific studies find both consensus and individuality. Static and video ratings of the same faces have shown high agreement when averaged across groups, including a reported Cronbach's alpha of 0.92 in one study. See the PLOS ONE comparison of video and static ratings.

That does not mean every person agrees. Germine and colleagues first gathered face preferences from roughly 35,000 participants, then studied twins. They concluded that individual face preferences are reliably measurable and shaped mainly by experiences unique to each individual rather than shared genes or family environment. See the 2015 Current Biology study.

Measurement design matters too. A 2024 psychometric study compared binary, 5-point, 7-point, 10-point, and 0-100 attractiveness scales. Its authors caution that common reliability measures can rise mechanically as more raters are added and that scales with very few response options can lose useful information. See the i-Perception study.

Together, those findings explain the WouldSwipe curve: averaging reduces noise and reveals shared signal, but no vote count removes personal taste.

A decision rule that matches the data

Stop at 5 when:

  • one photo is clearly behind by more than a point;
  • you only need an early warning about an obvious problem;
  • you plan to replace the image regardless.

Wait for 10 when:

  • you are choosing a lead photo for your profile;
  • the alternatives differ by roughly half a point or more;
  • you want a reasonable balance of speed and stability.

Wait for 20 when:

  • two photos are separated by only a few tenths;
  • the choice matters enough to justify more certainty;
  • you want to compare individual trait scores, especially attractiveness or fun.

Method and limitations

  • Data window: eligible votes dated January 30 through July 22, 2026; exported July 28.
  • Clean dataset: 985 ratings from 106 voters on 162 photos owned by 76 users.
  • Exclusions: known bot, example, and non-login seed accounts were removed as voters and owners.
  • Chronology: ratings were ordered by timestamp and vote ID. The running average after each threshold was compared with the final available average.
  • Eligible photos: only photos with at least one later rating after the threshold entered that threshold's comparison.
  • Important caveat: the final score includes the earlier votes, so the running and final estimates are not independent. The metric describes convergence, not prediction on an untouched test set.
  • Changing sample: photos that reached 20 votes may differ from photos that stopped earlier. Improvements across thresholds are therefore descriptive, not a controlled causal estimate.
  • Privacy: no submitted photo or user-level record is published. The hero contains a fictional, generated person.

Questions about vote counts

Are five dating photo ratings enough?

They are enough for an early signal, especially when the choice is obvious. They are weak evidence for a close decision: the combined average still moved by 0.46 points on average in this sample.

Why did my score drop after another vote?

A single new rating has a large effect when the denominator is small. One extra score represents one-sixth of a six-vote average but only one-twenty-first of a 21-vote average.

Should I ignore a photo that scores well with some people and poorly with others?

Not necessarily. A polarizing photo may work for the audience you want. The average answers "How broad is the appeal?" It does not fully answer "Will the right person like it?"

Get a decision you can use

Compare your real photos on attractiveness, trustworthiness, and fun.

Test My Photos

Read next