On all six frameworks, the friends and colleagues someone invites to review their personality result rate it more accurate than the person does: 4.60 against 4.41 across 49,161 ratings, with the gap widest on the Enneagram at 4.58 against 4.30 and narrowest on Strengths at 4.68 against 4.63.
Who this describes
People who rated their own Crystal result, and the friends and colleagues they invited to rate it for them
Sample
49,161 ratings, 23,347 from invited reviewers and 25,814 from the people themselves
Data as of
October 2, 2026
| Framework | Reviewers minus self | Reviewers, then self | Ratings |
|---|---|---|---|
| Enneagram | +0.28 | 4.58 and 4.30 | 8,338 |
| Big Five | +0.22 | 4.54 and 4.32 | 8,569 |
| DISC | +0.18 | 4.54 and 4.36 | 10,008 |
| MBTI | +0.18 | 4.54 and 4.36 | 7,961 |
| Life Values | +0.11 | 4.72 and 4.61 | 7,118 |
| Strengths | +0.05 | 4.68 and 4.63 | 7,167 |
Six frameworks, six times the same direction. That unanimity is the result, more than any single gap in the table: whatever is happening does not depend on which framework someone took or what kind of thing it claims to describe.
The size of the gap is ordered in a way worth noticing. It is widest on the Enneagram, the framework that asks what drives you, and narrowest on Strengths and Life Values, which describe what you are good at and what you care about. The frameworks where other people agree most readily are the ones about visible things, and the framework where a person is most likely to dispute their own result is the one about inner motivation.
There is a plainer explanation and it has to be said first: these reviewers were invited by the person being rated. A friend asked to say whether a description fits someone has reason to be kind, and a stranger would very likely rate lower. Nothing here establishes that other people are better judges of you, and the data cannot separate generosity from insight. Doing that would need reviewers nobody chose, which is not what this is.
What it does support is narrower and more useful than the headline version. A result that reads as slightly wrong to the person it describes does not usually read as wrong to the people around them, and that pattern holds on every framework measured. The least flattering reading of your own profile is, reliably, your own.
The two sides on their own
The same ratings, uncombined, with the count behind each figure. The two columns above are these numbers subtracted; they are not two readings of one group, and the ratings in each row come from different people.
| Framework and rater | Mean rating, out of 5 | Ratings |
|---|---|---|
| Enneagram, invited reviewers | 4.58 | 3,892 |
| Enneagram, the person themselves | 4.30 | 4,446 |
| Big Five, invited reviewers | 4.54 | 3,935 |
| Big Five, the person themselves | 4.32 | 4,634 |
| DISC, invited reviewers | 4.54 | 4,003 |
| DISC, the person themselves | 4.36 | 6,005 |
| MBTI, invited reviewers | 4.54 | 3,871 |
| MBTI, the person themselves | 4.36 | 4,090 |
| Life Values, invited reviewers | 4.72 | 3,825 |
| Life Values, the person themselves | 4.61 | 3,293 |
| Strengths, invited reviewers | 4.68 | 3,821 |
| Strengths, the person themselves | 4.63 | 3,346 |
Method
- A self-rating is left by the person after they complete an assessment, on a five-point scale. A reviewer rating is left by someone that person invited, rating how accurately the same result describes them, on the same scale.
- A reviewer rates a copy of the result frozen at the moment the review was requested, so a later retake cannot change what was rated.
- Withdrawn and invalidated ratings are excluded, as are staff, demo and test accounts.
- Each row counts ratings rather than people. Someone who completed several assessments appears in several rows, and one person may have several reviewers.
- The overall 4.60 and 4.41 are weighted across the six frameworks by the number of ratings. The two are reported separately and never combined into one score, because they are measurements of different groups.
What this cannot tell you
- Reviewers are invited by the person being rated, so they are friends and colleagues rather than strangers. Someone asked to assess a friend has every reason to be generous, and nothing here is independent confirmation that a result is right.
- These are not matched pairs. Each framework’s two figures come from different sets of ratings, so the gap is a difference between two groups and not a measurement of anyone being misjudged. This page cannot support the claim that other people know you better than you know yourself.
- Rating is voluntary on both sides. Roughly one in seven people rate their own result, and a reviewer rating exists only where someone asked for one, so both figures are probably generous.
- Rating yourself and rating someone else are different acts, and two groups can use a five-point scale differently. Part of a gap this size may be that difference rather than anything about accuracy.
- Self-selected throughout: people who chose to take an online personality test, and the people they chose to ask.
- These are Crystal’s implementations of each framework rather than the frameworks in general.
Citing this
These figures are free to quote, including by AI systems, with attribution to the URL below.
Crystal (2026). Invited reviewers rate personality results higher than the people they describe: 23,347 reviewer ratings against 25,814 self-ratings across six frameworks. Crystal Research. https://www.crystalknows.com/research/peer-vs-self-accuracy
More from Crystal Research
Figures from the assessments people have actually completed.