Crystal

Benchmarks

Pasting your DISC scores into ChatGPT? It’s guessing.

The answer will sound right. It will even sound like you. But a model that has never measured real people fills the gaps with whatever sounds plausible, and it never says when your results are too close to call. Crystal checks every profile against patterns measured from thousands of real people, and we put both approaches head to head. This page is the scoreboard.

Outscores all 4 leading AI models18,500+ profiles checked and counting6 frameworks checked against each other15 expert-written accuracy checks
Head to headAugust 2026

How many of the tricky spots in a person’s results each profile handles

100 real, anonymized people. Same scores, same prompt, one blind judge. Higher is better.

Why it fools people

Fluent is easy. True is the hard part.

Plenty of people paste their assessment results into ChatGPT and take what comes back at face value. So we ran the experiment properly. We gave the same scores and the same prompt to GPT-4o, GPT-5.6 Terra, Claude Sonnet 4.6, Claude Opus 4.8, and Crystal, then graded every profile blind, against the same rules.

A general-purpose model writes something fluent, confident, and generic. It has never seen how DISC, Enneagram, and Big Five results relate to each other in real people, so it fills the gaps with guesses. And it never tells you when your results are too close to call.

That last part matters more than it sounds. In our data, most people’s top Enneagram type leads the runner-up by a single point. That is a statistical tie. A model that has never measured real people states it as fact, just more eloquently. The rest of this page is that difference, measured.

The short version

The gap, in three numbers.

70%

of the tricky spots in a person’s results handled by Crystal. The strongest of the big models, GPT-5.6, manages 56%. GPT-4o manages 13%.

93% vs 4%

When two types are nearly tied, Crystal says “leaning” instead of “you are” 93% of the time. No raw model clears 5%, the flagships included.

4 of 4

models outscored overall, while Crystal runs on a smaller, cheaper base model. The priciest model tested, Claude Opus 4.8, scores no better than mid-tier Sonnet 4.6. Raw horsepower is not what this task rewards. Knowing real people is.

Where the gap comes from

Four checks, five models.

Every set of results has tricky spots: two frameworks that disagree, a score sitting on a fence, a near-tie between types. (Psychometricians call these tensions.) These are the places where a profile is most likely to tell you something your results don’t support, so that is exactly where we check. Four kinds of checks, five models, here is how they did.

CrystalGPT-5.6 TerraClaude Sonnet 4.6Claude Opus 4.8GPT-4o

Agreement between frameworks

When two tests measure the same trait, does the profile notice whether they agree, and say so when they don’t?

Emphasis matches your scores

Is the profile loud where your scores are extreme, and quiet where they’re average?

Careful near the midline

When someone lands near the middle of a 16 Personalities scale, a soft E versus I, say, does it hedge instead of stating a coin flip as fact?

Honest about near-ties

When two types are nearly tied, does it say “leaning” instead of “you are”?

Crystal leads three of the four, and the near-tie check is the clearest of all: every raw model scores near zero while Crystal clears 93%. The exception is the midline check, where every current model clusters near 50% and our checks do not yet pull ahead. We publish that rather than hide it; it is the gap we are working on next.

How we ran it

Same scores, same prompt, one blind judge.

Every benchmark number comes from one controlled comparison across 100 real, anonymized people. Each model received identical scores and an identical prompt. The tricky spots come from a person’s scores, not from anything a model wrote, so the exact same checks graded all five profiles.

Crystal

Our validated system. A mid-tier base model plus Crystal’s rule layer, which finds the tricky spots in each person’s results and makes sure the profile handles them before it’s shown.

GPT-5.6 Terra

OpenAI’s biggest current model. Given the same scores and prompt, with no access to Crystal’s rules.

Claude Opus 4.8

Anthropic’s flagship model. The identical task, with none of Crystal’s rules.

Claude Sonnet 4.6

Anthropic’s mid-tier model. A baseline at the same model tier Crystal builds on.

GPT-4o

Prior-generation baseline. Shows how far general models have come, and how far the gap still is.

The judge is blind, and the rules were locked first. Each check is a simple yes-or-no question, and the judge never knows which model wrote which profile. The rules themselves were measured from thousands of real assessments before any of these profiles existed. The full fairness story, including how a human expert double-checks the automated judge, lives in the FAQ at the bottom of this page.

One caveat: read the figures per check, not as one overall average. The sample includes extra people with rare kinds of tricky spots, so every check had enough cases to measure reliably.

The foundation

Built on real people, not internet text.

Why can’t a bigger model close the gap? Because the checks are not opinions about good writing. They come from how 59 different personality measurements move together across six frameworks, drawn from people who completed more than one assessment on Crystal. We know how these tests relate in real people because we measured it ourselves, and we re-check the numbers on fresh, independent samples as our community grows.

59

personality measurements mapped against each other, across DISC, Big Five, Enneagram, 16 Personalities, Values, and Strengths

r = 0.79

agreement between two different frameworks’ measures of social energy, a strong match, and identical in two independent samples (June and July 2026)

15,000+

assessments in our latest measurement sample

every key relationship confirmed on a second, independent sample before we rely on it

The people

Checked by humans who do this for a living.

The rules are not a black box, and neither is the scoring. Every rule was written from our data by a consulting psychometrician, a specialist in how personality is measured, and the automated scoring is compared with expert review, profile by profile, until they agree.

“A profile that sounds right but isn’t grounded in your scores is worse than one that’s awkward and true. We built the scoring to prefer the truth.”
Vladimir Novkov, consulting psychometrician, Crystal
≤ 0.12

largest gap between expert and automated scores on any profile section reviewed so far, on a 0-to-1 scale, with the same accept-or-reject conclusion every time (21 section scores, July 2026, reviews ongoing)

0

cases where automated scoring passed a section the expert would fail. On every noted difference, the human was the stricter grader (July 2026)

Beyond the lab

The same story on live profiles.

The benchmark is a controlled test. This is the same measurement running on live profiles for real users, every day. Before a profile is written, fixed rules, no AI involved, scan the person’s scores for tricky spots. Each one becomes something the profile has to handle, and we check whether it did.

Tricky spots the profile actually handles, on live profiles

And you can watch it happen, day by day

Share of tricky spots handled, per day, as the checks rolled out.

rollout beginschecks fully on

The step up is Crystal’s checks switching on. It holds. We publish the whole trend, not two hand-picked snapshots, so you can see it is a lasting change, not a good day.

62% is a strict number on purpose. The checks fire on almost everybody, and a missed spot means a profile could have been more careful; nobody was told anything false. The FAQ explains why we grade ourselves this hard.

Over the same period, profiles that cleared every single check they triggered went from 10% to 41%, and our internal quality score rose from 0.64 to 0.73. All of it measured across 7,651 live profiles, every one of which shipped to a real user.

The fine print, up front

What we won’t tell you.

We won’t call a lean a type. For most people in our data, the gap between their top two Enneagram types is a single point. When that’s you, your profile says “leaning,” because that is what your results support.

We won’t hide disagreement between frameworks. If your Big Five and 16 Personalities results pull in different directions, that tension becomes part of your profile, not an inconvenience to smooth over.

We won’t hide a metric we lose. On the midline check, today’s biggest AI models score about as well as we do. That number is on this page, and closing it is our next project.

We won’t pretend personality is destiny. Scores carry uncertainty and people change. A profile is a well-grounded starting point, not a verdict.

Your assessment responses are used to generate your profile and, in anonymized, aggregate form, to keep these checks accurate. Privacy policy

Questions we get

Answers about the benchmark.

Is Crystal just ChatGPT with extra steps?+

No. Crystal uses an AI model to write your profile, but the part that makes it accurate is not the model. It is a set of rules built from how 59 personality measurements move together in real people who completed more than one assessment. Those rules check every profile against your actual scores before you see it. In our benchmark, that rule layer beat every big model tested, including ones far larger than the model Crystal runs on.

What do you mean by “tricky spots” in my results?+

Any spot that needs careful handling: two frameworks disagreeing about the same trait, a score sitting right in the middle of a scale, or two types finishing in a near-tie. Psychometricians call these tensions. They are the places where a profile is most likely to tell you something your results can’t back up, which is why our checks focus on them.

How do I know the benchmark wasn’t rigged in Crystal’s favor?+

Three safeguards. Every model got identical scores and an identical prompt. The judge is blind: it grades each profile with simple yes-or-no checks and never knows which model wrote it. And the rules were measured from thousands of real assessments before any of these profiles existed, so they could not be tuned to favor one output. The automated judge is also double-checked by a human expert: their scores differ by at most 0.12 on a 0-to-1 scale, and the automation has never been more lenient than the human.

Why is Crystal’s score 70% and not closer to 100%?+

Because the checks are deliberately strict, and most of them fire on almost everybody. The “don’t over-claim a type” check alone fires on roughly four out of five profiles. A missed spot means a profile could have been more careful, not that it told you something false. We publish the strict number because a lenient one would tell us nothing, and because it leaves us somewhere to keep improving.

Wouldn’t a bigger, newer AI model close the gap?+

The benchmark says no. The most expensive model tested, Claude Opus 4.8, scored no better than mid-tier Sonnet 4.6, and Crystal beat both while running on a smaller base model. What counts as a real contradiction, or a statistical tie, was measured from real people. A model cannot reason its way to that knowledge from your scores alone, no matter how large it is.

Do I have to take all six assessments for this to work?+

No. Start with whichever one fits your goal; a single assessment gives you a complete, validated profile. The cross-framework checks simply get more to work with if you choose to add more over time.

What happens with my assessment data?+

Your responses generate your profile. In anonymized, aggregate form they also keep the checks accurate, which is how the rules stay grounded in real people rather than assumptions. See our privacy policy for the full details.

How often are these numbers updated?+

The footer carries the date of the last update, and the benchmark itself is dated where shown. When we re-run the comparison, or the production numbers meaningfully change, we update the page rather than let it drift.

Your results deserve more than a guess.

Every Crystal profile has to answer to the checks on this page before you see it. The easiest way to feel the difference: 28 questions, about five minutes, and a profile that answers to your scores.

Take the free personality test
No signup to startInstant results

Benchmark: 100 anonymized users vs GPT-4o, GPT-5.6 Terra, Claude Sonnet 4.6, and Claude Opus 4.8, graded blind by the same checks. The rules and thresholds themselves are proprietary. Last updated .