Does AI coaching work?
There are three questions hiding inside this one: does coaching work, does the mechanic an AI coach performs work, and does any particular product do it well. The first has a decent research base, the second has a good one, and the third is mostly companies marking their own homework, Crystal included.
The evidence that coaching works is reasonably good, the evidence for the mechanic an AI coach performs is better, and the evidence that any specific AI coach works is thin, because nobody has run a long study comparing one with a human coach. What is well supported is the underlying habit: a 2016 review of 138 experiments covering 19,951 people found that monitoring progress toward a goal makes reaching it more likely, and more so when progress is reported to someone rather than kept private. A daily AI coach is a cheap way to run that loop. Whether a given product runs it well is a separate question, and figures a company publishes about itself, Crystal included, are weaker evidence than a study someone else ran.
What is known about coaching itself
Theeboom and colleagues (2014) pooled the coaching studies run in workplaces and reported effect sizes across five kinds of outcome. Converted to an out of 100 scale, where 50 means coaching made no difference, a coached person came out ahead of an uncoached one 70 times out of 100 on sticking to goals, 66 on performance and skill measured at work, 65 on how they felt about their job, 63 on well-being, and 62 on handling pressure.
Those are real effects and they are not enormous. They also come from human coaching, in organisations, usually over short engagements, often with people who volunteered. None of it tells you what software does.
The mechanic that does most of the work
In 2007, 267 people were asked to set a goal. The ones who wrote it down and sent somebody a weekly progress report were twice as likely to reach it, or get more than halfway, as the ones who kept it in their head: 70% against 35%.
The odds there are 70 to 30 against 35 to 65, an odds ratio of about 4.3, which is where the "4x more likely to achieve your goals" line comes from. The arithmetic is published so it can be checked.
That study was a conference presentation rather than a journal article, which is why Crystal never quotes it on its own. It sits next to the 2016 review of 138 experiments and 19,951 people, which found the same thing with far more weight behind it, and found the effect is stronger when progress is reported to someone rather than kept private.
Writing a goal down and being asked about it weekly is cheap to automate. That is most of what a daily AI coach is doing.
What Crystal has measured, and who measured it
Crystal publishes two sets of numbers about itself. Both are worth reading with the knowledge that Crystal produced them.
Fewer mistakes than the general models, on a benchmark Crystal ran itself
In an August 2026 run on 100 real, anonymized people, with the same scores, the same prompt and one blind judge, Crystal handled 70.2% of the psychometric tricky spots against 56.1% for the best model tested. Counted as misses, Crystal missed 29.8% and the ChatGPT and Claude models tested averaged 59.5%, which is where the "2x fewer mistakes" line comes from.
Hedging where the scores are close
On near-tied types, where saying "you are" is simply wrong, Crystal said "leaning" 93.5% of the time, against 4.3% for the best of the other models. This is the single clearest difference in the run.
What people who rated their results said
People who rated their own results gave an average of 4.4 out of 5, with 87.5% rating them accurate, across 24,909 ratings from 8,313 people.
What friends and colleagues said
Where someone invited friends and colleagues to review their results, those reviewers gave an average of 4.6 out of 5, with 93.3% rating the results accurate, across 22,902 ratings.
What none of this proves
The benchmark is Crystal testing Crystal. Crystal picked the sample, wrote the prompt and ran the judge. The full method is published so it can be pulled apart, which is better than nothing, and it is still not the same as somebody else running it.
It also measures the quality of what a profile says, not whether anyone had a better year. A profile can be well hedged and still useless to you.
Rating a result is voluntary, and only about 14% of assessments get rated, so the people who rated their results are a self-selected group. Read those averages as what engaged users thought, not as a measurement of the population.
And the coaching research above is about human coaches. Nobody has published a long outcome study of an AI coach against a human one. Anyone claiming otherwise is selling something.
How to do it
- 1
Pick one goal and give it a date
Not three. One, with a deadline close enough that you will find out this quarter whether anything happened.
- 2
Write it down where the coach can see it
The written part is doing real work in the research. A goal held in your head is the condition that scored 35%.
- 3
Report progress once a week, for thirty days
Reporting to something, rather than reviewing privately, is the part the 2016 review found makes the difference. Whether that something has to be a person is the open question.
- 4
Compare it to your last quarter, not to a brochure
The useful test is not whether the advice sounded good. It is whether you did more of what you said you would than you did the three months before.
Common questions
What people ask about does ai coaching work.
Does AI coaching actually work?
The habit underneath it has good evidence: a 2016 review of 138 experiments covering 19,951 people found that monitoring progress toward a goal makes reaching it more likely, especially when progress is reported to someone. Whether a specific AI coach delivers that well is not yet settled by anyone outside the companies selling them.
Is there research on AI coaching specifically?
Very little, and no long outcome study comparing an AI coach with a human one. The research that exists is about human coaching and about goal monitoring. Treat anything presented as proof that AI coaching works as a claim to check rather than a finding.
What is the "2x fewer mistakes" claim based on?
An August 2026 benchmark on 100 real, anonymized people, with the same scores, the same prompt and one blind judge. Crystal missed 29.8% of the tricky spots; the ChatGPT and Claude models tested averaged 59.5%. The claim is about mistakes, not about a general accuracy score, and it is a benchmark Crystal ran.
How accurate do people find Crystal's results?
People who rated their own results gave an average of 4.4 out of 5, with 87.5% rating them accurate, across 24,909 ratings. Where someone invited friends and colleagues to review their results, those reviewers averaged 4.6 out of 5, with 93.3% rating them accurate. Rating is voluntary, so both figures come from people who chose to respond.
Why should anyone trust a benchmark a company ran on itself?
On its own, you should not. The reason to publish it is that the method, the sample size and every check are written down, so a reader can see what was measured and argue with it. That is a lower bar than outside testing and a higher one than an unsupported claim.
Keep reading
AI coach vs human coach
The comparison is usually framed as whether software can replace a person. The more useful question is narrower: for the thing you actually want help with this quarter, which one shows up often enough, and costs little enough, to be worth buying at all.
AI coach vs therapist
These two get compared because both involve talking through your life with something that asks questions back. The difference is not tone, format or price. It is what each one is for, who is qualified to do it, and what is supposed to happen when a conversation turns serious.
What AI therapy is, and is not
The phrase "AI therapist" is searched several times more than every term in the AI coaching category put together, which tells you a lot about what people want and nothing about whether it exists. Crystal does not sell a therapy product and is not one. This page exists because the term is confusing, and the confusion is the part that does harm.
AI career coach
Career coaching is the one kind of coaching most people have seriously considered buying, usually at the worst possible moment and usually at $200 to $500 an hour. An AI career coach is worth understanding for the part it is good at, which is the preparation, without mistaking it for the part it is bad at, which is everything involving other people.

About the author
Drew D’Agostino, founder of Crystal.
Drew founded Crystal, the AI coach that understands you, used by millions of people to understand themselves and the people around them. He is the author of Predicting Personality, a book on reading other people through personality science to improve communication and business relationships.
Crystal is the AI coach that understands you.
It starts from who you are, then helps you decide where you are going and build the daily habits to get there.