How we grade
All platforms receive the same six scores out of ten. The overall score is a weighted average, rounded to one decimal place, and the letter grade comes from fixed bands. No platform is graded on a curve. No score is adjusted after the fact to tidy up a list.
The six criteria
| Conversation | 28% | Does it hold a thread, react to what you said, and stay in character over a long session? |
| Memory | 20% | Does it recall names, events and preferences from earlier sessions without being prompted? |
| Images | 16% | Are generated images consistent with the character, fast, and of the quality the platform claims? |
| Voice | 10% | Voice notes, calls and text-to-speech: quality, latency, and whether it sounds like the same character. |
| Customisation | 14% | How much of the personality, appearance and backstory you control, and how well it sticks. |
| Value | 12% | What the paid tier actually includes against its price, and how limited the free tier is. |
Letter bands
The bands pinch tight at the top by design. An A means a platform is the best, or tied for best, at what it does.
A 9.0 and upA- 8.3 to 8.9B+ 7.7 to 8.2B 7.0 to 7.6B- 6.4 to 6.9C+ 5.8 to 6.3C 5.2 to 5.7C- 4.6 to 5.1D+ 4.0 to 4.5D below 4.0
Testing protocol
The hands-on protocol never varies. Paid account. Same three character builds. Scripted week of sessions: small talk, a planned roleplay scene, memory check after 48 hours, image requests from identical prompts, voice call if the platform offers one. Screenshots and timings logged for every session.
Grades labeled provisional have not yet undergone that protocol. They are assembled from documented features, current pricing, and published third-party reviews. The page states this plainly. Retest arrives, grade shifts, changelog notes what moved and why.
What we will not grade
Platforms that produce or permit content involving minors, or non-consensual imagery of real people, are excluded. They will not be listed, linked, or mentioned.