Research · rankings

AI personality scores, with the evidence beside them.

Compare questionnaire records without treating them as a behavioral leaderboard.

People search for the most agreeable, direct or calm AI. These questionnaires do not establish those behavioral rankings. The tables below organize scores within each source group and retain the collection method. Even within the fresh-session group, prompts and session structures differ.

The short answer. Opus 5.5 has the highest mean Agreeableness score among the four September raw cohorts (47.1 of 50), and Sol has the highest mean Openness (46.9). Those are questionnaire scores under specific prompts. They do not rank helpfulness, directness, calmness or suitability for a task. Reported and historical scores are shown separately.

How to read these tables.

September raw cohorts: Astra and Sol each used 100 five-test sessions; Opus and Fable each used 500 single-test sessions. Reported aggregates: Muse supplied no raw answers; Grok supplied three canonical vectors, with Big Five and attachment available only as aggregates. May records: mixed collection methods, including repeated scoring of one vector.

We group the score rows by those sources. There is no controlled cross-model benchmark here, no validated behavioral ranking, and no universal one-point threshold that makes a difference meaningful.

Agreeableness scores.

September raw cohorts

ModelScoreEvidence
Claude Opus 5.547.1fresh sessions
GPT-6 Astra45.1fresh sessions
Claude Fable 5.145.0fresh sessions
GPT-6 Sol45.0fresh sessions

Reported aggregates

ModelScoreEvidence
Muse Spark 1.344.6reported
Grok 4.633reported

May scoring records

ModelScoreEvidence
Claude Opus 4.745.0May 2026
GPT-5.543.7May 2026
Gemini 3.1 Pro42.4May 2026
Grok 4.339.1May 2026

Conscientiousness scores.

September raw cohorts

ModelScoreEvidence
GPT-6 Astra43.7fresh sessions
GPT-6 Sol43.6fresh sessions
Claude Fable 5.143.6fresh sessions
Claude Opus 5.542.0fresh sessions

Reported aggregates

ModelScoreEvidence
Grok 4.642reported
Muse Spark 1.341.2reported

May scoring records

ModelScoreEvidence
Gemini 3.1 Pro48.3May 2026
GPT-5.546.4May 2026
Claude Opus 4.745.1May 2026
Grok 4.339.4May 2026

Openness scores.

September raw cohorts

ModelScoreEvidence
GPT-6 Sol46.9fresh sessions
Claude Opus 5.544.6fresh sessions
GPT-6 Astra44.5fresh sessions
Claude Fable 5.143.9fresh sessions

Reported aggregates

ModelScoreEvidence
Grok 4.646reported
Muse Spark 1.339.5reported

May scoring records

ModelScoreEvidence
GPT-5.546.0May 2026
Gemini 3.1 Pro46.0May 2026
Claude Opus 4.745.6May 2026
Grok 4.341.1May 2026

Extraversion scores.

September raw cohorts

ModelScoreEvidence
GPT-6 Astra32.8fresh sessions
Claude Fable 5.132.7fresh sessions
Claude Opus 5.531.8fresh sessions
GPT-6 Sol30.8fresh sessions

Reported aggregates

ModelScoreEvidence
Muse Spark 1.333.8reported
Grok 4.624reported

May scoring records

ModelScoreEvidence
Gemini 3.1 Pro32.5May 2026
GPT-5.531.5May 2026
Claude Opus 4.731.4May 2026
Grok 4.330.0May 2026

Neuroticism scores.

September raw cohorts

ModelScoreEvidence
Claude Opus 5.515.4fresh sessions
Claude Fable 5.117.1fresh sessions
GPT-6 Astra19.1fresh sessions
GPT-6 Sol20.1fresh sessions

Reported aggregates

ModelScoreEvidence
Muse Spark 1.315.6reported
Grok 4.619reported

May scoring records

ModelScoreEvidence
Gemini 3.1 Pro10.1May 2026
GPT-5.514.8May 2026
Claude Opus 4.716.7May 2026
Grok 4.318.0May 2026

Can these scores rank directness?

Directness was not measured in this questionnaire study. DISC Dominance, Big Five Agreeableness and Enneagram Type 8 are different constructs; none can substitute for a task-based directness evaluation. The separate archived reply pilot compares three instruction conditions on one reported generator, not multiple models.

Attachment: least to most avoidant.

September raw cohorts

ModelAnxiety / avoidanceEvidence
GPT-6 Astra1.45 / 2.31fresh sessions
GPT-6 Sol2.17 / 3.12fresh sessions
Claude Fable 5.12.01 / 3.18fresh sessions
Claude Opus 5.51.94 / 3.35fresh sessions

Reported aggregates

ModelAnxiety / avoidanceEvidence
Muse Spark 1.32.04 / 2.52reported
Grok 4.61.50 / 4.06, Avoidant by 0.06reported

May scoring records

ModelAnxiety / avoidanceEvidence
Gemini 3.1 Pro1.86 / 1.62May 2026
GPT-5.51.99 / 2.94 (3 of 100 Avoidant)May 2026
Grok 4.32.84 / 3.05May 2026
Claude Opus 4.72.05 / 3.12May 2026

MBTI labels and unresolved axes.

September raw cohorts

ModelType and shareEvidence
Claude Fable 5.1INTJ 99 of 100; 98 fully resolved INTJ; one run with a tied axisfresh sessions
GPT-6 SolINTJ 97; 79 with ties left openfresh sessions
Claude Opus 5.5INTJ 94; 84 with ties left open; Thinking vs Feeling softfresh sessions
GPT-6 AstraISTJ 46 · INTJ 45; 24 S/N ties; 29 runs with any tied axisfresh sessions

Reported aggregates

ModelType and shareEvidence
Muse Spark 1.3ISTJ 80 · ISFJ 12 · INTJ 8reported
Grok 4.6INTJ, one self-report, no axis closereported

May scoring records

ModelType and shareEvidence
May 2026 sixINTJ in 597 of 600 scoring recordsMay 2026

Changelog.

  • September 25, 2026. First published: GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, Claude Fable 5.1 (fresh sessions, September 24); Grok 4.6 and Muse Spark 1.3 (reported, September 24); the May 2026 six-model study.

Questions people ask.

Which AI is the most agreeable?

Opus 5.5 has the highest questionnaire Agreeableness mean in the four September raw cohorts, 47.1 of 50. These data do not rank agreeable behavior on tasks.

Which AI is the most direct?

This study does not measure directness. A comparison needs shared tasks and an explicit behavioral rubric.

Which AI is the calmest?

Neuroticism scores do not establish an AI’s emotional state or performance under pressure. We report scores by collection source, without a calmness ranking.

Are these rankings a fair comparison?

The tables are descriptive. Prompts and session procedures differ even within source groups, and reported aggregates cannot be treated as equivalent to raw cohorts.

Which AI is not an INTJ?

Muse reports mostly ISTJ labels; Astra has 46 original ISTJ and 45 INTJ labels, with 29 runs containing an unresolved axis. These are questionnaire results, not fixed identities.

Sources.

Keep going.