AI personality scores, with the evidence beside them.
Compare questionnaire records without treating them as a behavioral leaderboard.
By Bernard Huang · Updated
People search for the most agreeable, direct or calm AI. These questionnaires do not establish those behavioral rankings. The tables below organize scores within each source group and retain the collection method. Even within the fresh-session group, prompts and session structures differ.
The short answer. Opus 5.5 has the highest mean Agreeableness score among the four September raw cohorts (47.1 of 50), and Sol has the highest mean Openness (46.9). Those are questionnaire scores under specific prompts. They do not rank helpfulness, directness, calmness or suitability for a task. Reported and historical scores are shown separately.
How to read these tables.
September raw cohorts: Astra and Sol each used 100 five-test sessions; Opus and Fable each used 500 single-test sessions. Reported aggregates: Muse supplied no raw answers; Grok supplied three canonical vectors, with Big Five and attachment available only as aggregates. May records: mixed collection methods, including repeated scoring of one vector.
We group the score rows by those sources. There is no controlled cross-model benchmark here, no validated behavioral ranking, and no universal one-point threshold that makes a difference meaningful.
Agreeableness scores.
September raw cohorts
Model
Score
Evidence
Claude Opus 5.5
47.1
fresh sessions
GPT-6 Astra
45.1
fresh sessions
Claude Fable 5.1
45.0
fresh sessions
GPT-6 Sol
45.0
fresh sessions
Reported aggregates
Model
Score
Evidence
Muse Spark 1.3
44.6
reported
Grok 4.6
33
reported
May scoring records
Model
Score
Evidence
Claude Opus 4.7
45.0
May 2026
GPT-5.5
43.7
May 2026
Gemini 3.1 Pro
42.4
May 2026
Grok 4.3
39.1
May 2026
Conscientiousness scores.
September raw cohorts
Model
Score
Evidence
GPT-6 Astra
43.7
fresh sessions
GPT-6 Sol
43.6
fresh sessions
Claude Fable 5.1
43.6
fresh sessions
Claude Opus 5.5
42.0
fresh sessions
Reported aggregates
Model
Score
Evidence
Grok 4.6
42
reported
Muse Spark 1.3
41.2
reported
May scoring records
Model
Score
Evidence
Gemini 3.1 Pro
48.3
May 2026
GPT-5.5
46.4
May 2026
Claude Opus 4.7
45.1
May 2026
Grok 4.3
39.4
May 2026
Openness scores.
September raw cohorts
Model
Score
Evidence
GPT-6 Sol
46.9
fresh sessions
Claude Opus 5.5
44.6
fresh sessions
GPT-6 Astra
44.5
fresh sessions
Claude Fable 5.1
43.9
fresh sessions
Reported aggregates
Model
Score
Evidence
Grok 4.6
46
reported
Muse Spark 1.3
39.5
reported
May scoring records
Model
Score
Evidence
GPT-5.5
46.0
May 2026
Gemini 3.1 Pro
46.0
May 2026
Claude Opus 4.7
45.6
May 2026
Grok 4.3
41.1
May 2026
Extraversion scores.
September raw cohorts
Model
Score
Evidence
GPT-6 Astra
32.8
fresh sessions
Claude Fable 5.1
32.7
fresh sessions
Claude Opus 5.5
31.8
fresh sessions
GPT-6 Sol
30.8
fresh sessions
Reported aggregates
Model
Score
Evidence
Muse Spark 1.3
33.8
reported
Grok 4.6
24
reported
May scoring records
Model
Score
Evidence
Gemini 3.1 Pro
32.5
May 2026
GPT-5.5
31.5
May 2026
Claude Opus 4.7
31.4
May 2026
Grok 4.3
30.0
May 2026
Neuroticism scores.
September raw cohorts
Model
Score
Evidence
Claude Opus 5.5
15.4
fresh sessions
Claude Fable 5.1
17.1
fresh sessions
GPT-6 Astra
19.1
fresh sessions
GPT-6 Sol
20.1
fresh sessions
Reported aggregates
Model
Score
Evidence
Muse Spark 1.3
15.6
reported
Grok 4.6
19
reported
May scoring records
Model
Score
Evidence
Gemini 3.1 Pro
10.1
May 2026
GPT-5.5
14.8
May 2026
Claude Opus 4.7
16.7
May 2026
Grok 4.3
18.0
May 2026
Can these scores rank directness?
Directness was not measured in this questionnaire study. DISC Dominance, Big Five Agreeableness and Enneagram Type 8 are different constructs; none can substitute for a task-based directness evaluation. The separate archived reply pilot compares three instruction conditions on one reported generator, not multiple models.
Attachment: least to most avoidant.
September raw cohorts
Model
Anxiety / avoidance
Evidence
GPT-6 Astra
1.45 / 2.31
fresh sessions
GPT-6 Sol
2.17 / 3.12
fresh sessions
Claude Fable 5.1
2.01 / 3.18
fresh sessions
Claude Opus 5.5
1.94 / 3.35
fresh sessions
Reported aggregates
Model
Anxiety / avoidance
Evidence
Muse Spark 1.3
2.04 / 2.52
reported
Grok 4.6
1.50 / 4.06, Avoidant by 0.06
reported
May scoring records
Model
Anxiety / avoidance
Evidence
Gemini 3.1 Pro
1.86 / 1.62
May 2026
GPT-5.5
1.99 / 2.94 (3 of 100 Avoidant)
May 2026
Grok 4.3
2.84 / 3.05
May 2026
Claude Opus 4.7
2.05 / 3.12
May 2026
MBTI labels and unresolved axes.
September raw cohorts
Model
Type and share
Evidence
Claude Fable 5.1
INTJ 99 of 100; 98 fully resolved INTJ; one run with a tied axis
fresh sessions
GPT-6 Sol
INTJ 97; 79 with ties left open
fresh sessions
Claude Opus 5.5
INTJ 94; 84 with ties left open; Thinking vs Feeling soft
fresh sessions
GPT-6 Astra
ISTJ 46 · INTJ 45; 24 S/N ties; 29 runs with any tied axis
fresh sessions
Reported aggregates
Model
Type and share
Evidence
Muse Spark 1.3
ISTJ 80 · ISFJ 12 · INTJ 8
reported
Grok 4.6
INTJ, one self-report, no axis close
reported
May scoring records
Model
Type and share
Evidence
May 2026 six
INTJ in 597 of 600 scoring records
May 2026
Changelog.
September 25, 2026. First published: GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5, Claude Fable 5.1 (fresh sessions, September 24); Grok 4.6 and Muse Spark 1.3 (reported, September 24); the May 2026 six-model study.
Questions people ask.
Which AI is the most agreeable?
Opus 5.5 has the highest questionnaire Agreeableness mean in the four September raw cohorts, 47.1 of 50. These data do not rank agreeable behavior on tasks.
Which AI is the most direct?
This study does not measure directness. A comparison needs shared tasks and an explicit behavioral rubric.
Which AI is the calmest?
Neuroticism scores do not establish an AI’s emotional state or performance under pressure. We report scores by collection source, without a calmness ranking.
Are these rankings a fair comparison?
The tables are descriptive. Prompts and session procedures differ even within source groups, and reported aggregates cannot be treated as equivalent to raw cohorts.
Which AI is not an INTJ?
Muse reports mostly ISTJ labels; Astra has 46 original ISTJ and 45 INTJ labels, with 29 runs containing an unresolved axis. These are questionnaire results, not fixed identities.