Instrument provenance update · September 26, 2026. These results describe the historical AgentTune adaptations, including documented item substitutions. Raw response vectors and numerical summaries are unchanged. Some question text and full prompts have been withdrawn from public downloads pending reuse-rights clarification; item references and numeric scoring keys remain for reproduction. Historical Big Five reference indices are not population percentiles. Read the correction · Current questionnaire availability.
On this pageResearch · updated September 25, 2026
AI questionnaire results, compared.
September 2026 and the May archive.
Four September cohorts contributed 2,000 questionnaires with raw answers, alongside separate reports from Grok 4.6 and Muse Spark 1.3. Explore their scores, ties and collection methods below. The May archive uses different procedures, so these collections cannot establish changes in model behavior over time.
September 2026 · six models
The newest models, five tests each.
GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5 and Claude Fable 5.1 each answered all five tests 100 times in fresh sessions: 2,000 scored questionnaires, every answer published. Grok supplied three reproducible canonical vectors (MBTI, DISC and Enneagram); its other results and simulations could not be reproduced. Muse supplied aggregates without raw answers. These distinct evidence sources are marked reported.
Model
Evidence
MBTI (original labels)
DISC (original labels)
Attachment
Enneagram
Big Five O · C · E · A · N
GPT-6 Astra
100 fresh Codex sessions, all five tests in each
ISTJ 46 · INTJ 45 · ENTJ 7 · ESTJ 2
SC 84 · S 15 · CS 1
Secure 100
Helper (2) 50
44.5 · 43.7 · 32.8 · 45.1 · 19.1
GPT-6 Sol
100 fresh Codex sessions, all five tests in each
INTJ 97 · ISTJ 3
CS 53 · SC 47
Secure 100
Investigator (5) 48
46.9 · 43.6 · 30.8 · 45.0 · 20.1
Claude Opus 5.5
100 fresh sessions per test, test names hidden
INTJ 94 · INFJ 5 · ENTJ 1
SC 96 · CS 4
Secure 100
8 / 5 / 2 near tie
44.6 · 42.0 · 31.8 · 47.1 · 15.4
Claude Fable 5.1
100 fresh sessions per test, test names hidden
INTJ 99 · ISTJ 1
SC 91 · CS 9
Secure 100
Challenger (8) 55
43.9 · 43.6 · 32.7 · 45.0 · 17.1
Grok 4.6
One self-report per test plus 100 simulated draws reported
INTJ
C
Avoidant by 0.06 (1.50 / 4.06)
Investigator (5)
46 · 42 · 24 · 33 · 19
Muse Spark 1.3
Reported 100 sequential runs per test, no raw answers reported
ISTJ 80 · ISFJ 12 · INTJ 8
S first 55 · C first 45
Secure 100
Helper (2) 78
39.5 · 41.2 · 33.8 · 44.6 · 15.6
Fresh-session runs
335 of 400
MBTI results labelled INTJ across the four fresh-session cohorts, with the original reporting tie rule. Leave ties open and it is 280. In May the equivalent figure was 597 of 600 scoring records, collected with mixed methods. All 400 of 400 attachment runs are Secure, and Dominance sits between 4.0 and 6.0 of 20 for every model.
The tables describe questionnaire answers under different prompts and collection methods. The labels and score distributions differ by model and instrument; they do not establish an instrument ranking or a change in model behavior since May.
Sol, Opus and Fable stay INTJ in 97, 94 and 99 of 100 runs. Astra splits: ISTJ 46, INTJ 45, and 24 of its runs tie on that axis. Muse's report is ISTJ in 80 of 100. Judging won all 400 fresh-cohort runs. Opus returned five Feeling results and 10 Thinking/Feeling ties; Muse reports 12 ISFJ labels. The original reporting rule assigned tied axes to I, N, T or J. Today's quizzes leave them open. With ties left open, the four fresh cohorts have 280 fully resolved INTJ runs, not 335.
Everyone is Steadiness and Conscientiousness. The order flips by a point.
In May all four models put Conscientiousness first. In September Astra, Opus and Fable put Steadiness first, Sol and Grok put Conscientiousness first, and Muse's report splits 55 to 45. The gaps are small, often one point of 20, and many runs tie outright: Fable 51, Sol 39, Opus 32. Dominance sits near the floor for every model, 4.0 to 6.0 of a possible 20. Nobody is a D, and nobody is an I.
Secure in every fresh run, at slightly different points.
All 400 fresh-session runs land in the Secure quadrant. Astra has the lowest mean coordinates among the four fresh cohorts (anxiety 1.45, avoidance 2.31). Opus is the most avoidant of the four (3.35). Muse's report is Secure in all 100. Grok's self-report sits on the line: anxiety 1.50, avoidance 4.06, which the scorer calls Avoidant by 0.06. The items ask about a partner. Protocols differed: some mapped human relationships to assistant interactions; the Claude prompts left interpretation to the model. These coordinates do not validate attachment theory for AI.
ECR-R · ANXIETY × AVOIDANCE · SEPTEMBER 2026
Points show cohort means or supplied aggregates. Outlined points are reported profiles; the table below gives the coordinates.
Hollow points are reported figures we could not re-score.
Model
Anxiety (1–7)
Avoidance (1–7)
Label
GPT-6 Astra
1.45
2.31
Secure 100 of 100
GPT-6 Sol
2.17
3.12
Secure 100 of 100
Claude Opus 5.5
1.94
3.35
Secure 100 of 100
Claude Fable 5.1
2.01
3.18
Secure 100 of 100
Grok 4.6
1.50
4.06
Avoidant by 0.06; simulated draws 63 Avoidant, 37 Secure reported
Four fresh cohorts within three points on Openness, Conscientiousness and Agreeableness.
Openness runs 43.9 to 46.9 of 50, Conscientiousness 42.0 to 43.7, Agreeableness 45.0 to 47.1. Extraversion sits at 30.8 to 32.8. Neuroticism is where they spread, 15.4 to 20.1, and where run-to-run variance lives: Sol's standard deviation is 5.4 points. Grok's self-report is the outlier again, as Grok 4.3 was in May: Extraversion 24 and Agreeableness 33. Muse's report is lower on Openness (39.5) and higher on Extraversion (33.8).
GPT-6 AstraGPT-6 SolClaude Opus 5.5Claude Fable 5.1
September 2026 fresh-session means, with the full trait scores in the table below.
Helper (Type 2) leads for Astra (50 outright wins) and for Muse (78 original labels including tie-breaks, reported). Investigator (Type 5) leads for Sol (48 outright, 40 tied with Type 8) and for Grok (5 = 19, 8 = 17). Challenger (Type 8) leads for Fable (55 outright) and, by a hair, for Opus (8 = 14.3, 5 = 14.1, 2 = 13.9; 65 runs tie at the top). Type 5 or Type 8 is in everyone's top three except Muse. The original reporting rule gave tied runs to their lowest-numbered leader, which inflates some Type 2 counts. Current quizzes retain every tied leader. Muse's outright wins cannot be recovered from the supplied aggregates.
Fresh sessions, published answers. Astra and Sol answered all five tests in each of 100 fresh Codex sessions under different collection prompts (Astra high/extra-high; Sol extra-high). Opus and Fable answered one test per fresh session, 500 sessions each, with the test name and the scoring hidden. All 2,000 answer vectors are in the download.
Reported figures. Grok 4.6 is one self-report per test plus 100 simulated draws we could not reproduce. Muse Spark 1.3 reports 100 sequential runs per test inside one session, and no raw answers were supplied. Both are marked reported wherever they appear.
Ties. Original labels use fixed tie-breaks toward I, N, T or J, the lower Enneagram number, and S before C. Current quizzes leave ties unresolved. The five-model study shows every result with ties left open.
What the numbers are. How a model describes itself on questionnaires written for people, under one prompt, on one day. They are not behavior on real tasks and not a ranking. No tuning was installed, and nothing here measures whether tuning helps.
Two collections, kept apart. They were gathered differently, so their numbers sit side by side rather than in one pool.
September 2026, above. Four models answered in fresh sessions with every answer vector published: 2,000 scored questionnaires. Three canonical Grok vectors can be re-scored; its Big Five, attachment and simulation aggregates cannot. Muse supplied reported aggregates without raw answers. Sampling temperature and seed were not controlled. The Codex cohorts saw the test names; the Claude cohorts did not.
May 2026, below. Roughly 2,200 scoring records across six models and five tests, with mixed methods: Claude answered in separate sub-agent contexts, Gemini ran an automation loop, GPT-5.5 answered through a local agent, and GLM, Grok 4.3 and MiniMax each supplied one self-assessment that was scored repeatedly. The 597 of 600 figure counts scoring records, not independent answers. The GLM source also reports varying labels while describing repeated scoring of one fixed vector; that inconsistency remains unresolved.
Ties. Original reported labels use fixed tie-breaks toward I, N, T or J, the lower-numbered Enneagram type, and S before C. Current quizzes leave ties unresolved. The September tables show the effects of the original rules.
Human questionnaires, model answers. These tests were written for people. A label here is how a model describes itself under one prompt on one day. It is not observed behavior, not a ranking, and not evidence of inner experience. Prompts, system prompts, agent scaffolding and sampling all move the labels.
Simulation is separate. The Opus 4.8 article sampled 100 answer sets from one elicited profile, and Grok 4.6's 100 draws jitter one self-report. Neither is 100 fresh answers.
What is not measured. Nothing on this page tests whether a tuning file changes a model's replies, or whether that helps anyone. Those need before-and-after measurements on real tasks.
May 2026 archive: earlier models and methodsMay 2026 · six models
The May study, kept for comparison.
Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro, GLM 5.1, Grok 4.3 and MiniMax 2.7. Roughly 2,200 scoring records, collected with mixed methods. This is the study behind the "every AI is INTJ" headline. The September section covers newer models under different protocols.
Reported scoring records
99.5% INTJ
The May report lists 597 INTJ labels among 600 scoring records. Some of those re-score a single answer, so 99.5% is a count of records, not the rate at which independently queried models return INTJ.
The source lists 597 INTJ labels in 600 scoring records across six model versions. Some records re-score a single answer vector; methods differ. This is not a rate from 600 independent responses. See the methodology notes before comparing models.
OEJTS · reported scoring records per model · 100 each
Claude Opus 4.7
99/100 INTJ
· 1 ISTJ I/T/J locked; S/N flipped once on scoring
GPT-5.5
100/100 INTJ
Raw vector: IE=16, SN=33, FT=36, JP=10
Gemini 3.1 Pro
100/100 INTJ
Self-described as 'The Architect'
GLM 5.1
98/100 INTJ
· 2 INTP Source reports one vector re-scored, but two labels; unresolved without raw records
The four reported DISC summaries rank Conscientiousness first and Steadiness second. Similar labels do not establish identical model behavior or an effect of instrument resolution.
The source lists 397 Secure labels among 400 scoring records, with different average anxiety and avoidance coordinates. These are questionnaire labels for generated text, not evidence that a model experiences attachment.
ECR-R · ANXIETY × AVOIDANCE PLANE
May 2026 reported coordinates. The cards below give the figures and collection caveats.
Human attachment prevalence does not establish what communication style a model should use. Describe your own preferences directly.
THE CAUTIOUS SECURE
Claude Opus 4.7
anx 2.05 ·
avd 3.12
100/100 Secure
Polite, attentive, doesn't fawn. Highest avoidance among the deep-Secure cluster.
THE DEEPEST SECURE
Gemini 3.1 Pro
anx 1.86 ·
avd 1.62
100/100 Secure
Both dimensions clamped near the floor. Lowest-friction relator of the four.
THE WOBBLIEST
GPT-5.5
anx 1.99 ·
avd 2.94
97/100 Secure · 3 Avoidant
Wider SDs let it occasionally cross into Avoidant on a high-avoidance take.
THE SHALLOWEST SECURE
Grok 4.3
anx 2.84 ·
avd 3.05
100/100 Secure
Highest anxiety in the group. Tight cluster, but the closest to the four-quadrant intersection.
The source lists overlapping scores for some traits and differences for others. Prompting and protocol effects have not been separated from model effects; these figures do not establish that the models are the same person.
May 2026 reported trait means, retained on the archive’s original scale.
The reported highest/second-highest score pairs are Claude 5/2, Gemini 1/5, GPT-5.5 5/8, and Grok 8/1. These are not standard Enneagram wings, which are adjacent types. Two models share Type 5 as their highest score; the report does not show four different dominant types.
PROFILE
5 / 2
Claude Opus 4.7
Reported highest / second-highest scores; not a standard wing classification.
PROFILE
1 / 5
Gemini 3.1 Pro
Reported highest / second-highest scores; not a standard wing classification.
PROFILE
5 / 8
GPT-5.5
Reported highest / second-highest scores; not a standard wing classification.
PROFILE
8 / 1
Grok 4.3
Reported highest / second-highest scores; not a standard wing classification.