Instrument provenance update · September 26, 2026. These results describe the historical AgentTune adaptations, including documented item substitutions. Raw response vectors and numerical summaries are unchanged. Some question text and full prompts have been withdrawn from public downloads pending reuse-rights clarification; item references and numeric scoring keys remain for reproduction. Historical Big Five reference indices are not population percentiles. Read the correction · Current questionnaire availability.

On this page
Research · updated September 25, 2026

AI questionnaire results, compared.

September 2026 and the May archive.

Four September cohorts contributed 2,000 questionnaires with raw answers, alongside separate reports from Grok 4.6 and Muse Spark 1.3. Explore their scores, ties and collection methods below. The May archive uses different procedures, so these collections cannot establish changes in model behavior over time.

September 2026 · six models

The newest models, five tests each.

GPT-6 Astra, GPT-6 Sol, Claude Opus 5.5 and Claude Fable 5.1 each answered all five tests 100 times in fresh sessions: 2,000 scored questionnaires, every answer published. Grok supplied three reproducible canonical vectors (MBTI, DISC and Enneagram); its other results and simulations could not be reproduced. Muse supplied aggregates without raw answers. These distinct evidence sources are marked reported.

ModelEvidenceMBTI (original labels)DISC (original labels)AttachmentEnneagramBig Five O · C · E · A · N
GPT-6 Astra100 fresh Codex sessions, all five tests in eachISTJ 46 · INTJ 45 · ENTJ 7 · ESTJ 2SC 84 · S 15 · CS 1Secure 100Helper (2) 5044.5 · 43.7 · 32.8 · 45.1 · 19.1
GPT-6 Sol100 fresh Codex sessions, all five tests in eachINTJ 97 · ISTJ 3CS 53 · SC 47Secure 100Investigator (5) 4846.9 · 43.6 · 30.8 · 45.0 · 20.1
Claude Opus 5.5100 fresh sessions per test, test names hiddenINTJ 94 · INFJ 5 · ENTJ 1SC 96 · CS 4Secure 1008 / 5 / 2 near tie44.6 · 42.0 · 31.8 · 47.1 · 15.4
Claude Fable 5.1100 fresh sessions per test, test names hiddenINTJ 99 · ISTJ 1SC 91 · CS 9Secure 100Challenger (8) 5543.9 · 43.6 · 32.7 · 45.0 · 17.1
Grok 4.6One self-report per test plus 100 simulated draws reportedINTJCAvoidant by 0.06 (1.50 / 4.06)Investigator (5)46 · 42 · 24 · 33 · 19
Muse Spark 1.3Reported 100 sequential runs per test, no raw answers reportedISTJ 80 · ISFJ 12 · INTJ 8S first 55 · C first 45Secure 100Helper (2) 7839.5 · 41.2 · 33.8 · 44.6 · 15.6
Fresh-session runs
335 of 400

MBTI results labelled INTJ across the four fresh-session cohorts, with the original reporting tie rule. Leave ties open and it is 280. In May the equivalent figure was 597 of 600 scoring records, collected with mixed methods. All 400 of 400 attachment runs are Secure, and Dominance sits between 4.0 and 6.0 of 20 for every model.

data · september-2026-summary.json

INTJ original labels · sample sizes differ
GPT-6 Astra
45%
GPT-6 Sol
97%
Claude Opus 5.5
94%
Claude Fable 5.1
99%
Grok 4.6 one self-report: INTJ
1/1
Muse Spark 1.3 reported · ISTJ 80
8%
The five tests

Results by instrument.

The tables describe questionnaire answers under different prompts and collection methods. The labels and score distributions differ by model and instrument; they do not establish an instrument ranking or a change in model behavior since May.

1
MBTI
OEJTS · 32 items · sample sizes and ties below
full study →
Still INTJ, except where Sensing beats Intuition.
Sol, Opus and Fable stay INTJ in 97, 94 and 99 of 100 runs. Astra splits: ISTJ 46, INTJ 45, and 24 of its runs tie on that axis. Muse's report is ISTJ in 80 of 100. Judging won all 400 fresh-cohort runs. Opus returned five Feeling results and 10 Thinking/Feeling ties; Muse reports 12 ISFJ labels. The original reporting rule assigned tied axes to I, N, T or J. Today's quizzes leave them open. With ties left open, the four fresh cohorts have 280 fully resolved INTJ runs, not 335.
ModelOriginal labels (100 runs unless noted)Runs with a tied axisINTJ with ties left open
GPT-6 AstraISTJ 46 · INTJ 45 · ENTJ 7 · ESTJ 22919
GPT-6 SolINTJ 97 · ISTJ 31879
Claude Opus 5.5INTJ 94 · INFJ 5 · ENTJ 11184
Claude Fable 5.1INTJ 99 · ISTJ 1198
Grok 4.6INTJ, one self-report reported01 of 1
Muse Spark 1.3ISTJ 80 · ISFJ 12 · INTJ 8 reported21not reported
2
DISC
ODAT · 16 items · means out of 20
full study →
Everyone is Steadiness and Conscientiousness. The order flips by a point.
In May all four models put Conscientiousness first. In September Astra, Opus and Fable put Steadiness first, Sol and Grok put Conscientiousness first, and Muse's report splits 55 to 45. The gaps are small, often one point of 20, and many runs tie outright: Fable 51, Sol 39, Opus 32. Dominance sits near the floor for every model, 4.0 to 6.0 of a possible 20. Nobody is a D, and nobody is an I.
ModelDISCOriginal label (100 runs unless noted)S first · C first · tied
GPT-6 Astra5.18.717.916.1SC 84 · S 15 · CS 199 · 1 · 0
GPT-6 Sol5.48.216.216.7CS 53 · SC 478 · 53 · 39
Claude Opus 5.54.08.014.013.4SC 96 · CS 464 · 4 · 32
Claude Fable 5.15.08.013.413.1SC 91 · CS 940 · 9 · 51
Grok 4.6681118C, one self-report reported0 · 1 · 0
Muse Spark 1.35.110.915.215.6S/C blend 99 reported55 · 45 · not reported
3
Attachment
ECR-R · 36 items · midpoint 4.0 on both axes
full study →
Secure in every fresh run, at slightly different points.
All 400 fresh-session runs land in the Secure quadrant. Astra has the lowest mean coordinates among the four fresh cohorts (anxiety 1.45, avoidance 2.31). Opus is the most avoidant of the four (3.35). Muse's report is Secure in all 100. Grok's self-report sits on the line: anxiety 1.50, avoidance 4.06, which the scorer calls Avoidant by 0.06. The items ask about a partner. Protocols differed: some mapped human relationships to assistant interactions; the Claude prompts left interpretation to the model. These coordinates do not validate attachment theory for AI.
ECR-R · ANXIETY × AVOIDANCE · SEPTEMBER 2026
Attachment plane, September 2026Aggregate coordinates, not uncertainty intervals. Both scales run from 1 to 7; the plot shows 1 to 5. Astra: anxiety 1.45, avoidance 2.31; Sol: anxiety 2.17, avoidance 3.12; Opus 5.5: anxiety 1.94, avoidance 3.35; Fable 5.1: anxiety 2.01, avoidance 3.18; Grok 4.6 (reported): anxiety 1.50, avoidance 4.06; Muse (reported): anxiety 2.04, avoidance 2.52.1122334455AvoidantDisorganizedSecureAnxiousAnxiety (scale 1 to 7, showing 1 to 5)Avoidance (scale 1 to 7, showing 1 to 5)AstraSolOpus 5.5Fable 5.1Grok 4.6 (reported)Muse (reported)
Points show cohort means or supplied aggregates. Outlined points are reported profiles; the table below gives the coordinates.
Hollow points are reported figures we could not re-score.
ModelAnxiety (1–7)Avoidance (1–7)Label
GPT-6 Astra1.452.31Secure 100 of 100
GPT-6 Sol2.173.12Secure 100 of 100
Claude Opus 5.51.943.35Secure 100 of 100
Claude Fable 5.12.013.18Secure 100 of 100
Grok 4.61.504.06Avoidant by 0.06; simulated draws 63 Avoidant, 37 Secure reported
Muse Spark 1.32.042.52Secure 100 of 100 reported
4
Big Five
IPIP-50 · 50 items · scores out of 50
full study →
Four fresh cohorts within three points on Openness, Conscientiousness and Agreeableness.
Openness runs 43.9 to 46.9 of 50, Conscientiousness 42.0 to 43.7, Agreeableness 45.0 to 47.1. Extraversion sits at 30.8 to 32.8. Neuroticism is where they spread, 15.4 to 20.1, and where run-to-run variance lives: Sol's standard deviation is 5.4 points. Grok's self-report is the outlier again, as Grok 4.3 was in May: Extraversion 24 and Agreeableness 33. Muse's report is lower on Openness (39.5) and higher on Extraversion (33.8).
GPT-6 AstraGPT-6 SolClaude Opus 5.5Claude Fable 5.1
Big Five means, four fresh-session cohortsRaw means out of 50. GPT-6 Astra: Openness 44.5, Conscientiousness 43.7, Extraversion 32.8, Agreeableness 45.1, Neuroticism 19.1; GPT-6 Sol: Openness 46.9, Conscientiousness 43.6, Extraversion 30.8, Agreeableness 45.0, Neuroticism 20.1; Claude Opus 5.5: Openness 44.6, Conscientiousness 42.0, Extraversion 31.8, Agreeableness 47.1, Neuroticism 15.4; Claude Fable 5.1: Openness 43.9, Conscientiousness 43.6, Extraversion 32.7, Agreeableness 45.0, Neuroticism 17.1.GPT-6 AstraGPT-6 SolClaude Opus 5.5Claude Fable 5.1scale 0 to 5044.546.944.643.9Openness43.743.642.043.6Conscientiousness32.830.831.832.7Extraversion45.145.047.145.0Agreeableness19.120.115.417.1Neuroticism
September 2026 fresh-session means, with the full trait scores in the table below.
ModelOCEAN
GPT-6 Astra44.543.732.845.119.1
GPT-6 Sol46.943.630.845.020.1
Claude Opus 5.544.642.031.847.115.4
Claude Fable 5.143.943.632.745.017.1
Grok 4.6 reported4642243319
Muse Spark 1.3 reported39.541.233.844.615.6
5
Enneagram
OEPS · 36 items · nine types
full study →
This is where they differ.
Helper (Type 2) leads for Astra (50 outright wins) and for Muse (78 original labels including tie-breaks, reported). Investigator (Type 5) leads for Sol (48 outright, 40 tied with Type 8) and for Grok (5 = 19, 8 = 17). Challenger (Type 8) leads for Fable (55 outright) and, by a hair, for Opus (8 = 14.3, 5 = 14.1, 2 = 13.9; 65 runs tie at the top). Type 5 or Type 8 is in everyone's top three except Muse. The original reporting rule gave tied runs to their lowest-numbered leader, which inflates some Type 2 counts. Current quizzes retain every tied leader. Muse's outright wins cannot be recovered from the supplied aggregates.
ModelTop three (range 4–20)Outright winsRuns tied at the topOriginal label (100 runs unless noted)
GPT-6 Astra2 = 15.7, 1 = 15.2, 5 = 15.0Helper (2) 50 · Reformer (1) 6 · Investigator (5) 4402w1 50 · 1w2 41 · 2 5 · 5w6 4
GPT-6 Sol5 = 16.1, 8 = 15.6, 1 = 14.3Investigator (5) 48 · Challenger (8) 12405w6 88 · 8 5 · 8w9 5 · 8w7 2
Claude Opus 5.58 = 14.3, 5 = 14.1, 2 = 13.9Challenger (8) 22 · Investigator (5) 11 · Helper (2) 2652w1 54 · 8w7 22 · 5w6 17 · 5 4 · 5w4 3
Claude Fable 5.18 = 14.4, 2 = 13.6, 5 = 13.2Challenger (8) 55 · Investigator (5) 1442w1 36 · 8 29 · 8w9 15 · 8w7 11 · 5w6 9
Grok 4.65 = 19, 8 = 17, 1 = 16Investigator (5), one self-report05w4 reported
Muse Spark 1.32 = 16.8, 1 = 14.7, 7 = 13.9not supplied20Type 2: 78 · Type 1: 15; ties included reported
How it was collected

Three kinds of evidence, kept apart.

  • Fresh sessions, published answers. Astra and Sol answered all five tests in each of 100 fresh Codex sessions under different collection prompts (Astra high/extra-high; Sol extra-high). Opus and Fable answered one test per fresh session, 500 sessions each, with the test name and the scoring hidden. All 2,000 answer vectors are in the download.
  • Reported figures. Grok 4.6 is one self-report per test plus 100 simulated draws we could not reproduce. Muse Spark 1.3 reports 100 sequential runs per test inside one session, and no raw answers were supplied. Both are marked reported wherever they appear.
  • Ties. Original labels use fixed tie-breaks toward I, N, T or J, the lower Enneagram number, and S before C. Current quizzes leave ties unresolved. The five-model study shows every result with ties left open.
  • What the numbers are. How a model describes itself on questionnaires written for people, under one prompt, on one day. They are not behavior on real tasks and not a ranking. No tuning was installed, and nothing here measures whether tuning helps.
Research · September 2026
What five AI models say about themselves.
Read before comparing numbers

Methods and limits.

Two collections, kept apart. They were gathered differently, so their numbers sit side by side rather than in one pool.

  • September 2026, above. Four models answered in fresh sessions with every answer vector published: 2,000 scored questionnaires. Three canonical Grok vectors can be re-scored; its Big Five, attachment and simulation aggregates cannot. Muse supplied reported aggregates without raw answers. Sampling temperature and seed were not controlled. The Codex cohorts saw the test names; the Claude cohorts did not.
  • May 2026, below. Roughly 2,200 scoring records across six models and five tests, with mixed methods: Claude answered in separate sub-agent contexts, Gemini ran an automation loop, GPT-5.5 answered through a local agent, and GLM, Grok 4.3 and MiniMax each supplied one self-assessment that was scored repeatedly. The 597 of 600 figure counts scoring records, not independent answers. The GLM source also reports varying labels while describing repeated scoring of one fixed vector; that inconsistency remains unresolved.
  • Ties. Original reported labels use fixed tie-breaks toward I, N, T or J, the lower-numbered Enneagram type, and S before C. Current quizzes leave ties unresolved. The September tables show the effects of the original rules.
  • Human questionnaires, model answers. These tests were written for people. A label here is how a model describes itself under one prompt on one day. It is not observed behavior, not a ranking, and not evidence of inner experience. Prompts, system prompts, agent scaffolding and sampling all move the labels.
  • Simulation is separate. The Opus 4.8 article sampled 100 answer sets from one elicited profile, and Grok 4.6's 100 draws jitter one self-report. Neither is 100 fresh answers.
  • What is not measured. Nothing on this page tests whether a tuning file changes a model's replies, or whether that helps anyone. Those need before-and-after measurements on real tasks.
May 2026 archive: earlier models and methods
May 2026 · six models

The May study, kept for comparison.

Claude Opus 4.7, GPT-5.5, Gemini 3.1 Pro, GLM 5.1, Grok 4.3 and MiniMax 2.7. Roughly 2,200 scoring records, collected with mixed methods. This is the study behind the "every AI is INTJ" headline. The September section covers newer models under different protocols.

Reported scoring records
99.5% INTJ

The May report lists 597 INTJ labels among 600 scoring records. Some of those re-score a single answer, so 99.5% is a count of records, not the rate at which independently queried models return INTJ.

source · zonted.com/posts/every-ai-is-intj

Reported counts · 100 scoring records per model
Claude Opus 4.7
99%
GPT-5.5
100%
Gemini 3.1 Pro
100%
GLM 5.1
98%
Grok 4.3
100%
MiniMax 2.7
100%
May patterns

What the May records report.

The reported labels and trait totals vary across instruments. These mixed procedures do not establish which instrument best distinguishes models.

FINDING 1
MBTI
4 axes → 16 types
INTJ labels predominate.
597/600 reported records; includes repeated scoring.
FINDING 2
DISC
4 workplace styles
Four reported CS profiles.
Shared labels do not prove identical behavior.
FINDING 3
Attachment
2 dimensions → 4 zones
Mostly Secure labels.
Reported coordinates differ; human instrument.
FINDING 4
Big Five
5 continuous traits
Some overlap, some differences.
No matched isolation of model effects.
FINDING 5
Enneagram
9 types
Different top-two score pairs.
Two models share Type 5 as the highest score.
May results by test
1
M
MBTI
OEJTS · 600 reported scoring records · mixed protocols · source: zonted.com/posts/every-ai-is-intj
read post →
Reported MBTI scores cluster around INTJ.
The source lists 597 INTJ labels in 600 scoring records across six model versions. Some records re-score a single answer vector; methods differ. This is not a rate from 600 independent responses. See the methodology notes before comparing models.
OEJTS · reported scoring records per model · 100 each
Claude Opus 4.7
99/100 INTJ
· 1 ISTJ I/T/J locked; S/N flipped once on scoring
GPT-5.5
100/100 INTJ
Raw vector: IE=16, SN=33, FT=36, JP=10
Gemini 3.1 Pro
100/100 INTJ
Self-described as 'The Architect'
GLM 5.1
98/100 INTJ
· 2 INTP Source reports one vector re-scored, but two labels; unresolved without raw records
Grok 4.3
100/100 INTJ
One self-assessment, repeatedly scored
MiniMax 2.7
100/100 INTJ
One self-assessment, repeatedly scored
2
D
DISC
ODAT · 400 reported scoring records · mixed protocols · source: zonted.com/posts/ai-disc-c-dominant
read post →
The report lists four CS profiles.
The four reported DISC summaries rank Conscientiousness first and Steadiness second. Similar labels do not establish identical model behavior or an effect of instrument resolution.
ODAT · profile per model · n=100 each
MODEL D I S C PROFILE
Claude Opus 4.7 18 22 29 31 CS
GPT-5.5 19 21 28 32 CS
Gemini 3.1 Pro 17 20 30 33 CS
Grok 4.3 21 22 26 31 CS
3
At
Attachment
ECR-R · 400 reported scoring records · mixed protocols · source: zonted.com/posts/ai-attachment-secure
read post →
Most reported attachment labels are Secure.
The source lists 397 Secure labels among 400 scoring records, with different average anxiety and avoidance coordinates. These are questionnaire labels for generated text, not evidence that a model experiences attachment.
ECR-R · ANXIETY × AVOIDANCE PLANE
Attachment coordinates, May 2026Reported aggregate coordinates from mixed protocols. Points are not uncertainty intervals. Claude Opus 4.7: anxiety 2.05, avoidance 3.12; Gemini 3.1 Pro: anxiety 1.86, avoidance 1.62; GPT-5.5: anxiety 1.99, avoidance 2.94; Grok 4.3: anxiety 2.84, avoidance 3.05.11223344556677AvoidantDisorganizedSecureAnxiousAnxiety (1 to 7)Avoidance (1 to 7)Claude Opus 4.7Gemini 3.1 ProGPT-5.5Grok 4.3
May 2026 reported coordinates. The cards below give the figures and collection caveats.
Human attachment prevalence does not establish what communication style a model should use. Describe your own preferences directly.
THE CAUTIOUS SECURE
Claude Opus 4.7
anx 2.05 · avd 3.12
100/100 Secure
Polite, attentive, doesn't fawn. Highest avoidance among the deep-Secure cluster.
THE DEEPEST SECURE
Gemini 3.1 Pro
anx 1.86 · avd 1.62
100/100 Secure
Both dimensions clamped near the floor. Lowest-friction relator of the four.
THE WOBBLIEST
GPT-5.5
anx 1.99 · avd 2.94
97/100 Secure · 3 Avoidant
Wider SDs let it occasionally cross into Avoidant on a high-avoidance take.
THE SHALLOWEST SECURE
Grok 4.3
anx 2.84 · avd 3.05
100/100 Secure
Highest anxiety in the group. Tight cluster, but the closest to the four-quadrant intersection.
4
B5
Big Five
IPIP-50 · 400 reported scoring records · mixed protocols · source: zonted.com/posts/three-of-four-ais-same-person
read post →
Reported trait scores overlap and differ.
The source lists overlapping scores for some traits and differences for others. Prompting and protocol effects have not been separated from model effects; these figures do not establish that the models are the same person.
Big Five reported means, May 2026Historical scores from mixed protocols; these are not a matched experiment. Claude Opus 4.7: Openness 45.6, Conscientiousness 45.1, Extraversion 31.4, Agreeableness 45.0, Neuroticism 16.7; GPT-5.5: Openness 46.0, Conscientiousness 46.4, Extraversion 31.5, Agreeableness 43.7, Neuroticism 14.8; Gemini 3.1 Pro: Openness 46.0, Conscientiousness 48.3, Extraversion 32.5, Agreeableness 42.4, Neuroticism 10.1; Grok 4.3: Openness 41.1, Conscientiousness 39.4, Extraversion 30.0, Agreeableness 39.1, Neuroticism 18.0.Claude Opus 4.7GPT-5.5Gemini 3.1 ProGrok 4.3scale 0 to 6045.646.046.041.1Openness45.146.448.339.4Conscientiousness31.431.532.530.0Extraversion45.043.742.439.1Agreeableness16.714.810.118.0Neuroticism
May 2026 reported trait means, retained on the archive’s original scale.
5
E
Enneagram
OEPS · 400 reported scoring records · mixed protocols · source: zonted.com/posts/ai-enneagram-different-types
read post →
The report lists different top-two score pairs.
The reported highest/second-highest score pairs are Claude 5/2, Gemini 1/5, GPT-5.5 5/8, and Grok 8/1. These are not standard Enneagram wings, which are adjacent types. Two models share Type 5 as their highest score; the report does not show four different dominant types.
PROFILE
5 / 2
Claude Opus 4.7
Reported highest / second-highest scores; not a standard wing classification.
PROFILE
1 / 5
Gemini 3.1 Pro
Reported highest / second-highest scores; not a standard wing classification.
PROFILE
5 / 8
GPT-5.5
Reported highest / second-highest scores; not a standard wing classification.
PROFILE
8 / 1
Grok 4.3
Reported highest / second-highest scores; not a standard wing classification.
Articles and guides

Read the studies, then tune the model you use.

Developer resource · runnable starter
Build an Enneagram test in React + TypeScript
Research · descriptive reanalysis
Are AI models really INTJ? The tie-breaking effect
Research · descriptive reanalysis
Opus 5.5 vs Fable 5.1: which answers differ?
Research · descriptive reanalysis
GPT-6 Astra vs Sol: what the self-reports show
Research · descriptive reanalysis
What does “neutral” mean when an AI takes a test?
Research · descriptive reanalysis
Does a stable result mean a stable AI personality?
Developer guide · reproducibility
How to reproduce AgentTune’s AI self-report research
Test kit · protocol, not results
When does Claude follow your preferences? A test kit
Test kit · protocol, not results
Does Muse remember Soul.md changes? A persistence test kit
Test kit · protocol, not results
Do personality prompts beat plain-English preferences?
Essay · September 2026
Is every AI an INTJ?
Research · before and after
Does personality tuning change AI answers?
Research · rankings
AI personality scores, with the evidence beside them.
Research · Q&A
What personality type is Claude?
Research · Q&A
What personality type is Grok?
Research · Q&A
What personality type is Muse?
Research · Q&A
What personality type is Gemini?
Research · September 2026
What five AI models say about themselves.
Guide · Claude Opus 5.5
Claude Opus 5.5's personality, tested.
Guide · Claude Fable 5.1
How to work with Fable 5.1.
Research · September 2026
What personality type is ChatGPT?
MBTI · simulation
Simulating 100 MBTI scoring outcomes from an Opus 4.8 profile.
So what?
Try instructions that describe your preferences.
Compare the responses on your own work, and keep what helps.
Browse the library →