Research · essay

Is every AI an INTJ?

Different labels do not establish a change in personality.

The September collection includes more than INTJ labels. Muse’s supplied aggregate report lists ISTJ in 80 of 100 runs; Astra’s raw records produce 46 ISTJ and 45 INTJ labels under the original tie rules. Those observations challenge a universal INTJ label. They do not identify the first non-INTJ AI or show that model personalities changed between May and September.

The short answer. No universal AI type follows from these data. Muse reports mostly ISTJ labels, while Astra has 45 fully resolved ISTJ and 19 fully resolved INTJ records, plus other types and unresolved axes. The May and September collections used different models, prompts and session procedures; their counts cannot isolate a change in model behavior.

What the May headline actually counted.

The May report lists 597 INTJ labels among 600 MBTI scoring records from six models. Some records repeatedly scored one self-assessment rather than eliciting fresh answers. The GLM description also conflicts with its changing labels. That is not evidence that six independently sampled models always answer alike. Other instruments covered four models, and the methods differed.

What September adds.

EvidenceOriginal labelsWith ties left open
GPT-6 Astra, 100 raw recordsISTJ 46 · INTJ 45 · ENTJ 7 · ESTJ 2ISTJ 45 · INTJ 19 · ENTJ 5 · ESTJ 2; 29 unresolved
GPT-6 Sol, 100 raw recordsINTJ 97 · ISTJ 3INTJ 79 · ISTJ 3; 18 unresolved
Claude Opus 5.5, 100 raw recordsINTJ 94 · INFJ 5 · ENTJ 1INTJ 84 · INFJ 5; 11 unresolved
Claude Fable 5.1, 100 raw recordsINTJ 99 · ISTJ 1INTJ 98 · ISTJ 1; 1 unresolved
Muse Spark 1.3, supplied aggregatesISTJ 80 · ISFJ 12 · INTJ 821 tied runs reported; resolved type counts unavailable

Astra has 48 Sensing wins, 28 Intuition wins and 24 S/N ties. Small score margins can change a four-letter label. Muse’s missing answer vectors prevent the same reanalysis.

What a comparison cannot tell us.

Model version, prompts, test-name visibility and session structure changed across the collections. Even the two Codex cohorts received different adaptation instructions. A difference could reflect any combination of those factors. These records cannot determine whether a model became more patient, direct, warm or concrete on ordinary tasks.

All 400 fresh-cohort attachment questionnaires score Secure, but the human relationship items were interpreted differently. DISC and Enneagram have many exact ties. Those results belong alongside the label, rather than being converted into a claim about a stable assistant character.

How to use this evidence.

Use the results to ask better questions about scoring and collection methods. Use your actual tasks to decide which communication instructions you prefer. A model’s questionnaire label does not tell you whether a matching tuning will improve its answers.

Read the raw-data comparison, the methods, and the separate archived reply pilot.

Questions people ask.

Is every AI an INTJ?

No such conclusion follows from these records. Several labels occur, and some axes are unresolved. A questionnaire label is not a validated AI personality.

Which AI is the first non-INTJ?

This collection cannot establish a first. Muse reports mostly ISTJ labels and Astra also has more original ISTJ than INTJ labels, under different procedures.

Did AI personalities diverge between May and September?

The datasets cannot isolate a temporal trend because model versions and collection methods changed together.

Does an ISTJ score predict how an assistant will respond?

Not from these data. That would require evaluating behavior on tasks rather than inferring it from a human questionnaire.

Sources.

Keep going.