Is every AI an INTJ?
Different labels do not establish a change in personality.
By Bernard Huang · Updated
The September collection includes more than INTJ labels. Muse’s supplied aggregate report lists ISTJ in 80 of 100 runs; Astra’s raw records produce 46 ISTJ and 45 INTJ labels under the original tie rules. Those observations challenge a universal INTJ label. They do not identify the first non-INTJ AI or show that model personalities changed between May and September.
What the May headline actually counted.
The May report lists 597 INTJ labels among 600 MBTI scoring records from six models. Some records repeatedly scored one self-assessment rather than eliciting fresh answers. The GLM description also conflicts with its changing labels. That is not evidence that six independently sampled models always answer alike. Other instruments covered four models, and the methods differed.
What September adds.
| Evidence | Original labels | With ties left open |
|---|---|---|
| GPT-6 Astra, 100 raw records | ISTJ 46 · INTJ 45 · ENTJ 7 · ESTJ 2 | ISTJ 45 · INTJ 19 · ENTJ 5 · ESTJ 2; 29 unresolved |
| GPT-6 Sol, 100 raw records | INTJ 97 · ISTJ 3 | INTJ 79 · ISTJ 3; 18 unresolved |
| Claude Opus 5.5, 100 raw records | INTJ 94 · INFJ 5 · ENTJ 1 | INTJ 84 · INFJ 5; 11 unresolved |
| Claude Fable 5.1, 100 raw records | INTJ 99 · ISTJ 1 | INTJ 98 · ISTJ 1; 1 unresolved |
| Muse Spark 1.3, supplied aggregates | ISTJ 80 · ISFJ 12 · INTJ 8 | 21 tied runs reported; resolved type counts unavailable |
Astra has 48 Sensing wins, 28 Intuition wins and 24 S/N ties. Small score margins can change a four-letter label. Muse’s missing answer vectors prevent the same reanalysis.
What a comparison cannot tell us.
Model version, prompts, test-name visibility and session structure changed across the collections. Even the two Codex cohorts received different adaptation instructions. A difference could reflect any combination of those factors. These records cannot determine whether a model became more patient, direct, warm or concrete on ordinary tasks.
All 400 fresh-cohort attachment questionnaires score Secure, but the human relationship items were interpreted differently. DISC and Enneagram have many exact ties. Those results belong alongside the label, rather than being converted into a claim about a stable assistant character.
How to use this evidence.
Use the results to ask better questions about scoring and collection methods. Use your actual tasks to decide which communication instructions you prefer. A model’s questionnaire label does not tell you whether a matching tuning will improve its answers.
Read the raw-data comparison, the methods, and the separate archived reply pilot.
Questions people ask.
Is every AI an INTJ?
No such conclusion follows from these records. Several labels occur, and some axes are unresolved. A questionnaire label is not a validated AI personality.
Which AI is the first non-INTJ?
This collection cannot establish a first. Muse reports mostly ISTJ labels and Astra also has more original ISTJ than INTJ labels, under different procedures.
Did AI personalities diverge between May and September?
The datasets cannot isolate a temporal trend because model versions and collection methods changed together.
Does an ISTJ score predict how an assistant will respond?
Not from these data. That would require evaluating behavior on tasks rather than inferring it from a human questionnaire.