Scientists at the Hebrew University of Jerusalem have demonstrated that large language models such as ChatGPT can be used to generate, validate and even predict responses to personality questionnaires derived from any English source text. The team reported its work in the journal iScience under the title “Generating and analysing personality questionnaires using large language models.”
What the researchers did
The research group — Dr Rotem Monsa, Prof Aviv Zohar and Prof Shahar Arzy from Hebrew University–Hadassah Medical School and the university’s Faculty of Computer Science — developed a method for producing personality-assessment items from text using ChatGPT. They applied the approach to two very different sources: the Diagnostic and Statistical Manual of Mental Disorders, Fifth Edition (DSM‑5), the standard reference for psychiatric diagnoses, and, deliberately as an unconventional counterpoint, an astrology textbook.
The underlying idea rests on a simple observation: because large language models (LLMs) are trained on vast amounts of human-written text from the internet — including books, websites and social media — they may implicitly capture patterns that reflect human personality. The team argued that, “
Since personality traits are reflected in language, LLMs may have learned the structure of human personality as a natural byproduct of their training,”a point attributed to lead author Monsa in reporting by The Jerusalem Post.
Key findings and claims
According to the published paper, the method allowed the researchers to:
- Automatically generate questionnaire items from arbitrary English texts using ChatGPT;
- Validate those items with established psychometric procedures;
- Predict population-level response patterns before conducting empirical surveys.
The report emphasises that LLMs’ training on extensive human language corpora may give them an ability to approximate the relationships between language use and personality traits. In practical terms, the researchers suggest LLMs could speed the creation of assessment tools and provide prior estimates of how groups will respond to items derived from specific texts.
Context and implications
Personality assessment traditionally requires careful item design, piloting and psychometric validation — a process that can be time-consuming and costly. Automating parts of that workflow could increase efficiency for researchers and clinicians, but it also raises important caveats.
- Methodological caution: The study shows capability, not clinical readiness. Automated item generation still needs rigorous empirical validation before use in diagnosis or high‑stakes settings.
- Interpretation limits: Language patterns captured by LLMs reflect the data they were trained on; biases and gaps in training data may influence inferred personality structures.
- Ethical concerns: If LLMs can predict group responses from text, questions arise about privacy, consent and the potential for profiling when models are applied to social media or other personal texts.
The research contributes to ongoing debate about what LLMs “know” about humans. Are they merely statistical pattern-matchers, or have they internalised meaningful structure about human traits? The Hebrew University team presents empirical steps towards answering that question, while the broader community must decide how to apply such capabilities responsibly.
| Element | Details |
|---|---|
| Authors | Dr Rotem Monsa, Prof Aviv Zohar, Prof Shahar Arzy |
| Institution | Hebrew University of Jerusalem (Hadassah Medical School; Faculty of Computer Science) |
| Journal | iScience (Cell Press) |
| Text sources used | DSM‑5; an astrology textbook |
The study is an example of careful experimental work probing what contemporary AI models can do beyond straightforward question‑answering. It offers a promising technique for researchers seeking to build questionnaires quickly, but it also underlines the need for continued scrutiny of model biases, empirical validation of generated instruments, and ethical guidance on application.