// HACKER NEWS — CYBERSECURITY
The shrinking landscape of linguistic diversity in the age of LLMs
Nature Human Behaviour
(2026) Cite this article
Language is far more than a communication tool; it encodes a wealth of information about a person’s identity, psychological state and social context, providing valuable insights for diverse fields including psychology, marketing and healthcare. Across three studies spanning seven datasets in different domains and over 880,000 texts, we show that the widespread adoption of large language models (LLMs) as writing assistants is linked to declines in linguistic diversity, interfering with the societal and psychological insights language provides. While core content is retained when LLMs polish and rewrite texts, LLMs also homogenize writing styles, reducing writing-complexity variance by a statistically significant 21–50% across datasets and models (P ≤ 0.05), and amplify patterns associated with dominant characteristics while suppressing others, emphasizing conformity over individuality. These trends hold across different LLMs, prompts and contexts, with potential implications for diagnostic processes, personalization efforts, hiring assessments and cultural preservation.
This is a preview of subscription content, access via your institution
Access Nature and 54 other Nature Portfolio journals
Get Nature+, our best-value online-access subscription
Receive 12 digital issues and online access to articles
Prices may be subject to local taxes which are calculated during checkout
Of the seven datasets analysed in this Article, six are publicly available and one is access restricted. Publicly available datasets: Reddit r/WritingPrompts posts were obtained from Project Arctic Shift (https://github.com/ArthurHeitmann/arctic_shift), which mirrors publicly accessible Reddit content (excluding private subreddits and content for which the original poster has filed a removal request) and is itself a successor archive to the Pushshift Reddit dataset144. arXiv abstracts were obtained from the publicly released metadata collection92, distributed under a Creative Commons CC0 1.0 Universal Public Domain Dedication. Patch News articles were accessed via the public Patch.com archive. United States Congressional Records floor speeches were obtained from the publicly released parsed corpus119, with the underlying speeches in the public domain as works of the US federal government. YourMorals Facebook posts and Moral Foundations Questionnaire responses are publicly released by the original investigators20; the data were collected from participants who voluntarily completed self-report measures on yourmorals.org and consented at registration to having their Facebook posts accessed for research purposes, under University of Southern California Institutional Review Board protocol UP-07-00393-AM019 (M. Dehghani, PI). Empathic Conversations are publicly released by the original investigators128; the dataset was collected under University of Pennsylvania IRB protocol #826448 from Amazon Mechanical Turk workers who were informed that their conversations, demographic information, personality surveys and essays would be used for academic research and distributed externally for research purposes, and were compensated per HIT. Access-restricted dataset: The Essays corpus124 is not publicly redistributable and is governed by access conditions imposed by the dataset owner. The corpus can be requested directly from Prof. James W. Pennebaker (University of Texas at Austin; pennebaker@utexas.edu or jwpennebaker@gmail.com); the original data were collected at UT-Austin from undergraduate psychology students who participated for course credit, with informed consent and institutional research approval reported in the original paper. Access is granted at the discretion of the dataset owner and does not require additional ethics approval from the requester beyond standard institutional review. We obtained the corpus via direct request and were not permitted to redistribute it. Use of Reddit