Anderson의 관점
AI 모델의 ‘창의적’ 출력이 제공업체 간에 유사해지고 있다

새로운 연구에 따르면 경쟁 업체의 AI 모델들이 창의적인 답변에서 점점 더 비슷해지고 있어, 사용자가 접하게 되는 아이디어의 다양성이 좁혀질 가능성이 있습니다.
미국에서 진행된 최신 연구는 주요 AI 모델들의 ‘창의적’ 출력이 시간이 지남에 따라 점점 더 유사해지고 있음을 발견했습니다.
듀크 대학교 소속 연구진은 2023년부터 2026년까지 출시된 12개 제공업체의 69개 모델을 테스트했으며, 실제 세계의 개방형 질문과 표준 창의성 테스트 모두에서 출력 다양성이 통계적으로 유의하게 감소했음을 확인했습니다. 이는 주요 LLM 제공업체 전반에 걸쳐 ‘창의적’ 요청에 대한 답변이 점점 더 비슷해질 수 있음을 시사합니다.
테스트된 제공업체 그룹은 Anthropic; Cohere; DeepSeek; Google; Meta; MiniMax; Mistral AI; Moonshot AI; OpenAI; Qwen; xAI; 그리고 Z.ai*:

새 논문에서 사용된 모델 출시 일정(2023년 3월부터 2026년 7월까지). 이 차트는 12개 AI 제공업체의 69개 모델 버전을 포함하며, 3년 기간 동안 테스트된 모델이 어떻게 배포되었는지와 일부 제공업체의 출시 빈도가 점점 증가하고 있음을 보여줍니다. 출처 – https://arxiv.org/pdf/2608.19437 Source
The authors of the new paper state**:
‘We find that LLM responses to the open-ended prompts we [test] have become increasingly similar over time.
‘This suggests that LLMs are becoming less creative in tasks that involve generating open-ended responses, demanding scrutiny of their long-term usefulness as creative assistants.’
Although ‘algorithmic monoculture‘ has become an established line of study, an examination of the trends toward homogeneity in AI-generated creative outputs has not been undertaken until now, the authors assert.
하지만, 고립된 사례들은 이미 이러한 수렴 현상을 가리켜 왔으며, 그 중 하나는 Cornell University의 20026년 5월 연구에서 발견된 특이한 경우입니다. 이 연구에서는 다양한 LLM 제공업체가 ‘등대지기’와 같은 특정 고전적 이름(예: ‘Mara’, ‘Elias’)에 대해 이상하게도 충돌하는 집착을 보였습니다.

5월에 Cornell University의 새 논문은 ‘개방형 프롬프트’를 제공했을 때 LLM 제공업체들 사이에 이상한 유사성을 발견했지만, 이러한 현상의 원인에 대한 단서는 아직 다양하게 존재하는 훈련 데이터에서 명확히 드러나지 않았습니다.
새 연구는 보다 체계적입니다: 연구진은 실제 세계의 창의적 프롬프트와 일상 물건의 비정상적인 사용을 묻는 고전 심리학 테스트를 모두 적용해 3년간의 모델을 비교했습니다. 그런 다음 모델 답변 간 의미적 거리를 측정해 세대가 거듭될수록 그 거리가 줄어드는지를 추적했습니다.
The authors of the new work, titled Are LLMs becoming similarly creative? Evidence from three years of models, state†:
‘Our findings, though preliminary, raise concerns about the long-term usefulness of LLMs as creative partners. Even if models perform well on creative tasks, converging outputs could bound the range of possibilities LLM users are exposed to, and with it, the breadth of their own thinking.
‘If using an LLM for creative tasks like essay writing decreases one’s brain activity, could using increasingly less creative LLMs–the trend suggested by our study–further worsen LLMs’ effects on human creativity, as observed by this and other studies?’
Already, the diverse characteristics of LLM text output have become fodder for memes; and, as we reported in May this year, at least one author has already self-published a book apparently featuring the aforementioned ‘lighthouse’ fixation common to major models.
따라서 출판사들이 AI가 썼다고 의심되는 책을 점점 더 회수하고 있는 현 상황에서, 새로운 연구는 인간이 쓴 것으로 보이는 작품—특히 학생들이 제출하는 학술 논문—에서 ‘플롯 우연’이 증가할 가능성을 시사합니다.
As for where this convergence is coming from, the authors hypothesize that overlapping training data and increasingly similar optimization objectives may be pushing models toward comparable internal representations of concepts and semantic relationships – ultimately producing more similar answers†:
‘[It] remains unclear whether this homogeneity is a temporary byproduct of a still-developing technology or an inevitable–potentially compounding–feature of statistical language models.
‘Plausible forces point in both directions. A growing body of academic work suggests that generative models trained on overlapping data and optimized toward similar objectives will organize concepts and semantic relationships in increasingly similar ways, resulting in similar outputs.’
However, the authors also cite work indicating that as models evolve, they may diverge again into their own ring-fenced set of characteristics, in regard to creative output.
Method
To track whether AI models are becoming more alike, the researchers began with open-ended questions, curating answers from successive generations of models. Each answer was converted into a numerical representation of its meaning, making it possible to measure how similar or different the responses were.
Regression analysis was then used to establish whether these differences were shrinking as newer models were released:

AI 출력이 시간이 지남에 따라 더 유사해지는지를 측정하기 위한 저자들의 스키마. 개방형 창의적 프롬프트를 다양한 모델군에 적용하고, 답변을 의미에 따라 매핑해 거리(유사도)를 측정하며, 회귀 분석을 통해 세대가 바뀔 때마다 그 거리가 어떻게 변하는지 추적합니다.
Two sets of prompts were selected to test creativity from different angles. First, the Alternate Uses Task (AUT), a standard psychology test of divergent thinking, asks for unconventional uses for ordinary objects, with the study using objects such as a book, shoe and hammer. Its fixed format provided a controlled way to compare the ideas produced by different models.
A further 100 prompts were distilled from Infinity-Chat100, a collection derived from real-world conversations with language models. These covered less-constrained creative tasks involving content generation, problem-solving, brainstorming and ideation; tasks that would permit greater freedom in regard to the answers produced and the approaches taken.
The complete AUT and Infinity-Chat100 prompt sets were then sent to models released between 2023년 3월 and 2026년 7월, covering 12 providers and systems from Anthropic, Google, Meta, OpenAI, DeepSeek and Mistral AI. All responses were collected through OpenRouter’s API under the same sampling settings, with temperature (freedom to respond creatively) and top-p both set to one.
The models were arranged by release month to examine whether newer generations were producing more similar answers.
Since release dates don’t actually capture changes in training data, architecture or post-training, the authors treated the resulting trend as an association with model development, rather than evidence that the passage of time itself causes convergence.
The models’ answers were then converted into embeddings using the all-MiniLM-L6-v2 sentence-transformer. Each response was represented as a numerical vector, allowing answers with similar meanings to be placed in similar directions, and their semantic distance to be measured.
For the AUT, all uses suggested for each object were treated as a single response, producing ten embeddings per model. For Infinity-Chat100, each answer was embedded separately, producing 100 embeddings per model.
The 27 release months were grouped into nine time-periods, and only models from different providers were compared. Semantic distances between their answers were measured within the scope of each period, with smaller distances indicating greater similarity.
To prevent providers with more models from dominating the results, the analysis was repeated 1,000 times, using balanced samples from each provider.
Results
Linear regression was then used to track these distances over time, with a consistent negative slope indicating that models from different providers were becoming more similar:

초기 테스트 결과: 서로 다른 AI 제공업체의 창의적 응답이 시간이 지남에 따라 더 유사해지는지를 보여줍니다. ‘Alternate Uses Task’와 ‘Infinity-Chat100’ 프롬프트 모두에서 의미 거리(semantic distance)가 감소했으며, 반복된 균형 샘플링에서도 일관된 음의 추세가 나타나 세대가 거듭될수록 유사성이 증가함을 나타냅니다.
Of these results the authors emphasize:
‘For both the AUT and Infinity-Chat response sets, we observe a decline in cross-provider output distances over the observation period.
‘This suggest decreasing diversity–or increasing homogeneity–of LLM creative outputs over time.’
The difference between older and newer models was clearest in the Alternate Uses Task. The average distance between answers from different providers fell from about 0.50 for the earliest models to below 0.40 for the newest, meaning that their answers became substantially more alike. The same pattern was found with Infinity-Chat100, though the change was smaller, with average distance falling from about 0.34 to just above 0.32.

다양한 AI 제공업체의 답변이 시간이 지남에 따라 얼마나 빠르게 유사해지는지를 측정한 테스트 결과. ‘Alternate Uses’ 과제는 Infinity-Chat보다 모델 간 차이가 훨씬 크게 감소했으며, 두 결과 모두 연구진이 반복 샘플링을 수행한 전 과정에서 일관되었습니다.
The finding also survived all 1,000 rounds of balanced resampling. In every case, newer models produced answers that were closer together than those from older models. The size of these declines, together with the researchers’ 95% confidence intervals, are depicted above.
Regarding this, the paper states:
‘The magnitude of the AUT decline is particularly notable given the task’s purpose. The AUT explicitly tests divergent thinking by instructing models to produce uses that are as original and unexpected as possible, making it precisely the setting in which outputs would be expected to differ.’
The Infinity-Chat results show that the same trend also appeared across a much wider range of creative tasks. Its 100 prompts asked models to produce many different kinds of open-ended answers, and these answers became more similar over time.
The change was smaller than in the Alternate Uses Task, but became clearer among models released from 2025 onward. The authors suggest that this could be an early sign that creative outputs from different AI providers are becoming more alike generally, rather than only on one particular type of test.
Conclusion
Issues around LLM and VLM convergence are an ongoing and growing source of concern in both the research community and the consumers of the downstream models that issue from them. We are coming to the very end of the first and only ‘pure’ generation of data that the research scene will ever have access to; and the determination of model providers to exfiltrate the data of rivals inevitably risks convergence, since the data is becoming identical, and principles of model training are similar, if not identical, among providers.
Therefore, nothing quite as simple as replacing em-dashes is likely to emerge to combat the growing similarity of creative output among the model families, and cases of duplication are, instead, likely to come to light in rather more public and embarrassing ways.
* The complete model list is Claude-3-Haiku; Claude-Fable-5; Claude-Opus-4; Claude-Opus-4.1; Claude-Opus-4.5; Claude-Opus-4.6; Claude-Opus-4.7; Claude-Opus-4.8; Command-A; Command-R-08-2024; DeepSeek-Chat; DeepSeek-V3.1-Terminus; DeepSeek-V3.2; DeepSeek-V4-Pro; Gemini-2.5-Pro; Gemini-3-Flash-Preview; Gemini-3.1-Pro-Preview; Gemini-3.5-Flash; Llama-3.1-70B-Instruct; Llama-3.2-3B-Instruct; Llama-3.3-70B-Instruct; Llama-4-Maverick; MiniMax-01; MiniMax-M1; MiniMax-M2; MiniMax-M2.1; MiniMax-M2.5; MiniMax-M2.7; MiniMax-M3; Mistral-Large; Mistral-Large-2407; Mistral-Large-2512; Mistral-Medium-3; Mistral-Medium-3.5; Mistral-Medium-3.1; Mistral-Small-24B-Instruct-2501; Mistral-Small-2603; Mistral-Small-3.1-24B-Instruct; Mistral-Small-3.2-24B-Instruct; Mixtral-8x22B-Instruct; Kimi-K2; Kimi-K2-0905; Kimi-K2.5; Kimi-K2.6; GPT-3.5-Turbo; GPT-4; GPT-4.1; GPT-4o; GPT-5; GPT-5.1; GPT-5.2; GPT-5.3-Chat; GPT-5.4; GPT-5.5; GPT-5.6-Sol; Qwen-2.5-72B-Instruct; Qwen3-Max; Qwen3.5-Plus-20260420; Qwen3.6-Max-Preview; Qwen3.7-Max; Grok-4.20; Grok-4.3; Grok-4.5; GLM-4.5; GLM-4.6; GLM-4.7; GLM-5; GLM-5.1; and GLM-5.2
** Authors’ emphases, not mine.
† My conversion of the authors’ inline citations to hyperlinks.
First published Friday, 2026년 8월 21일 into Korean Translate to Korean. Write idiomatic, publication-quality native Korean for a professional web audience. Preserve the source meaning and facts faithfully without copying English syntax or translating word-for-word. Use contemporary spelling, grammar, punctuation, agreement, and established subject-matter terminology. Avoid false friends, awkward calques, hybrid words, obsolete diacritics, and unnecessary untranslated English. Preserve only genuine names, brands, products, URLs, code, and exact shortcode tokens. Before returning the response, silently proofread the entire translation for fluency, terminology consistency, missing text, duplicated text, and residual source-language prose. Use natural professional Korean editorial style and consistent spacing. while keeping a professional,natural,and SEO-optimized editorial tone.
Do NOT add new information,remove information,alter meaning,summarize,shorten,or provide advice.
Translate the full source from the first character through the final character.












