Andersonâun AÃ§ÄąsÄą
Yapay Zeka ‘Pembe Ãamur’ Haberlerini TanÄąyabilir

Agenda-driven opinion mills, designed more to sway public opinion than serve the public, might be harder to spot if AI is used to make them sound more original and rational. So the race is on to stay ahead in the âpink slime detectionâ game.
Â
Son yirmi yÄąlda geleneksel yerel medya kuruluÅlarÄąnÄąn fonlarÄąnÄąn kesilmesi, hem deÄiÅen medya trendleri hem de â son zamanlarda â ABD hÞkÞmeti politikasÄą nedeniyle, bÃķlgesel raporlamada bir boÅluk oluÅmuÅ ve bu boÅluk partizan organizasyonlar tarafÄąndan yapay zeka kullanarak gÞndemlerini ilerletmek için aceleyle doldurulmuÅtur.
âPartizanâ terimini baÄlam içinde kullanmak gerekirse (çÞnkÞ hiçbir haber kuruluÅu siyasi eÄilimlerden tamamen uzak deÄildir), petrol Åirketlerinin yerel haber sitelerini uzak yerlerden yÃķnettiÄini, yerel kaynaklara sahip olmadÄąÄÄąnÄą, ancak Åirketin kamu imajÄąnÄą savunma gÃķrevi olduÄunu, seçimlerden Ãķnce siyasi olarak motive edilmiÅ haber sitelerinin gelir akÄąÅÄąndan yoksun olduÄunu ve Cumhuriyetçi haber sitelerinin hiçbir yerden ortaya Ã§ÄąktÄąÄÄąnÄą konuÅuyoruz.
2024 yÄąlÄąnda, yapay zeka destekli pembe çamur haberlerinin gerçek haber sitelerini nihayet aÅmÄąÅ olabileceÄi tahmin edildi; o zaman, bir Avustralya anketi, tÞketicilerin %41âinin pembe çamur kaynaklarÄąnÄą âgerçekâ kaynaklardan daha çok tercih ettiÄini buldu.
Bu tÞr gizli seçim propagandasÄą, demokrasiye (siyasi olarak motive edilmiÅ siteler için) ve haberlerin makul standartlarÄąnda adil olunmasÄąna gÞven için bir varoluÅsal tehdit olarak gÃķrÞlebilir.
Pembe çamur yayÄąncÄąlarÄą ve yayÄąncÄąlarÄąnÄąn karakteristik Ã§ÄąktÄąsÄąnÄą geleneksel medya kuruluÅlarÄąndan ayÄąrt etmek için yÃķntemler geliÅtirilmesi bÞyÞk yardÄąm olurdu.
As it stands, the tropes and templates of real news organizations are very easy to mimic, and AI makes scalable publishing a current and affordable reality, using many of the same tricks being adopted by budget-stricken âold mediaâ publishers and broadcasters.
Sinyal ve GÞrÞltÞ
A new study from the US addresses this issue, by investigating the growing use of Large Language Models to make pink slime websites sound less generic and easy-to-spot, and by creating a learning framework designed to keep up with evolving changes in pink slime (PS) output.
Titled Pembe Ãamur GazeteciliÄini AÃ§ÄąÄa ÃÄąkarmak: Dilsel İmzalar ve LLM OluÅturulan Tehditlere KarÅÄą DayanÄąklÄą Tespit, the new work comes from five researchers at the University of Texas.
The new work investigates how mass-produced PS local news articles differ from legitimate reporting, focusing on their reliance on short, repetitive structures and templated phrasing with minimal variation; and the authors note that PS articles tend to reuse identical templates designed to manipulate public opinion, with appeals to emotion uppermost in the content:

Yeni çalÄąÅmadan â birden fazla yayÄąn aynÄą içerikleri sadece konum ayrÄąntÄąlarÄąnÄą deÄiÅtirerek yayÄąnlÄąyor, bu da meÅru yerel haberleri taklit eden içerikleri kitleler halinde Þretmek için kullanÄąlan bir kopyala-yapÄąÅtÄąr stratejisi ortaya koyuyor. Kaynak
Traditional detection models trained on these traits perform well against such content, but fail when the articles are rewritten using AI chatbots to appear more natural or sophisticated.
The authorsâ own tests indicate that even minor stylistic changes introduced by large language models can reduce detection accuracy by up to 40%. To mitigate this, they propose a continual learning framework that incrementally retrains detection models on both original and AI-rewritten articles, to adapt to shifting linguistic patterns.
YÃķntem
To establish data for the project, the authors used the Pembe Ãamur Veri Seti, featuring 7.9 million articles covering 1,093 outlets during 2021-2023, from which they obtained 9,472 pink slime articles after filtering. They also used the LIAR dataset, which contains annotated fake news, as well as the NELA-GT-2021 collection, which contains solely US articles*.
To prepare their training and testing sets, the authors first used the T-distributed Stochastic Neighbor Embedding (t-SNE) algorithm to reduce the article embeddings to two dimensions. They then applied the data clustering algorithm Density-Based-Spatial-Clustering-of-Applications-with-Noise (DBSCAN) to isolate clusters of similar pink-slime articles.
Each cluster was treated as a group of related stories, many of which still followed the same template, despite a concerted effort to address duplicates.
To prevent similar articles from appearing in both training and test sets, entire clusters were randomly selected, with 80% used for training and 20% for testing. Because the legitimate news articles did not form clear clusters, a random split was applied instead.
This process was repeated three times, to ensure consistency, and to reduce sampling bias.
Pembe Ãamurun Ãzellikleri
Commenting on the distinguishing traits of PS vs. regular news, the researchers assert that PS-style local news articles are significantly shorter and simpler than legitimate reporting, averaging fewer than dokuz cÞmle per article.
A higher proportion of simple sentences and a heavier reliance on adjectives are further hallmarks of pink slime, according to the paper, and indicate a penchant for repetitive, emotionally charged language.
Lexical richness was measured using the Root-Type-Token Ratio (RTTR), and found to be notably lower in the PS articles, which also exhibited far fewer unique noun phrases.
These patterns denote a limited vocabulary and formulaic style, in contrast to legitimate local news, which is characterized by complex part-of-speech patterns built around auxiliary verbs, pronouns, and conjunctions. Instead, the fake articles favor basic noun-preposition structures, with frequent use of punctuation-based trigrams, suggesting a less formal, more fragmented writing style.
Testler
To examine the associations between different types of news articles, based on linguistic and structural features, embeddings were generated using the 435-million parameter stella_en_400M_v5 model, and reduced with Principal Component Analysis (PCA), and t-SNE for visualization.
When projected into two dimensions, the fake local news articles formed small, dense clusters, each corresponding to narrowly focused topics such as crime statistics, stock updates, or charitable donations:

t-SNE projeksiyonundan elde edilen kÞmeleme kalÄąplarÄą, pembe çamur makalelerinin sÄąkÄą, tekrarlÄą gruplar oluÅturduÄunu, meÅru haberlerin ise daha geniÅ, daha çeÅitli daÄÄąlÄąmlar sergilediÄini gÃķsterir.
As we can see, to some extent, in the visualization above, this pattern suggests a rigid, template-driven format, with minimal variation between articles.
Interestingly, articles labeled as âfake newsâ diverged from the fake local content, showing a distribution more in line with real news, indicating that mass-produced local fakes may not be simply less truthful, but may also be mechanically distinct in form and composition.
By contrast, âlegitimateâ local news forms fewer and more widely spaced clusters, consistent with more diverse language and subject matter, while national news articles show even greater dispersion, reflecting broader topical range and looser stylistic consistency.

MeÅru yerel haberler ve pembe çamur içeriÄi arasÄąndaki Ãķzellikler karÅÄąlaÅtÄąrmasÄą, PS makalelerinin daha kÄąsa olduÄunu, daha basit cÞmle yapÄąlarÄąnÄą kullandÄąÄÄąnÄą, daha fazla sÄąfat içerdiÄini, daha dÞÅÞk sÃķzcÞk zenginliÄine sahip olduÄunu, temel sÃķzdizimi ÞçlÞlerini tercih ettiÄini ve daha az benzersiz isim terkibine sahip olduÄunu gÃķsterir.
Tespit
The researchers evaluated two main approaches to detect pink slime content: classification, based on handcrafted linguistic features; and transformer-based fine-tuning.
For the handcrafted approach, structural rather than semantic characteristics were emphasized, using sentence count; lexical richness; syntactic depth; part-of-speech co-occurrence probabilities; dependency tag co-occurrence probabilities; readability; and part-of-speech counts.
Three models were tested on this feature set: XGBoost; Random Forest; and Support Vector Machine (SVM) â with Random Forest showing slightly stronger results overall.
Both XGBoost and Random Forest assigned high predictive importance to features such as sentence count and the number of unique noun phrases. Readability and lexical richness measures also influenced classification strongly, though the models weighted these differently, with XGBoost favoring Flesch and RTTR, while Random Forest leaned on CTTR:

SHAP (SHapley Additive exPlanations) temelli Ãķzellik Ãķnem puanlarÄą, her bir girdi ÃķzelliÄinin Ãķrnekler genelinde model Ã§ÄąktÄąsÄąnÄą nasÄąl etkilediÄini vurgular. Bu durumda, SHAP deÄerleri, XGBoost ve Random Forestâun her ikisinin de pembe çamuru gerçek haberden ayÄąrmak için cÞmle sayÄąsÄąna ve benzersiz isim terkibine en çok gÞvendiÄini gÃķsterirken, sÃķzcÞk zenginliÄi ve okunabilirlik ÃķlçÞmlerine farklÄą aÄÄąrlÄąk verdiÄini gÃķsterir.
As seen in the earlier feature comparisons (above), pink slime articles favor sensationalism over detail, with a lower lexical richness, and with fewer distinct noun phrases â reinforcing the conclusion that this kind of content is highly templated and repetitive.
Part-of-speech trigram patterns confirm that legitimate local news tends to use more structurally complex forms, involving auxiliaries, pronouns, and conjunctions, while pink slime favors fragmented or minimal syntax. These patterns suggest a reliable foundation for distinguishing synthetic local content from real journalism.
A second round of testing involved fine-tuning transformer models on full article text, to capture both semantic content and syntactic structure.
BERT, XLNet, and Flan-T5 were tested, with BERT achieving the highest F1-score at 89.31% â although the paper notes that performance differences were not statistically significant. Unlike the handcrafted classifiers, these models update all weights during training, allowing them to learn task-specific representations directly from the data.
An additional embedding-based method used a fully-connected downstream classifier performed better than the handcrafted models, but did not match the accuracy of full fine-tuning.
These approaches benefited from broader linguistic context â advantageous, given that pink slime articles often recycle content across outlets. Handcrafted features, by contrast, rely on surface-level syntax alone, limiting their ability to generalize across sources.
LLMâlerle Pembe Ãamurun GÞçlendirilmesi
To test whether pink slime articles could be made harder to detect, the researchers rewrote them using large language models. Prompts were designed to improve the writing style by expanding content, adding sentence complexity, reducing overuse of adjectives, enforcing a neutral tone, using more varied vocabulary, and increasing the number of unique noun phrases, with the ambit of hiding the surface-level traits that most detection systems rely on.
The models tested were Gemmaâ2B; Mistralâ7B; Llamaâ1B; Llamaâ3B; Llamaâ8B; Qwenâ7B; GPTâ4o mini; and Claudeâ3.5âHaiku.

FarklÄą LLMâler tarafÄąndan yeniden yazÄąlmÄąÅ pembe çamur makalelerinde ince ayarlanmÄąÅ bir BERT sÄąnÄąflandÄąrÄącÄąnÄąn tespit performansÄą (F1). Daha dÞÅÞk puanlar daha baÅarÄąlÄą gizleme anlamÄąna gelir.
Of these results, the authors comment:
â[AÃ§Äąk kaynaklÄą modeller arasÄąnda] Qwen, tespitten kaÃ§Äąnmada en etkili olanÄąydÄą: daha Ãķnce ince ayarlanmÄąÅ bir BERT sÄąnÄąflandÄąrÄącÄąya (baÅlangÄąÃ§ F1 puanÄą: %89) karÅÄą test edildiÄinde, tespit performansÄąnda %28,34âlÞk bir dÞÅÞÅe neden oldu.â
âKapalÄą kaynaklÄą modeller çok daha gÞçlÞydÞ, GPT-4o-mini ve Claude-3.5-Haiku, F1 puanÄąnÄą ortalama %40 oranÄąnda dÞÅÞrdÞ, bu da yÞksek kaliteli LLM oluÅturulan gizlemenin oluÅturduÄu zorluÄu vurguladÄą.â
These results, the authors contend, show how easily LLMs can disguise pink slime content, making it much harder for current tools to catch**.
Sonuç
GÃķrÞŠBu araÅtÄąrma hattÄą bazÄą ilginç ikilemler içerir, bunlardan en az biri, birçok insanÄąn (en az bir anket tarafÄąndan bahsedildiÄi gibi) desteklediÄi pembe çamur içeriÄi, bu da pejoratif baÄlamÄą sorgulamaya yol açar. İnsanlarÄąn âSoylent Green insanlarÄąn etiâ olduÄunu bildikleri, ancak yine de yedikleri gibi gÃķrÞnÞyor; ya da liberal bir duyarlÄąlÄąktan bakÄąldÄąÄÄąnda bÃķyle gÃķrÞnÞyor.
Bu kamu ilgisizliÄi algoritmik haberlere karÅÄą derinleÅiyor gibi gÃķrÞnÞyor.
Beni bu makaleyi okurken etkileyen bir diÄer Åey, pembe çamur Ã§ÄąkÄąÅÄąnÄąn daha basit Þslubu ve indirgeyiciliÄinin bir eksiklik olarak gÃķrÞlmesi, ancak bu minimalizm, duygusallÄąk ve sÄąnÄąrlÄą sÃķzcÞk daÄarcÄąÄÄąnÄąn aslÄąnda kasÄątlÄą olduÄudur.
EÄer pembe çamurun arkasÄąndaki çeÅitli Ã§Äąkar gruplarÄą daha entelektÞel veya liberal bir kitleye ulaÅmaya çalÄąÅÄąyorsa (bu, gÞçlerine gÃķre olmayabilir), mevcut platformlarda deÄil, hedef kitlelerine daha yakÄąn bir yerde kamp kurmalarÄą daha muhtemeldir.
Â
* Maalesef, makaledeki bazÄą biçimlendirme sorunlarÄą nedeniyle, yerel haber makalelerinin ek kaynaÄÄą net bir atÄąfta bulunmuyor. LÞtfen kaynak makaleye baÅvurun ve hangi âHorneâ referansÄąnÄąn uygulanacaÄÄąnÄą tahmin edin.
** Burada, sonuç bÃķlÞmÞnÞn sonundaki ikincil, tamamlayÄącÄą deneylerin ayrÄąntÄąlarÄą için okuyucuyu kaynak makaleye yÃķnlendiriyoruz.
İlk olarak 12 AralÄąk 2025 Cuma gÞnÞ yayÄąnlandÄą.












