Anderson’un Açısı
Yapay Zeka ‘Pembe Çamur’ Haberlerini Tanıyabilir

Agenda-driven opinion mills, designed more to sway public opinion than serve the public, might be harder to spot if AI is used to make them sound more original and rational. So the race is on to stay ahead in the ‘pink slime detection’ game.
Son yirmi yılda geleneksel yerel medya kuruluşlarının fonlarının kesilmesi, hem değişen medya trendleri hem de – son zamanlarda – ABD hükümeti politikası nedeniyle, bölgesel raporlamada bir boşluk oluşmuş ve bu boşluk partizan organizasyonlar tarafından yapay zeka kullanarak gündemlerini ilerletmek için aceleyle doldurulmuştur.
‘Partizan’ terimini bağlam içinde kullanmak gerekirse (çünkü hiçbir haber kuruluşu siyasi eğilimlerden tamamen uzak değildir), petrol şirketlerinin yerel haber sitelerini uzak yerlerden yönettiğini, yerel kaynaklara sahip olmadığını, ancak şirketin kamu imajını savunma görevi olduğunu, seçimlerden önce siyasi olarak motive edilmiş haber sitelerinin gelir akışından yoksun olduğunu ve Cumhuriyetçi haber sitelerinin hiçbir yerden ortaya çıktığını konuşuyoruz.
2024 yılında, yapay zeka destekli pembe çamur haberlerinin gerçek haber sitelerini nihayet aşmış olabileceği tahmin edildi; o zaman, bir Avustralya anketi, tüketicilerin %41’inin pembe çamur kaynaklarını “gerçek” kaynaklardan daha çok tercih ettiğini buldu.
Bu tür gizli seçim propagandası, demokrasiye (siyasi olarak motive edilmiş siteler için) ve haberlerin makul standartlarında adil olunmasına güven için bir varoluşsal tehdit olarak görülebilir.
Pembe çamur yayıncıları ve yayıncılarının karakteristik çıktısını geleneksel medya kuruluşlarından ayırt etmek için yöntemler geliştirilmesi büyük yardım olurdu.
As it stands, the tropes and templates of real news organizations are very easy to mimic, and AI makes scalable publishing a current and affordable reality, using many of the same tricks being adopted by budget-stricken ‘old media’ publishers and broadcasters.
Sinyal ve Gürültü
A new study from the US addresses this issue, by investigating the growing use of Large Language Models to make pink slime websites sound less generic and easy-to-spot, and by creating a learning framework designed to keep up with evolving changes in pink slime (PS) output.
Titled Pembe Çamur Gazeteciliğini Açığa Çıkarmak: Dilsel İmzalar ve LLM Oluşturulan Tehditlere Karşı Dayanıklı Tespit, the new work comes from five researchers at the University of Texas.
The new work investigates how mass-produced PS local news articles differ from legitimate reporting, focusing on their reliance on short, repetitive structures and templated phrasing with minimal variation; and the authors note that PS articles tend to reuse identical templates designed to manipulate public opinion, with appeals to emotion uppermost in the content:

Yeni çalışmadan – birden fazla yayın aynı içerikleri sadece konum ayrıntılarını değiştirerek yayınlıyor, bu da meşru yerel haberleri taklit eden içerikleri kitleler halinde üretmek için kullanılan bir kopyala-yapıştır stratejisi ortaya koyuyor. Kaynak
Traditional detection models trained on these traits perform well against such content, but fail when the articles are rewritten using AI chatbots to appear more natural or sophisticated.
The authors’ own tests indicate that even minor stylistic changes introduced by large language models can reduce detection accuracy by up to 40%. To mitigate this, they propose a continual learning framework that incrementally retrains detection models on both original and AI-rewritten articles, to adapt to shifting linguistic patterns.
Yöntem
To establish data for the project, the authors used the Pembe Çamur Veri Seti, featuring 7.9 million articles covering 1,093 outlets during 2021-2023, from which they obtained 9,472 pink slime articles after filtering. They also used the LIAR dataset, which contains annotated fake news, as well as the NELA-GT-2021 collection, which contains solely US articles*.
To prepare their training and testing sets, the authors first used the T-distributed Stochastic Neighbor Embedding (t-SNE) algorithm to reduce the article embeddings to two dimensions. They then applied the data clustering algorithm Density-Based-Spatial-Clustering-of-Applications-with-Noise (DBSCAN) to isolate clusters of similar pink-slime articles.
Each cluster was treated as a group of related stories, many of which still followed the same template, despite a concerted effort to address duplicates.
To prevent similar articles from appearing in both training and test sets, entire clusters were randomly selected, with 80% used for training and 20% for testing. Because the legitimate news articles did not form clear clusters, a random split was applied instead.
This process was repeated three times, to ensure consistency, and to reduce sampling bias.
Pembe Çamurun Özellikleri
Commenting on the distinguishing traits of PS vs. regular news, the researchers assert that PS-style local news articles are significantly shorter and simpler than legitimate reporting, averaging fewer than dokuz cümle per article.
A higher proportion of simple sentences and a heavier reliance on adjectives are further hallmarks of pink slime, according to the paper, and indicate a penchant for repetitive, emotionally charged language.
Lexical richness was measured using the Root-Type-Token Ratio (RTTR), and found to be notably lower in the PS articles, which also exhibited far fewer unique noun phrases.
These patterns denote a limited vocabulary and formulaic style, in contrast to legitimate local news, which is characterized by complex part-of-speech patterns built around auxiliary verbs, pronouns, and conjunctions. Instead, the fake articles favor basic noun-preposition structures, with frequent use of punctuation-based trigrams, suggesting a less formal, more fragmented writing style.
Testler
To examine the associations between different types of news articles, based on linguistic and structural features, embeddings were generated using the 435-million parameter stella_en_400M_v5 model, and reduced with Principal Component Analysis (PCA), and t-SNE for visualization.
When projected into two dimensions, the fake local news articles formed small, dense clusters, each corresponding to narrowly focused topics such as crime statistics, stock updates, or charitable donations:

t-SNE projeksiyonundan elde edilen kümeleme kalıpları, pembe çamur makalelerinin sıkı, tekrarlı gruplar oluşturduğunu, meşru haberlerin ise daha geniş, daha çeşitli dağılımlar sergilediğini gösterir.
As we can see, to some extent, in the visualization above, this pattern suggests a rigid, template-driven format, with minimal variation between articles.
Interestingly, articles labeled as ‘fake news’ diverged from the fake local content, showing a distribution more in line with real news, indicating that mass-produced local fakes may not be simply less truthful, but may also be mechanically distinct in form and composition.
By contrast, ‘legitimate’ local news forms fewer and more widely spaced clusters, consistent with more diverse language and subject matter, while national news articles show even greater dispersion, reflecting broader topical range and looser stylistic consistency.

Meşru yerel haberler ve pembe çamur içeriği arasındaki özellikler karşılaştırması, PS makalelerinin daha kısa olduğunu, daha basit cümle yapılarını kullandığını, daha fazla sıfat içerdiğini, daha düşük sözcük zenginliğine sahip olduğunu, temel sözdizimi üçlülerini tercih ettiğini ve daha az benzersiz isim terkibine sahip olduğunu gösterir.
Tespit
The researchers evaluated two main approaches to detect pink slime content: classification, based on handcrafted linguistic features; and transformer-based fine-tuning.
For the handcrafted approach, structural rather than semantic characteristics were emphasized, using sentence count; lexical richness; syntactic depth; part-of-speech co-occurrence probabilities; dependency tag co-occurrence probabilities; readability; and part-of-speech counts.
Three models were tested on this feature set: XGBoost; Random Forest; and Support Vector Machine (SVM) – with Random Forest showing slightly stronger results overall.
Both XGBoost and Random Forest assigned high predictive importance to features such as sentence count and the number of unique noun phrases. Readability and lexical richness measures also influenced classification strongly, though the models weighted these differently, with XGBoost favoring Flesch and RTTR, while Random Forest leaned on CTTR:

SHAP (SHapley Additive exPlanations) temelli özellik önem puanları, her bir girdi özelliğinin örnekler genelinde model çıktısını nasıl etkilediğini vurgular. Bu durumda, SHAP değerleri, XGBoost ve Random Forest’un her ikisinin de pembe çamuru gerçek haberden ayırmak için cümle sayısına ve benzersiz isim terkibine en çok güvendiğini gösterirken, sözcük zenginliği ve okunabilirlik ölçümlerine farklı ağırlık verdiğini gösterir.
As seen in the earlier feature comparisons (above), pink slime articles favor sensationalism over detail, with a lower lexical richness, and with fewer distinct noun phrases – reinforcing the conclusion that this kind of content is highly templated and repetitive.
Part-of-speech trigram patterns confirm that legitimate local news tends to use more structurally complex forms, involving auxiliaries, pronouns, and conjunctions, while pink slime favors fragmented or minimal syntax. These patterns suggest a reliable foundation for distinguishing synthetic local content from real journalism.
A second round of testing involved fine-tuning transformer models on full article text, to capture both semantic content and syntactic structure.
BERT, XLNet, and Flan-T5 were tested, with BERT achieving the highest F1-score at 89.31% – although the paper notes that performance differences were not statistically significant. Unlike the handcrafted classifiers, these models update all weights during training, allowing them to learn task-specific representations directly from the data.
An additional embedding-based method used a fully-connected downstream classifier performed better than the handcrafted models, but did not match the accuracy of full fine-tuning.
These approaches benefited from broader linguistic context – advantageous, given that pink slime articles often recycle content across outlets. Handcrafted features, by contrast, rely on surface-level syntax alone, limiting their ability to generalize across sources.
LLM’lerle Pembe Çamurun Güçlendirilmesi
To test whether pink slime articles could be made harder to detect, the researchers rewrote them using large language models. Prompts were designed to improve the writing style by expanding content, adding sentence complexity, reducing overuse of adjectives, enforcing a neutral tone, using more varied vocabulary, and increasing the number of unique noun phrases, with the ambit of hiding the surface-level traits that most detection systems rely on.
The models tested were Gemma‑2B; Mistral‑7B; Llama‑1B; Llama‑3B; Llama‑8B; Qwen‑7B; GPT‑4o mini; and Claude‑3.5‑Haiku.

Farklı LLM’ler tarafından yeniden yazılmış pembe çamur makalelerinde ince ayarlanmış bir BERT sınıflandırıcının tespit performansı (F1). Daha düşük puanlar daha başarılı gizleme anlamına gelir.
Of these results, the authors comment:
‘[Açık kaynaklı modeller arasında] Qwen, tespitten kaçınmada en etkili olanıydı: daha önce ince ayarlanmış bir BERT sınıflandırıcıya (başlangıç F1 puanı: %89) karşı test edildiğinde, tespit performansında %28,34’lük bir düşüşe neden oldu.’
‘Kapalı kaynaklı modeller çok daha güçlüydü, GPT-4o-mini ve Claude-3.5-Haiku, F1 puanını ortalama %40 oranında düşürdü, bu da yüksek kaliteli LLM oluşturulan gizlemenin oluşturduğu zorluğu vurguladı.’
These results, the authors contend, show how easily LLMs can disguise pink slime content, making it much harder for current tools to catch**.
Sonuç
Görüş Bu araştırma hattı bazı ilginç ikilemler içerir, bunlardan en az biri, birçok insanın (en az bir anket tarafından bahsedildiği gibi) desteklediği pembe çamur içeriği, bu da pejoratif bağlamı sorgulamaya yol açar. İnsanların “Soylent Green insanların eti” olduğunu bildikleri, ancak yine de yedikleri gibi görünüyor; ya da liberal bir duyarlılıktan bakıldığında böyle görünüyor.
Bu kamu ilgisizliği algoritmik haberlere karşı derinleşiyor gibi görünüyor.
Beni bu makaleyi okurken etkileyen bir diğer şey, pembe çamur çıkışının daha basit üslubu ve indirgeyiciliğinin bir eksiklik olarak görülmesi, ancak bu minimalizm, duygusallık ve sınırlı sözcük dağarcığının aslında kasıtlı olduğudur.
Eğer pembe çamurun arkasındaki çeşitli çıkar grupları daha entelektüel veya liberal bir kitleye ulaşmaya çalışıyorsa (bu, güçlerine göre olmayabilir), mevcut platformlarda değil, hedef kitlelerine daha yakın bir yerde kamp kurmaları daha muhtemeldir.
* Maalesef, makaledeki bazı biçimlendirme sorunları nedeniyle, yerel haber makalelerinin ek kaynağı net bir atıfta bulunmuyor. Lütfen kaynak makaleye başvurun ve hangi ‘Horne’ referansının uygulanacağını tahmin edin.
** Burada, sonuç bölümünün sonundaki ikincil, tamamlayıcı deneylerin ayrıntıları için okuyucuyu kaynak makaleye yönlendiriyoruz.
İlk olarak 12 Aralık 2025 Cuma günü yayınlandı.












