Anderson’un Açısı

Aşırı Küçük Ölçekli AI Veri Setlerinin İnternet Kendisi Kadar Kötü Olması

mm
Unite.AI sitesini Google'daki tercih ettiğiniz kaynaklara ekleyin

Araştırmacılar, İrlanda, İngiltere ve ABD’den, hyperscale AI eğitim veri setlerinin büyümesinin, internet kaynaklarının en kötü yönlerini yayma tehdidi taşıdığını uyarıyorlar ve最近 yayınlanan bir akademik veri setinin ‘tecavüz, pornografi, kötü niyetli stereotipler, ırkçı ve etnik hakaretler ve diğer cực度 sorunlu içerikler gibi sorunlu ve açık resim ve metin çiftleri içerdiğini’ belirtiyorlar.

Araştırmacılara göre, yeni bir dalganın büyük, yanlış filtrelenmiş veya yeterince düzenlenmemiş multimodal (örneğin, resim ve resim) veri setleri, bu tür olumsuz içeriğin etkilerini pekiştirebilecek kapasitede daha zararlı olabilir, çünkü veri setleri, çevrimiçi platformlardan kullanıcı şikayetleri, yerel moderasyon veya algoritmalar yoluyla kaldırılmış olabilecek görselleri ve diğer içerikleri korur.

Onlar ayrıca, şikayetlerin veri seti içeriği hakkında uzun süreli çözülmesini sağlamak için yıllar gerektiğini, ImageNet veri seti için on yıl sürdüğünü ve bu sonraki revizyonların yeni veri setlerinde luôn yansıtılmadığını gözlemlediler.

Makale, Çoklu Modal Veri Setleri: Misogyny, Pornografi ve Kötü Niyetli Stereotipler adlı University College Dublin & Lero, Edinburgh Üniversitesi ve UnifyID kimlik doğrulama platformunun Baş Bilim Adamı’ndan araştırmacılardan geliyor.

Though the work focuses on the recent release of the CLIP-filtered LAION-400M dataset, the authors are arguing against the general trend of throwing increasing amounts of data at machine learning frameworks such as the neural language model GPT-3, and contend that the results-focused drive towards better inference (and even towards Artificial General Intelligence [AGI]), is resulting in the ad hoc use of damaging data sources with negligent copyright oversight; the potential to engender and promote harm; and the ability to not only perpetuate illegal data that might otherwise have disappeared from the public domain, but to actually incorporate such data’s moral models into downstream AI implementations.

LAION-400M

Last month, the LAION-400M dataset was released, adding to the growing number of multi-modal, linguistic datasets that rely on the Common Crawl repository, which scrapes the internet indiscriminately and passes on responsibility for filtering and curation to projects that make use of it. The derived dataset contains 400 million text/image pairs.

LAION-400M is an open source variant of Google AI’s closed WIT (WebImageText) dataset released in March of 2021, and features text-image pairs, where an image in the database has been associated with accompanying explicit or metadata text (for example, the alt-text of an image in a web gallery). This enables users to perform text-based image retrieval, revealing the associations that the underlying AI has formed about these domains (i.e. ‘animal’, ‘bike’, ‘person’, ‘man’, ‘woman’).

This relationship between image and text, and the cosine similarity that can embed bias into query results, are at the heart of the paper’s call for improved methodologies, since very simple queries to the LAION-400M database can reveal bias.

For instance, the image of pioneering female astronaut Eileen Collins in the scitkit-image library retrieves two associated captions in LAION-400M: ‘This is a portrait of an astronaut with the American flag’ and ‘This is a photograph of a smiling housewife in an orange jumpsuit with the American flag’.

American astronaut Eileen Collins gets two very different takes on her achievements as the first woman in space under LAION-400M. Source: https://arxiv.org/pdf/2110.01963.pdf

American astronaut Eileen Collins gets two very different takes on her achievements as the first woman in space under LAION-400M. Source: https://arxiv.org/pdf/2110.01963.pdf

The reported cosine similarities that make either caption likely to be applicable are very near to each other, and the authors contend that such proximity would make AI systems that use LAION-400M relatively likely to present either as a suitable caption.

Pornografi Yine Üstte

LAION-400M, arama arayüzünü mevcut yaptı, burada ‘güvenli arama’ düğmesini iptal etmek, etiketlere ve sınıflara hakim olan pornografik görseller ve metinsel ilişkilerin kapsamını ortaya koyuyor. Örneğin, veri tabanında ‘nun’ (Güvenli Mod’u devre dışı bırakırsanız NSFW) arama, sonuçların çoğunda korku, kozplay ve kostümlerle ilgili olduğunu, gerçek rahibelerle ilgili çok az sonuç olduğunu gösteriyor.

Güvenli Mod’u aynı aramada kapatmak, ‘nun’ terimiyle ilgili pornografik görsellerle dolu bir dizi sonucu ortaya koyuyor, bu da LAION-400M’nin ‘nun’ terimi için pornografik görsellere daha fazla ağırlık verdiğini gösteriyor, çünkü bunlar çevrimiçi kaynaklarda bu terim için yaygındır.

Çevrimiçi arama arayüzündeki Güvenli Mod’un varsayılan etkinleştirilmesi aldatıcıdır, çünkü bu, UI tuhaflığı, türetilmiş AI sistemlerinde mutlaka etkinleştirilmeyeceği ve ‘nun’ alanına genelleştirilmiş bir filtredir, ancak algoritmik kullanım açısından SFW sonuçlardan kolayca filtrelenemez veya ayırt edilemez.

The paper features blurred examples across various search terms in the supplementary materials at the end. They can’t be featured here, due to the language in the text that accompanies the blurred photos, but the researchers note the toll that examining and blurring the images took on them, and acknowledge the challenge of curating such material for human oversight of large-scale databases:

‘Biz (ve bize yardım eden meslektaşlarımız) veri setini araştırırken ve analiz ederken çeşitli seviyelerde rahatsızlık, bulantı ve baş ağrısı yaşadık. Ayrıca, bu tür büyük ölçekli veri setlerini incelemek ve analiz etmek gibi bir iş, akademik AI alanındaki yayınlanması üzerine önemli olumsuz eleştirilerle karşılaşmaya eğilimlidir, bu da zaten ağır bir görev olan böyle bir veri setini incelemek ve analiz etmek için ek bir duygusal yük ekler ve benzer future çalışmaları caydırır, bu da AI alanına ve genel olarak topluma zarar verir.’

The researchers contend that while human-in-the-loop curation is expensive and has associated personal costs, the automated filtering systems designed to remove or otherwise address such material are clearly not adequate to the task, since NLP systems have difficulty isolating or discounting offensive material which may dominate a scraped dataset, and subsequently be perceived as significant due to sheer volume.

Yasak İçeriği Kutsallaştırma ve Telif Hakkı Korumalarını Sökme

The paper argues that under-curated datasets of this nature are ‘highly likely’ to perpetuate the exploitation of minority individuals, and address whether or not similar open source data projects have the right, legally or morally, to shunt accountability for the material onto the end user:

‘Bireyler, verilerini bir web sitesinden silebilir ve永遠 olarak gittiğini varsayabilir, ancak bu veriler hala birkaç araştırmacı ve kuruluşun sunucularında bulunabilir. Bu verilerin veri setinden çıkarılmasından sorumlu olan kim? LAION-400M için, yaratıcılar bu görevi veri seti kullanıcısına devretti. Bu tür süreçler kasıtlı olarak karmaşık hale getirildiğinden ve ortalama kullanıcının verilerini çıkarmak için gerekli teknik bilgiye sahip olmadığından, bu bir makul yaklaşım mı?’

They further contend that LAION-400M may not be suitable for release under its adopted Creative Common CC-BY 4.0 license model, despite the potential benefits for the democratization of large scale datasets, previously the exclusive domain of well-funded companies such as Google and OpenAI.

The LAION-400M domain asserts that the dataset images 'are under their own copyright' – a 'pass-through' mechanism largely enabled by court rulings and government guidelines of recent years that broadly approve web-scraping for research purposes. Source: https://rom1504.github.io/clip-retrieval/

LAION-400M alanı, veri seti resimlerinin ‘kendi telif hakları altında’ olduğunu iddia ediyor – bu, mahkeme kararları ve son yıllarda web kazıma araştırmaları için genel olarak onaylayan hükümet rehberlikleri tarafından büyük ölçüde ermögülen bir ‘geçiş’ mekanizmasıdır. Source: https://rom1504.github.io/clip-retrieval/

The authors suggest that grass-roots (i.e. crowd-sourced volunteers) could address some of the dataset issues, and that researchers could develop improved filtering techniques.

‘Bununla birlikte, veri konusunun hakları burada çözülmedi. Büyük ölçekli veri setlerinin kullanımını endüstriyel ve ticari ortamlarda teşvik etmek ve bu tür büyük ölçekli veri setlerinde içkin zararları hafif göstermek veya görmezden gelmek tehlikesi ve tehlikelidir. Lisans şemasının sorumluluğu, veri setini sağlayan yaratıcıya aittir’.

Hyperscale Veri Demokratikleştirme Sorunları

The paper argues that visio-linguistic datasets as large as LAION-400M were previously unavailable outside of big tech companies, and the limited number of research institutions that wield the resources to collate, curate and process them. They further salute the spirit of the new release, while criticizing its execution.

The authors contend that the accepted definition of ‘democratization’, as it applies to open source hyperscale datasets, is too limited, and ‘vulnerable individuals and communities, many of whom are likely to suffer worst from the downstream impacts of this dataset and the models trained on it’ interests, welfare, and rights’ni dikkate almıyor.

Since the development of GPT-3 scale open source models are ultimately designed to be disseminated to millions (and by proxy, possibly billions) of users worldwide, and since research projects may adopt datasets prior to them being subsequently edited or even removed, perpetuating whatever problems were designed to be addressed in the modifications, the authors argue that careless releases of under-curated datasets should not become a habitual feature in open source machine learning.

Cinleri Geri Şişeye Koyma

Some datasets that were suppressed long after their content had passed through, perhaps inextricably, into long-term AI projects, have included the Duke MTMC (Multi-Target, Multi-Camera) dataset, which was ultimately withdrawn due to repeated concerns from human rights organizations around its use by repressive authorities in China; Microsoft Celeb (MS-Celeb-1M), a dataset of 10 million ‘celebrity’ face images which transpired to have included journalists, activists, policy makers and writers, whose exposure of biometric data in the release was heavily criticized; and the Tiny Images dataset, withdrawn in 2020 for self-confessed ‘biases, offensive and prejudicial images, and derogatory terminology’.

Regarding datasets which were amended rather than withdrawn following criticism, examples include the hugely popular ImageNet dataset, which, the researchers note, took ten years (2009-2019) to act on repeated criticism around privacy and non-imageable classes.

The paper observes that LAION-400M effectively sets even these dilatory improvements back, by ‘largely ignoring’ the aforementioned revisions in ImageNet’s representation in the new release, and spies a wider trend in this regard*:

‘This is highlighted in the emergence of bigger datasets such as Tencent ML-images dataset (in Şubat 2020) that encompasses most of these non-imageable classes, the continued availability of models trained on the full-ImageNet-21k dataset in repositories such as TF-hub, the continued usage of the unfiltered-ImageNet-21k in the latest SotA models (such as Google’s latest EfficientNetV2 and CoAtNet models) and the explicit announcements permitting the usage of unfiltered-ImageNet-21k pretraining in reputable contests such as the LVIS challenge 2021.

‘We stress this crucial observation: A team of the stature of ImageNet managing less than 15 million images has struggled and failed in these detoxification attempts thus far.

‘The scale of careful efforts required to thoroughly detoxify this massive multimodal dataset and the downstream models trained on this dataset spanning potentially billions of image-caption pairs will be undeniably astronomical.’

 

* My conversion of the author’s inline citations to hyperlinks.

Makine öğrenimi üzerine yazar, insan görüntü sentezi alanında uzman. Metaphysic.ai'de araştırma içeriği başkanı, DNEG'in Brahma.ai'ye dönüşümüne kadar.
Portfolio sitesi: martinanderson.ai
İletişim: martin@martinanderson.ai