Wikipedia talk:WikiProject AI Cleanup

This is the talk page for discussing WikiProject AI Cleanup and anything related to its purposes and tasks.

Put new text under old text. Click here to start a new topic.
New to Wikipedia? Welcome! Learn to edit; get help.

Archives: 1: 90 days

To help centralise discussions and keep related topics together, all non-archive subpages of this talk page redirect here.

This page has been mentioned by multiple media organizations:

Maiberg, Emanuel (9 October 2024). "The Editors Protecting Wikipedia from AI Hoaxes". 404 Media. Retrieved 9 October 2024.
Nine, Adrianna (9 October 2024). "People Are Stuffing Wikipedia with AI-Generated Garbage". ExtremeTech. Retrieved 10 October 2024.
Harrison Dupré, Maggie (10 October 2024). "Wikipedia Declares War on AI Slop". The Byte. Retrieved 10 October 2024.

I wanted to share a helpful tip for spotting AI generated articles on Wikipedia

If you look up several buzzwords associated with ChatGPT and limit the results to Wikipedia, it will bring up articles with AI-generated text. For example I looked up "vibrant" "unique" "tapestry" "dynamic" site:en.wikipedia.org and I found some (mostly) low-effort articles. I'm actually surprised most of these are articles about cultures (see Culture of Indonesia, Culture of Qatar, or Culture of Indonesia). 95.18.76.205 (talk) 01:54, 2 September 2024 (UTC)[reply]

Thanks! That matches with Wikipedia:WikiProject AI Cleanup/AI Catchphrases, feel free to add any new buzzwords you find! Chaotic Enby (talk · contribs) 02:00, 2 September 2024 (UTC)[reply]

A new WMF thing

Y'all might be interested in m:Future Audiences/Experiment:Add a Fact. Charlotte (Queen of Hearts • talk) 21:46, 26 September 2024 (UTC)[reply]

Is it possible to specifically tell LLM-written text from encyclopedically written articles?

The WikiProject page says "Automatic AI detectors like GPTZero are unreliable and should not be used." However, those detectors are full of false positives because LLM-written text stylistically overlap with human-written text. But Wikipedia doesn't seek to cover all breadth of human writing, only a very narrow strand (encyclopedic writing) that is very far from natural conservation. Is it possible to specifically train a model on (high-quality) Wikipedia text vs. average LLM output? Any false positive would likely be unencyclopedic and needing to be fixed regardless. MatriceJacobine (talk) 13:29, 10 October 2024 (UTC)[reply]

That would definitely be a possibility, as the two output styles are stylistically different enough to be reliably distinguished most of the time. If we can make a good corpus of both (from output of the most common LLMs on Wikipedia-related prompts on one side, and Wikipedia articles on the other), which should definitely be feasible, we could indeed train such a detector. I'd be more than happy to help work on this! Chaotic Enby (talk · contribs) 14:50, 10 October 2024 (UTC)[reply]

That is entirely possible, a corpus of both "Genuine" Articles and articles generated by LLMs would be better though, as the writing style of for example ChatGPT can still vary depending on prompting. Someone should collect/archive articles found to be certainly generated by Language Models and open-source it so the community can contribute. 92.105.144.184 (talk) 15:10, 10 October 2024 (UTC)[reply]

We do have Wikipedia:WikiProject AI Cleanup/List of uses of ChatGPT at Wikipedia and User:JPxG/LLM dungeon which could serve as a baseline, although it is still quite small for a corpus. A way to scale it would be to find the kind of prompts being used and use variations of them to generate more samples. Chaotic Enby (talk · contribs) 15:24, 10 October 2024 (UTC)[reply]

GPTZero etc

I have never used an automatic AI detector, but I would be interested to know why the advice is "Automatic AI detectors like GPTZero are unreliable and should not be used."

Obviously, we shouldn't just tag/delete any article that GPTZero flags, but I would have thought it could be useful to highlight to us articles that might need our attention. I can even imagine a system like WP:STiki that has a backlog of edits sorted by likelihood to be LLM-generated and then feeds those edits to trusted editors for review.

Yaris678 (talk) 14:30, 11 October 2024 (UTC)[reply]

It could indeed be useful to flag potential articles, assuming we keep in mind the risk that editors might over-rely on the flagging as a definitive indicator, given the risk of both false positives and false negatives. I would definitely support brainstorming such a backlog system, but with the usual caveats – notably, that a relatively small false positive rate can easily be enough to drown true positives. Which means, it should be emphasized that editorial judgement shouldn't be primarily based on GPTZero's assessment.

Regarding the advice as currently written, the issue is that AI detectors will lag behind the latest LLMs themselves, and will often only be accurate on older models on which they have been trained. Indeed, their inaccuracy has been repeatedly pointed out. Chaotic Enby (talk · contribs) 14:54, 11 October 2024 (UTC)[reply]

404 Media article

https://www.404media.co/email/d516cf7f-3b5f-4bf4-93da-325d9522dd79/?ref=daily-stories-newsletter Seananony (talk) 00:44, 12 October 2024 (UTC)[reply]

Question about To-Do List

I went to 3 of the articles listed, Petite size, I Ching, and Pension, and couldn't find any templates in the articles about AI generation. Is the list outdated? Seananony (talk) 02:21, 12 October 2024 (UTC)[reply]

Seananony: The to-do list page hasn't been updated since January. Wikipedia:WikiProject AI Cleanup § Categories automatically catches articles with the {{AI-generated}} tag. Chaotic Enby, Queen of Hearts: any objections to unlinking the outdated to-do list? — ClaudineChionh (she/her · talk · contribs · email) 07:08, 12 October 2024 (UTC)[reply]

Fine with me! Chaotic Enby (talk · contribs) 11:14, 12 October 2024 (UTC)[reply]