Artificial intelligence versus Maya Angelou: Experimental evidence that people cannot differentiate AI-generated from human-written poetry

Искусственный интеллект против Майи Энджелоу: Экспериментальные данные о том, что люди не могут отличить стихи, созданные искусственным интеллектом, от стихов, написанных человеком
Nils Köbis, Luca Mossink
2020-09-08

AI-generated poetryGPT-2Turing Testhuman-agent experimentsnatural language generation
The release of openly available, robust natural language generation algorithms (NLG) has spurred much public attention and debate. One reason lies in the algorithms' purported ability to generate humanlike text across various domains. Empirical evidence using incentivized tasks to assess whether people (a) can distinguish and (b) prefer algorithm-generated versus human-written text is lacking. We conducted two experiments assessing behavioral reactions to the state-of-the-art Natural Language Generation algorithm GPT-2 (Ntotal = 830). Using the identical starting lines of human poems, GPT-2 produced samples of poems. From these samples, either a random poem was chosen (Human-out-of-theloop) or the best one was selected (Human-in-the-loop) and in turn matched with a human-written poem. In a new incentivized version of the Turing Test, participants failed to reliably detect the algorithmicallygenerated poems in the Human-in-the-loop treatment, yet succeeded in the Human-out-of-the-loop treatment. Further, people reveal a slight aversion to algorithm-generated poetry, independent on whether participants were informed about the algorithmic origin of the poem (Transparency) or not (Opacity). We discuss what these results convey about the performance of NLG algorithms to produce human-like text and propose methodologies to study such learning algorithms in human-agent experimental settings.
1
Participants failed to reliably identify GPT-2 poems when humans selected the best generated sample, but succeeded when samples were randomly selected.
2
Participants showed a slight aversion to algorithm-generated poetry, regardless of whether its algorithmic origin was disclosed.
3
The findings demonstrate that human selection substantially improves the perceived human-likeness of NLG poetry and motivate experimental methods for studying human-agent interactions.
4
Two experiments with 830 participants tested whether people could distinguish GPT-2-generated poetry from human-written poetry using incentivized tasks.

GPT-2-generated and human-written poetry

People’s ability to distinguish and preference for algorithm-generated versus human-written poetry, including effects of selection procedure and transparency

Publication Details
Publication Date
2020-09-08
Journal
Publisher
ISSN
Access Type
Author Information
Authors
Nils Köbis
Luca Mossink
Explore further
Open the scid.ai AI chat with a ready-made request: it will find papers on a similar topic and help build a literature review.
Find similar papers in the chat →
Make a presentation
100%