Can you tell a poem’s author from its lines?
Shakespeare’s sonnets, Keats’s lush odes, Whitman’s long breathless lines, Dickinson’s dashes: these voices are so familiar that we feel we would know them anywhere. So what happens when an AI imitates the same voice?
In this experiment you read two short excerpts for each of four poets. Each excerpt is either from the poet’s own work or by an AI writing in their style. First you say how much you like the excerpt, then you guess who wrote it. On the results screen you can see the sources of all the excerpts and which one fooled the crowd most.
Porter and Machery’s experiment
Brian Porter and Édouard Machery (2024) collected five poems by each of ten English-language poets, from Chaucer to Dorothea Lasky, and asked ChatGPT-3.5 to write five poems in each poet’s style. They used the first five poems the AI produced as they were, without picking the best. Each of 1,634 participants read five real and five AI poems for one poet and guessed who wrote each one.
Accuracy was 46.6%: slightly below chance. More strikingly, participants said “written by a human” more often about the AI poems than about the real ones. The five poems least often judged human were all by real poets; four of the five most often judged human were written by AI.
We like them more, but look down on them once we know
In the second study, 696 participants rated the same poems on 14 qualities such as rhythm, imagery, beauty and meaning. The AI poems scored higher than the real poems on most of these qualities. But when participants were told a poem was written by AI, the ratings for that same poem fell.
The authors explain it like this: non-expert readers like AI poems, which convey emotions and ideas in more direct and understandable language; but they expect AI poetry to be bad. So they assume the poems they like were written by humans, and attribute the complex poems they struggle with to AI.
Older models, newer models
A few years ago the picture was different. Köbis and Mossink (2021) compared poems written with GPT-2 against real poems. When a random poem was chosen from the AI’s output, participants could tell it apart; but when a human selected the best one, they couldn’t. Participants also showed a slight aversion to AI-written poems.
Porter and Machery’s results show that with a newer model the distinction disappears even when no human makes a selection. The AI excerpts in this experiment were written by an AI model (Claude, Anthropic) and selected by us, which is closer to the “human in the loop” condition.
Why is an AI studio running this experiment?
ba creative labs is an AI studio and this lab is its showcase. The most honest way to show what AI can do is to give people an experiment they can test themselves on, and then explain everything. That is why we mark every AI excerpt in the solution and give the sources of the real ones.
Whatever the result, one thing doesn’t change: Shakespeare’s “Bare ruin’d choirs, where late the sweet birds sang” is the product of a life and a view of the world four centuries old. An AI can imitate that voice; being able to imitate it doesn’t make the voice worth less. Perhaps the real question isn’t “can we tell them apart?” but “why does telling them apart matter to us?”