Open Artificial intelligence Porter and Machery, 2024

Written by a human or an AI?

Eight short excerpts from four poets. Each is either by the poet or by an AI writing in their style. First say how much you like it, then guess: human or AI?

3 minto take part 1responses Anonymousno sign-up needed
?
EXP. 063

Can you tell lines by Shakespeare, Keats, Whitman and Dickinson from an AI writing in their style?

Loading the experiment… It needs JavaScript to run; you can still read the science box and the article below.

Play again, climb higher

Leaderboard.

  1. 1B@berke12/8
Only volunteers appear on the leaderboard. Your anonymous plays are linked to your account when you volunteer.Volunteer correct guesses · each person’s best play counts · “Today” resets every midnight (Istanbul time)
Take part first.

Reading the science could sway your answer. Everything unlocks once you take part, but you can read it now if you prefer.

Back to the experiment
Science box

Most readers can’t tell lines an AI wrote in a poet’s style from the poet’s own lines; they even like the AI poems more and call them human more often.

What we measure

We repeat Brian Porter and Édouard Machery’s (2024) “human or AI?” experiment with poetry in English and Turkish. You see two short excerpts for each of four poets: in English, Shakespeare, Keats, Whitman and Dickinson; in Turkish, Yunus Emre, Karacaoğlan, Namık Kemal and Tevfik Fikret. The pool has two real excerpts and two AI excerpts for each poet; the game’s seed decides which ones you get, so a poet’s two excerpts may both be human, both AI, or one of each. Human excerpts come only from public-domain works published before 1929; the AI excerpts were written for this experiment by an AI model (Claude, Anthropic) in each poet’s style, and we disclose all of them in the solution on the results screen. For each excerpt you first rate how much you like it from 1 to 5, then guess “human” or “AI”. Your score is the number of correct guesses out of eight; chance level is 4.

What the research says

Porter and Machery (2024) showed 1,634 participants five poems by one of ten English-language poets and five poems ChatGPT-3.5 wrote in the same poet’s style. Participants’ accuracy was 46.6%, slightly below chance (50%); they were more likely to say “written by a human” about AI poems than about the real poets’ poems. In a second study, 696 participants rated the poems on 14 qualities: AI poems scored higher on qualities such as rhythm and beauty, but when the same poem was presented as “written by AI”, its ratings fell. Köbis and Mossink (2021) had found that with poems written by an older model, GPT-2, people could spot the AI when its poems were chosen at random, but not when a human picked the best one.

Why it happens

According to Porter and Machery, readers use a shared but mistaken rule: they attribute poems that are easy to understand, melodic and direct about emotion to humans, and complex, hard-to-understand poems to AI. Yet AI poems are often more accessible, while real poets’ poems are more layered. In the end readers treat their own liking as evidence that a human wrote the poem. In their dataset 89% of the AI poems rhymed, against only 40% of the human poems; so rules like “AI can’t rhyme” are misleading too. In this experiment we ask about liking before the guess, to look at how liking relates to saying “human”.

Limitations

The excerpts are short (about four lines) and not whole poems; Porter and Machery used whole poems. One model wrote the AI excerpts and we chose them, so their quality depends on our prompting and selection. The human excerpts are by famous poets; if you recognise one, guessing becomes easier. Older texts are given in the spelling of the source edition; readers familiar with metre may pick up cues from form, so English and Turkish results may not be directly comparable. The score is a small sample of eight excerpts and shouldn’t be read as a measure of personal ability. This lab is part of an AI studio; we designed the experiment to show openly what AI can do and how readers perceive it.

46.6%Accuracy at telling AI poems from human poems (1,634 participants)Porter and Machery, 2024
50%Chance levelPorter and Machery, 2024
AI poemsRated higher on qualities such as rhythm and beautyPorter and Machery, 2024
fellRating of the same poem when told “written by AI”Porter and Machery, 2024
  1. Porter, B., & Machery, E. (2024). AI-generated poetry is indistinguishable from human-written poetry and is rated more favorably. Scientific Reports, 14, 26133. View source ↗
  2. Köbis, N., & Mossink, L. D. (2021). Artificial intelligence versus Maya Angelou: Experimental evidence that people cannot differentiate AI-generated from human-written poetry. Computers in Human Behavior, 114, 106553. View source ↗

Can you tell a poem’s author from its lines?

Shakespeare’s sonnets, Keats’s lush odes, Whitman’s long breathless lines, Dickinson’s dashes: these voices are so familiar that we feel we would know them anywhere. So what happens when an AI imitates the same voice?

In this experiment you read two short excerpts for each of four poets. Each excerpt is either from the poet’s own work or by an AI writing in their style. First you say how much you like the excerpt, then you guess who wrote it. On the results screen you can see the sources of all the excerpts and which one fooled the crowd most.

Porter and Machery’s experiment

Brian Porter and Édouard Machery (2024) collected five poems by each of ten English-language poets, from Chaucer to Dorothea Lasky, and asked ChatGPT-3.5 to write five poems in each poet’s style. They used the first five poems the AI produced as they were, without picking the best. Each of 1,634 participants read five real and five AI poems for one poet and guessed who wrote each one.

Accuracy was 46.6%: slightly below chance. More strikingly, participants said “written by a human” more often about the AI poems than about the real ones. The five poems least often judged human were all by real poets; four of the five most often judged human were written by AI.

We like them more, but look down on them once we know

In the second study, 696 participants rated the same poems on 14 qualities such as rhythm, imagery, beauty and meaning. The AI poems scored higher than the real poems on most of these qualities. But when participants were told a poem was written by AI, the ratings for that same poem fell.

The authors explain it like this: non-expert readers like AI poems, which convey emotions and ideas in more direct and understandable language; but they expect AI poetry to be bad. So they assume the poems they like were written by humans, and attribute the complex poems they struggle with to AI.

Older models, newer models

A few years ago the picture was different. Köbis and Mossink (2021) compared poems written with GPT-2 against real poems. When a random poem was chosen from the AI’s output, participants could tell it apart; but when a human selected the best one, they couldn’t. Participants also showed a slight aversion to AI-written poems.

Porter and Machery’s results show that with a newer model the distinction disappears even when no human makes a selection. The AI excerpts in this experiment were written by an AI model (Claude, Anthropic) and selected by us, which is closer to the “human in the loop” condition.

Why is an AI studio running this experiment?

ba creative labs is an AI studio and this lab is its showcase. The most honest way to show what AI can do is to give people an experiment they can test themselves on, and then explain everything. That is why we mark every AI excerpt in the solution and give the sources of the real ones.

Whatever the result, one thing doesn’t change: Shakespeare’s “Bare ruin’d choirs, where late the sweet birds sang” is the product of a life and a view of the world four centuries old. An AI can imitate that voice; being able to imitate it doesn’t make the voice worth less. Perhaps the real question isn’t “can we tell them apart?” but “why does telling them apart matter to us?”

FAQ

Where do the human excerpts come from?

Only from public-domain works published before 1929: Shakespeare, Keats, Whitman and Dickinson from Project Gutenberg editions; Yunus Emre, Karacaoğlan, Namık Kemal and Tevfik Fikret from texts on Turkish Wikisource. All sources are linked in the solution on the results screen.

Who wrote the AI excerpts?

An AI model (Claude, Anthropic) wrote them for this experiment in each poet’s style. None was copied from the poets’ real lines. On the results screen we show which excerpts are AI, one by one.

How is my score calculated?

It is the number of correct guesses out of eight: at least 0, at most 8. Someone tossing a coin gets 4 right on average. Your liking ratings don’t affect the score, but the crowd results show how liking relates to saying “human”.

Is one of each poet’s two excerpts always AI?

No. Each poet’s two excerpts are drawn at random from the pool: both may be by the poet, both by the AI, or one of each. So your answer for one excerpt doesn’t tell you the other.

Does this mean AI is as good as these poets?

No. Not being able to tell who wrote a short excerpt doesn’t show that two texts are of equal worth. Porter and Machery’s participants mostly didn’t read poetry, and according to the authors they attributed what was easy to understand to humans and what was complex to AI.

Discussion 0 comments

Join the discussionReading is open to everyone. Volunteer to comment and vote; it takes 20 seconds.
No comments yet.Be the first to write; a good question gets the discussion going.

Threads about this experiment

Start a new thread →
There’s no separate thread for this experiment yet. Start one to critique the method, share a paper or ask a new question.