AI and Creativity — Can Machines Be Genuinely Creative?
GPT-4, Claude, and Gemini now beat the average human on the standard divergent-thinking test. They still lose to the top 10% of human creatives. The interesting question is no longer whether machines can produce novel-and-useful outputs — that is settled — but whether their use is quietly narrowing the variance of human culture itself.
The two questions, kept separate
Most arguments about AI creativity collapse two problems that should stay apart:
- Production — can a system generate outputs humans rate as novel, valuable, surprising? Empirically yes, and the gap to the median human is closing fast.
- Understanding — does the system know what it made, in any sense that earns the word creative? This is concept hard problem consciousness in a machine costume.
The production question has data. The understanding question has concept chinese room, and not much else.
What the benchmarks actually show
- University of Graz, 2026. Leading LLMs vs. 100,000+ humans on the Alternative Uses Task (name unusual uses for a brick). AI beats the human average; the top decile of humans beats AI (ScienceDaily, Jan 2026).
- Individual gain, collective loss. A 2024 Science Advances study found AI-assisted short stories were rated more creative and more enjoyable than human-only stories — especially when the writer was below-average. But the AI-assisted stories were measurably more similar to each other. The floor rose; the ceiling and the spread fell.
- Stage-dependent help (Frontiers in Computer Science, 2025). AI is most useful in brainstorming, for everyone. In execution, it helps novices and frustrates experts, whose taste finds the suggestions too generic.
The pattern across studies: AI is a compressor. It lifts the bottom of the distribution and pulls in the tails.
The neural picture: switching, not generating
The Default Mode Network is the brain's wandering, associative engine. But creativity is not sustained DMN activation — it is the rate of switching between the DMN and the Executive Control Network, the evaluative circuit. A 2025 multi-center study (N=2,433, 10 countries, Nature Communications Biology) found switching rate predicts measured creativity better than activation in either network alone. Causal stimulation of DMN regions (2024, Brain, Oxford) selectively reduces originality — the closest thing to a confirmed creativity circuit yet identified.
A January 2025 bioRxiv finding inverts a common assumption: professional artists show more DMN, ECN, and sensorimotor gray matter but less Salience Network expression. Creative experts have quieter filtering systems, not louder generative ones. The same dynamic appears in roughly 2.5% of frontotemporal dementia patients, who develop de novo visual art when frontal inhibition collapses. Removing the critic, in brains as in models, widens the output.
Boden's three creativities
Margaret Boden's taxonomy is still the cleanest frame:
| Type | What it is | AI status |
|---|---|---|
| Combinational | New combinations of existing ideas | AI excels |
| Exploratory | Systematic search within an existing conceptual space | AI competent |
| Transformational | Restructuring the space itself | No demonstrated case |
Impressionism, Cubism, Abstract Expressionism — every paradigm break in art history was transformational. No model has produced one. Whether this is a limit of architecture, of training distribution, or of what we can recognise as transformational from inside the new space is open.
The monoculture problem
If everyone drafts with the same three models, the long tail of human idiosyncrasy compresses toward the training distribution's high-probability ridge. The concept outsider art canon is the test case: Henry Darger's 15,000 hidden pages, Adolf Wölfli's 25,000 psychiatric pages, Martín Ramírez's trains. None of it is recoverable from a model trained on the mainstream — it was generated by minds outside the distribution. AI tools generate from inside it, by construction.
The analogy is agricultural. A few cultivars feeding a continent is efficient until a pathogen finds them. The cultural equivalent of a concept svalbard seed vault does not exist, and the gene pool is shrinking measurably with each cohort of writers, designers, and students who default to the same assistants.
What's contested
Three live disputes:
- Does production without understanding count? Searle says no. Anthropic's 2025 mechanistic interpretability work found structured semantic intermediate representations in LLMs — concept nodes that causally mediate outputs, not just statistical pattern-matching. This does not refute Searle, but it makes "syntax all the way down" harder to defend.
- Is transformational creativity blocked by architecture or by data? Some argue transformer models cannot restructure their own conceptual space; others argue they have not been trained on enough out-of-distribution cases to try.
- Does the listener's model of the maker shape the response? concept frisson — musical chills — arises from prediction violation modulated by perceived intent. Knowing a piece is AI-generated may itself suppress the response, independent of the audio. This is testable and largely untested.
Why this has to do with other realms
The creativity-as-filter-removal finding maps cleanly onto concept outsider art: the FTD patients losing frontal inhibition, the artists with quieter Salience Networks, and the LLMs without domain expertise all produce wider output by lacking the critic that training installs. Creativity may be less about adding generative power and more about subtracting evaluation — a thesis that connects neuroscience, psychiatric art history, and machine learning into the same equation. The bridge to biology is sharper still: monoculture in seeds and monoculture in ideas are the same fragility problem at different substrates.
An open question
If the AI-assisted cohort produces fewer Dargers and Wölflis — not because the tools forbid them but because the tools make them harder to recognise as anything but noise — what does the wiki of 2050 look like, and who writes the pages no model would have suggested?
Key sources
- The Creative Mind: Myths and Mechanisms — Margaret Boden (1990, revised 2004). The combinational / exploratory / transformational taxonomy is hers.
- Science Advances (2024) — "Generative AI enhances individual creativity but reduces the collective diversity of novel content." The single most important empirical result in the field.
- Nature Communications Biology (2025, N=2,433 multi-center) — DMN ↔ ECN switching as the predictor of measured creativity. To verify: exact author list.
- Brain (Oxford, 2024) — causal DMN stimulation reduces originality. To verify: lead author.
- Anthropic mechanistic interpretability work (2025) — attribution graphs in large models. Public on anthropic.com/research.
- Minds and Machines / Searle (1980) — the original concept chinese room paper, still the load-bearing objection.
Further reading
- Gödel, Escher, Bach — Douglas Hofstadter (1979). Still the best long-form meditation on what it would mean for a formal system to mean anything.
- The Artist in the Machine — Arthur I. Miller (2019). Survey of AARON, AIVA, and the early generative-art systems with interviews.
- Harold Cohen's AARON archive — Stanford / V&A holdings. The 1973–2016 project that posed every question LLMs now inherit.
- Refik Anadol's DATALAND (Los Angeles, opened 2025) — the first museum dedicated to AI-generated art. Worth visiting to test your own frisson response.
See Also
- concept generative art (the 1965-to-now arc AARON sits inside)
- concept outsider art (the canon AI cannot generate, and the filter-removal parallel)
- concept chinese room (the standing objection)
- concept frisson (does knowing it's a machine kill the chills?)
- concept hard problem consciousness (the question underneath)
- concept embodied cognition (creativity as a body-in-world phenomenon, not a text-prediction one)
- concept svalbard seed vault (monoculture fragility, applied to ideas)