Human-AI Co-Creativity: Advances, Opportunities, and Challenges
Adish Singla, Abhilasha Ravichander, Liwei Jiang, Alexander Spangher, Alice Oh, Jiho Jin, Jun Seong Kim, Changyoon Lee, Manh Hung Nguyen, Chao Wen
cs.CY
2026-09-08
A recap of the ICML 2026 human-AI co-creativity workshop: 80 submissions, 49 accepted, six invited talks. It organizes five research gaps and reports no new model or numbers.
Generative models now sit inside everyday design, writing, and scientific ideation. The upside is obvious: people can use a model as a brainstorming medium. The downside is equally concrete: homogenized outputs, design fixation, and muddy copyright and authorship. The ICML 2026 workshop in Seoul, organized by researchers at MPI-SWS, the University of Washington, Stanford, and KAIST, tried to gather the people who care about those tensions. This article is not a method paper. It is a map of the workshop: talks, accepted papers, and a short list of open problems.
Most model interfaces were built for automation. Dropped into open-ended creative work, they pin users to average answers. At population scale the same pull shows up as an “artificial hivemind”: outputs collapse both within a model and across models.
The call for papers ran on two tracks. One asked how generative models can support open-ended work: field reports, diversity methods, co-creation architectures, and evaluation. The other asked what breaks after the model enters a workflow: fixation, idea homogeneity, authenticity checks, and credit.
Reviewing was double-blind, with eight reviewers and at least three reviews per paper. Out of 80 submissions, 49 were accepted as non-archival reports, presented as posters plus lightning talks. Six invited talks ran up to 30 minutes each, plus a 60-minute panel. The talks were deliberately mixed. Zhao discussed technical debts that must be paid before generated content can sit beside human work, AI slop in music streaming, and Glaze as a defense for artists’ styles. Elliott asked what originality means once AI art is mainstream. Hong argued that industrial design should move from delegation to orchestration, and should insert friction, for example by casting the model as a naive mentee. Chung described post-training and user simulators that push language models from executing intents toward discovering them. Mehrotra treated creativity and learning as two faces of rebuilding a mental model. van der Plas argued that computational creativity is still under-studied in NLP.
The write-up then groups the papers and discussion into five directions: creativity metrics, creative-task benchmarks, improving generative diversity, authenticity and authorship, and human-AI co-creation interfaces.
There is no new model and no controlled experiment. The usable “results” are the workshop’s structure and the gaps it names.
| Item | Number |
| Submissions / accepted | 80 / 49 |
| Reviewers / reviews per paper | 8 / ≥3 |
| Invited talks | 6 × 30 min |
| Panel | 60 min |
| Research directions listed | 5 |
On metrics, the authors note that AUT, RAT, and TTCT were built for people; a high or low score does not automatically describe a model’s creative skill. Expert consensus scoring is closer to how creative work is judged, but aligning an automated rater with experts is still open. On benchmarks, current suites mostly probe combinatorial or exploratory creativity and almost never measure transformational creativity or a human in the loop. On methods, the same tension repeats: post-training raises task accuracy and collapses output diversity. On authenticity, the problem moves from “did a model write this” to process-level provenance: who contributed what inside a session, and how to keep the ledger when work is remixed later. On interaction, the cited evidence on fixation and over-reliance is used to argue for orchestration and intentional friction, not a smoother autocomplete.
If a product team is still scoring a writing or design assistant with MMLU-style accuracy, this recap locates the mismatch: the tasks people actually use models for are barely in the benchmarks. Diversity collapse after post-training, cross-model homogeneity, and detector bias against non-native writers already have papers behind them. The practical claim is sharper than a vibe: raising a model’s solo “creativity score” may not transfer to better co-creation. That transfer has to be measured on the joint process.
Treat the article as a community map. It does not ship a runnable system. It does make the bottleneck list short: evaluation, diversity, and credit.
This is a workshop recap, not a systematic survey. Most of the 49 papers appear as titles, so evidence quality is mixed. The five “future directions” are a wish list without ranking or a quantitative failure census. Claims about music-industry slop and the artificial hivemind are imported from cited work and are not re-run here. Non-archival workshop papers can change or vanish. The organizers also publish in this area, so the framing leans toward metrics and diversity training they already know.