Article Contents
REVIEW   Open Access     Cite

Extracting creativity from hallucination: rethinking large language models

    Show all affliationsShow less
More Information
  • †These authors contributed equally

  • Corresponding authors: wangyuanzhuo@ict.ac.cn (Y.W.);  guojian@idea.edu.cn (J.G.)
  • DownLoad: Full size image
    1. Large language model (LLM) hallucinations can be errors in factual tasks but may aid openended ideation.

      Creative value arises only after hallucination-derived candidates pass tests of novelty, usefulness, and task fit.

      Divergent methods expand the candidate space, while convergent methods filter risks and select useful outputs.

      This review maps methods, evaluation criteria, application domains, and responsible-use boundaries.

  • Hallucinations in large language models (LLMs) are always seen as limitations. However, could they also be a source of creativity? This survey explores this possibility, suggesting that hallucinations may contribute to LLM application by fostering creativity. Hallucinations are not treated as creativity perse; rather, they are considered candidate materials that require novelty, usefulness, task alignment, semantic structure, safety boundaries, and explicit verification before they can contribute to creative outcomes. This survey begins with a review of the taxonomy of hallucinations and their negative impact on LLM reliability in critical applications. Then, through historical examples and recent relevant theories, the survey explores the potential creative benefits of hallucinations in LLMs. To elucidate the value and evaluation criteria of this connection, we delve into the definitions and assessment methods of creativity. Following the framework of divergent and convergent thinking phases, the survey systematically reviews the literature on transforming and harnessing hallucinations for creativity in LLMs. Finally, the survey discusses future research directions, emphasizing the need to further explore and refine the application of hallucinations in creative processes within LLMs.
  • 加载中
  • [1] Ye H., Liu T., Zhang A., et al. (2023). Cognitive mirage: A review of hallucinations in large language models. arXiv preprint arXiv: 2309.06794.

    View in Article Google Scholar

    [2] Sui P., Duede E., Wu S., et al. (2024). Confabulation: The Surprising Value of Large Language Model Hallucinations. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 14274−84. DOI:10.18653/v1/2024.acl-long.770

    View in Article CrossRef Google Scholar

    [3] Zhao Y., Zhang R., Li W., et al. (2025). Assessing and understanding creativity in large language models. Machine Intelligence Research 22:417−36. DOI:10.1007/s11633-025-1546-4

    View in Article CrossRef Google Scholar

    [4] Touvron H., Lavril T., Izacard G., et al. (2023). Llama: Open and efficient foundation language models. arXiv preprint arXiv: 2302.13971.

    View in Article Google Scholar

    [5] Ngo R., Chan L. and Mindermann S. (2024). The alignment problem from a deep learning perspective. International Conference on Learning Representations 2024:7474−501.

    View in Article Google Scholar

    [6] Ji Z., Lee N., Frieske R., et al. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys.

    View in Article Google Scholar

    [7] Dziri N., Madotto A., Zaïane O., et al. (2021). Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path Grounding. Proc. of EMNLP.

    View in Article Google Scholar

    [8] Huang L., Yu W., Ma W., et al. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems 43:1−55. DOI:10.1145/3703155

    View in Article CrossRef Google Scholar

    [9] Guilford J. P. (2017). Creativity: A quarter century of progress. Perspectives in creativity 37−59.

    View in Article Google Scholar

    [10] Pressing J. (1998). Psychological constraints on improvisational expertise and communication. In the course of performance: Studies in the world of musical improvisation.

    View in Article Google Scholar

    [11] Liang T., He Z., Jiao W., et al. (2024). Encouraging divergent thinking in large language models through multi-agent debate. Proceedings of the 2024 conference on empirical methods in natural language processing 17889−904.

    View in Article Google Scholar

    [12] Dhuliawala S., Komeili M., Xu J., et al. (2024). Chain-of-Verification Reduces Hallucination in Large Language Models. Findings of the Association for Computational Linguistics: ACL 2024 3563−78. DOI:10.18653/v1/2024.findings-acl.212

    View in Article CrossRef Google Scholar

    [13] Manakul P., Liusie A. and Gales M. J. F. (2023). SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing 9004−17. DOI:10.18653/v1/2023.emnlp-main.557

    View in Article CrossRef Google Scholar

    [14] Rawte V., Sheth A. and Das A. (2023). A survey of hallucination in large foundation models. ArXiv preprint.

    View in Article Google Scholar

    [15] Li J., Cheng X., Zhao W. X., et al. (2023). Halueval: A large-scale hallucination evaluation benchmark for large language models. Proc. of EMNLP.

    View in Article Google Scholar

    [16] Lin S., Hilton J. and Evans O. (2022). TruthfulQA: Measuring How Models Mimic Human Falsehoods. Proc. of ACL.

    View in Article Google Scholar

    [17] Pal A., Umapathi L. K. and Sankarasubbu M. (2023). Med-halt: Medical domain hallucination test for large language models. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL) 314−34.

    View in Article Google Scholar

    [18] Tonmoy S., Zaman S., Jain V., et al. (2024). A comprehensive survey of hallucination mitigation techniques in large language models. arXiv preprint arXiv: 2401.01313.

    View in Article Google Scholar

    [19] Agrawal G., Kumarage T., Alghamdi Z., et al. (2024). Can knowledge graphs reduce hallucinations in llms?: A survey. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) 3947−60.

    View in Article Google Scholar

    [20] Shi X., Zhu Z., Zhang Z., et al. (2023). Hallucination Mitigation in Natural Language Generation from Large-Scale Open-Domain Knowledge Graphs. Proc. of EMNLP.

    View in Article Google Scholar

    [21] Sui Y., He Y., Ding Z., et al. (2025). Can knowledge graphs make large language models more trustworthy? an empirical study over open-ended question answering. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 12685−701.

    View in Article Google Scholar

    [22] Zhang Y., Li Y., Cui L., et al. (2025). Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models. Computational Linguistics 51:1373−418. DOI:10.1162/COLI.a.16

    View in Article CrossRef Google Scholar

    [23] National Institute of Standards and Technology (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. DOI: 10.6028/NIST.AI.600-1

    View in Article Google Scholar

    [24] Benedek M., Jauk E., Fink A., et al. (2014). To create or to recall? Neural mechanisms underlying the generation of creative new ideas. NeuroImage.

    View in Article Google Scholar

    [25] Beaty R. E. (2015). The neuroscience of musical improvisation. Neuroscience & Biobehavioral Reviews 51:108−17. DOI:10.1016/j.neubiorev.2015.01.004

    View in Article CrossRef Google Scholar

    [26] Salesforce (2023). A world without AI is becoming unthinkable.

    View in Article Google Scholar

    [27] Lee M. (2023). A mathematical investigation of hallucination and creativity in gpt models. Mathematics.

    View in Article Google Scholar

    [28] Bostrom K. and Durrett G. (2020). Byte Pair Encoding is Suboptimal for Language Model Pretraining. Findings of the Association for Computational Linguistics: EMNLP 2020 4617−24. DOI:10.18653/v1/2020.findings-emnlp.414

    View in Article CrossRef Google Scholar

    [29] Wang F. (2024). LightHouse: A Survey of AGI Hallucination. arXiv preprint arXiv: 2401.06792.

    View in Article Google Scholar

    [30] He Z., Zhang B. and Cheng L. (2025). Shakespearean Sparks: The Dance of Hallucination and Creativity in LLMs' Decoding Layers. arXiv preprint arXiv: 2503.02851.

    View in Article Google Scholar

    [31] Treffinger D. J. (1998). Creativity, Creative Thinking, and Critical Thinking: In Search of Definitions. Gifted and Talented International.

    View in Article Google Scholar

    [32] Couger J. D., Higgins L. F. and McIntyre S. C. (1993). (Un)Structured Creativity in Information Systems Organizations. MIS Q..

    View in Article Google Scholar

    [33] Torrance E. P. (1977). Creativity in the Classroom; What Research Says to the Teacher.

    View in Article Google Scholar

    [34] Mednick S. A. (1962). The associative basis of the creative process. Psychological review.

    View in Article Google Scholar

    [35] Khatena J. and Torrance E. P. (1973). Thinking creatively with sounds and words: Normstechnical manual. Res. ed.) Bensenville, IL: Scholastic Testing Service.

    View in Article Google Scholar

    [36] Gardner H. (2011). Creating minds: An anatomy of creativity seen through the lives of Freud, Einstein, Picasso, Stravinsky, Eliot, Graham, and Ghandi. (Civitas books).

    View in Article Google Scholar

    [37] Kaufman J. C. and Beghetto R. A. (2009). Beyond Big and Little: The Four C Model of Creativity. Review of General Psychology.

    View in Article Google Scholar

    [38] Olson J. A., Nahas J., Chmoulevitch D., et al. (2021). Naming unrelated words predicts creativity. Proceedings of the National Academy of Sciences.

    View in Article Google Scholar

    [39] Marko M., Michalko D. and Rievcansky I. (2018). Remote associates test: An empirical proof of concept. Behavior Research Methods.

    View in Article Google Scholar

    [40] Gianotti L. R. R., Mohr C., Pizzagalli D. A., et al. (2001). Associative processing and paranormal belief. Psychiatry and Clinical Neurosciences.

    View in Article Google Scholar

    [41] Amabile T. M. (1982). Social psychology of creativity: A consensual assessment technique. Journal of Personality and Social Psychology.

    View in Article Google Scholar

    [42] Baer J. (2015). The Importance of Domain-Specific Expertise in Creativity. Roeper Review.

    View in Article Google Scholar

    [43] Niu W. and Sternberg R. J. (2001). CULTURAL INFLUENCES ON ARTISTIC CREATIVITY AND ITS EVALUATION. International Journal of Psychology.

    View in Article Google Scholar

    [44] Stevenson C., Smal I., Baas M., et al. (2022). Putting GPT-3's Creativity to the Alternative Uses Test.

    View in Article Google Scholar

    [45] Summers-Stay D., Voss C. R. and Lukin S. M. (2023). Brainstorm, then Select: a Generative Language Model Improves Its Creativity Score. The AAAI-23 Workshop on Creative AI Across Modalities.

    View in Article Google Scholar

    [46] Cropley D. (2023). Is artificial intelligence more creative than humans? : ChatGPT and the Divergent Association Task. Learning Letters.

    View in Article Google Scholar

    [47] Góes L. F., Volpe M., Sawicki P., et al. (2023). Pushing gpt’s creativity to its limits: Alternative uses and torrance tests.

    View in Article Google Scholar

    [48] Guzik E. E., Byrge C. and Gilde C. (2023). The originality of machines: AI takes the Torrance Test. Journal of Creativity.

    View in Article Google Scholar

    [49] Wenger E. and Kenett Y. N. (2026). Large language models are homogeneously creative. PNAS nexus 5:pgag042. DOI:10.1093/pnasnexus/pgag042

    View in Article CrossRef Google Scholar

    [50] Wang H., Zou J., Mozer M., et al. (2024). Can AI be as creative as humans?. arXiv preprint arXiv: 2401.01623.

    View in Article Google Scholar

    [51] Chuang Y. S., Xie Y., Luo H., et al. (2024). DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models. International Conference on Learning Representations.

    View in Article Google Scholar

    [52] Pressing J. (2007). Improvisation: Methods and models. Physical Theatres: A Critical Reader 66−78.

    View in Article Google Scholar

    [53] Tian Y., Ravichander A., Qin L., et al. (2024). MacGyver: Are Large Language Models Creative Problem Solvers?. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) 5303−24.

    View in Article Google Scholar

    [54] Colin T. R., Belpaeme T., Cangelosi A., et al. (2016). Hierarchical reinforcement learning as creative problem solving. Robotics and Autonomous Systems.

    View in Article Google Scholar

    [55] Yuan S., Lyu X., Wang S., et al. (2025). FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models. Advances in Neural Information Processing Systems.

    View in Article Google Scholar

    [56] Beaty R., Smeekens B., Silvia P., et al. (2013). A First Look at the Role of Domain-General Cognitive and Creative Abilities in Jazz Improvisation. Psychomusicology: Music, Mind, & Brain.

    View in Article Google Scholar

    [57] Hubert K., Awa K. and Zabelina D. (2023). Artificial intelligence is more creative than humans: A cognitive science perspective on the current state of generative language models.

    View in Article Google Scholar

    [58] Koivisto M. and Grassini S. (2023). Best humans still outperform artificial intelligence in a creative divergent thinking task. Scientific reports.

    View in Article Google Scholar

    [59] Weber S., Kordyaka B., Palombo R., et al. (2023). Is a Fool With a (n AI) Tool Still a Fool? An Empirical Study of the Creative Quality of Human–AI Collaboration. ACIS 2023 Proceedings.

    View in Article Google Scholar

    [60] Rick S. R., Giacomelli G., Wen H., et al. (2023). Supermind Ideator: Exploring generative AI to support creative problem-solving. ArXiv preprint.

    View in Article Google Scholar

    [61] Zhang C. (2023). User-controlled knowledge fusion in large language models: Balancing creativity and hallucination. arXiv preprint arXiv: 2307.16139.

    View in Article Google Scholar

    [62] Banerjee M., Wangsajaya N. Y., Alsagoff S. A. R., et al. (2025). Does Less Hallucination Mean Less Creativity? An Empirical Investigation in LLMs. arXiv preprint arXiv: 2512.11509.

    View in Article Google Scholar

    [63] Argese A., Lisena P. and Troncy R. (2026). Hallucination or Creativity: How to Evaluate AI-Generated Scientific Stories?. arXiv preprint arXiv: 2602.02290.

    View in Article Google Scholar

    [64] Yuan S., Qu Z., Kangen A. Y., et al. (2025). Can Hallucinations Help? Boosting LLMs for Drug Discovery. arXiv preprint arXiv: 2501.13824.

    View in Article Google Scholar

    [65] Silvia P. J., Winterstein B., Willse J. T., et al. (2008). Assessing creativity with divergent thinking tasks: exploring the reliability and validity of new subjective scoring methods. Psychology of Aesthetics, Creativity, and the Arts.

    View in Article Google Scholar

    [66] Chan C. M., Chen W., Su Y., et al. (2024). Chateval: Towards better llm-based evaluators through multi-agent debate. International conference on learning representations 2024:9079−93.

    View in Article Google Scholar

    [67] Zhang L., Zhang M., Wang W. L., et al. (2025). Simulation as Reality? The Effectiveness of LLM-Generated Data in Open-ended Question Assessment. arXiv preprint arXiv: 2502.06371.

    View in Article Google Scholar

    [68] Karpowicz M. and others (2025). On the Fundamental Impossibility of Hallucination Control in Large Language Models. arXiv preprint arXiv: 2506.06382.

    View in Article Google Scholar

    [69] Qiu Z. and Hu R. (2025). Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing 10870−83.

    View in Article Google Scholar

    [70] Hu Y., Liu B., Kasai J., et al. (2023). TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering. Proceedings of the IEEE/CVF International Conference on Computer Vision 20406−17.

    View in Article Google Scholar

    [71] Lim Y., Choi H. and Shim H. (2025). Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering. Proceedings of the AAAI Conference on Artificial Intelligence 39:26290−8. DOI:10.1609/aaai.v39i25.34827

    View in Article CrossRef Google Scholar

    [72] Huang K., Duan C., Sun K., et al. (2025). T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-Image Generation. IEEE Transactions on Pattern Analysis and Machine Intelligence 47:3563−79. DOI:10.1109/TPAMI.2025.3531907

    View in Article CrossRef Google Scholar

    [73] Lin Z., Pathak D., Li B., et al. (2024). Evaluating Text-to-Visual Generation with Image-to-Text Generation. European Conference on Computer Vision.

    View in Article Google Scholar

    [74] Liu Y., Cun X., Liu X., et al. (2024). EvalCrafter: Benchmarking and Evaluating Large Video Generation Models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 22139−49.

    View in Article Google Scholar

    [75] Huang Z., He Y., Yu J., et al. (2024). VBench: Comprehensive Benchmark Suite for Video Generative Models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 21807−18.

    View in Article Google Scholar

    [76] Zheng D., Huang Z., Liu H., et al. (2025). VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness. arXiv preprint arXiv: 2503.21755.

    View in Article Google Scholar

    [77] Liu J., Liu G., Liang J., et al. (2025). Improving Video Generation with Human Feedback. Advances in Neural Information Processing Systems.

    View in Article Google Scholar

  • Cite this article:

    Jiang X., Liu Y., Shen Y., et al. (2026). Extracting creativity from hallucination: rethinking large language models. AI Plus 1:100008. https://doi.org/10.59717/ipj.aiplus.2026.100008
    Jiang X., Liu Y., Shen Y., et al. (2026). Extracting creativity from hallucination: rethinking large language models. AI Plus 1:100008. https://doi.org/10.59717/ipj.aiplus.2026.100008

Welcome!

To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.

Figures(4)     Tables(3)

Share

  • Share the QR code with wechat scanning code to friends and circle of friends.

Article Metrics

Article views(239) PDF downloads(99)

Relative Articles

Cited by

Catalog

    /

    DownLoad:  Full-Size Img  PowerPoint