Large language model (LLM) hallucinations can be errors in factual tasks but may aid openended ideation.
Creative value arises only after hallucination-derived candidates pass tests of novelty, usefulness, and task fit.
Divergent methods expand the candidate space, while convergent methods filter risks and select useful outputs.
This review maps methods, evaluation criteria, application domains, and responsible-use boundaries.
| [1] | Ye H., Liu T., Zhang A., et al. (2023). Cognitive mirage: A review of hallucinations in large language models. arXiv preprint arXiv: 2309.06794. |
| [2] | Sui P., Duede E., Wu S., et al. (2024). Confabulation: The Surprising Value of Large Language Model Hallucinations. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 14274−84. DOI:10.18653/v1/2024.acl-long.770 |
| [3] | Zhao Y., Zhang R., Li W., et al. (2025). Assessing and understanding creativity in large language models. Machine Intelligence Research 22:417−36. DOI:10.1007/s11633-025-1546-4 |
| [4] | Touvron H., Lavril T., Izacard G., et al. (2023). Llama: Open and efficient foundation language models. arXiv preprint arXiv: 2302.13971. |
| [5] | Ngo R., Chan L. and Mindermann S. (2024). The alignment problem from a deep learning perspective. International Conference on Learning Representations 2024:7474−501. |
| [6] | Ji Z., Lee N., Frieske R., et al. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys. |
| [7] | Dziri N., Madotto A., Zaïane O., et al. (2021). Neural Path Hunter: Reducing Hallucination in Dialogue Systems via Path Grounding. Proc. of EMNLP. |
| [8] | Huang L., Yu W., Ma W., et al. (2025). A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Transactions on Information Systems 43:1−55. DOI:10.1145/3703155 |
| [9] | Guilford J. P. (2017). Creativity: A quarter century of progress. Perspectives in creativity 37−59. |
| [10] | Pressing J. (1998). Psychological constraints on improvisational expertise and communication. In the course of performance: Studies in the world of musical improvisation. |
| [11] | Liang T., He Z., Jiao W., et al. (2024). Encouraging divergent thinking in large language models through multi-agent debate. Proceedings of the 2024 conference on empirical methods in natural language processing 17889−904. |
| [12] | Dhuliawala S., Komeili M., Xu J., et al. (2024). Chain-of-Verification Reduces Hallucination in Large Language Models. Findings of the Association for Computational Linguistics: ACL 2024 3563−78. DOI:10.18653/v1/2024.findings-acl.212 |
| [13] | Manakul P., Liusie A. and Gales M. J. F. (2023). SelfCheckGPT: Zero-Resource Black-Box Hallucination Detection for Generative Large Language Models. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing 9004−17. DOI:10.18653/v1/2023.emnlp-main.557 |
| [14] | Rawte V., Sheth A. and Das A. (2023). A survey of hallucination in large foundation models. ArXiv preprint. |
| [15] | Li J., Cheng X., Zhao W. X., et al. (2023). Halueval: A large-scale hallucination evaluation benchmark for large language models. Proc. of EMNLP. |
| [16] | Lin S., Hilton J. and Evans O. (2022). TruthfulQA: Measuring How Models Mimic Human Falsehoods. Proc. of ACL. |
| [17] | Pal A., Umapathi L. K. and Sankarasubbu M. (2023). Med-halt: Medical domain hallucination test for large language models. Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL) 314−34. |
| [18] | Tonmoy S., Zaman S., Jain V., et al. (2024). A comprehensive survey of hallucination mitigation techniques in large language models. arXiv preprint arXiv: 2401.01313. |
| [19] | Agrawal G., Kumarage T., Alghamdi Z., et al. (2024). Can knowledge graphs reduce hallucinations in llms?: A survey. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) 3947−60. |
| [20] | Shi X., Zhu Z., Zhang Z., et al. (2023). Hallucination Mitigation in Natural Language Generation from Large-Scale Open-Domain Knowledge Graphs. Proc. of EMNLP. |
| [21] | Sui Y., He Y., Ding Z., et al. (2025). Can knowledge graphs make large language models more trustworthy? an empirical study over open-ended question answering. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) 12685−701. |
| [22] | Zhang Y., Li Y., Cui L., et al. (2025). Siren’s Song in the AI Ocean: A Survey on Hallucination in Large Language Models. Computational Linguistics 51:1373−418. DOI:10.1162/COLI.a.16 |
| [23] | National Institute of Standards and Technology (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile. DOI: 10.6028/NIST.AI.600-1 |
| [24] | Benedek M., Jauk E., Fink A., et al. (2014). To create or to recall? Neural mechanisms underlying the generation of creative new ideas. NeuroImage. |
| [25] | Beaty R. E. (2015). The neuroscience of musical improvisation. Neuroscience & Biobehavioral Reviews 51:108−17. DOI:10.1016/j.neubiorev.2015.01.004 |
| [26] | Salesforce (2023). A world without AI is becoming unthinkable. |
| [27] | Lee M. (2023). A mathematical investigation of hallucination and creativity in gpt models. Mathematics. |
| [28] | Bostrom K. and Durrett G. (2020). Byte Pair Encoding is Suboptimal for Language Model Pretraining. Findings of the Association for Computational Linguistics: EMNLP 2020 4617−24. DOI:10.18653/v1/2020.findings-emnlp.414 |
| [29] | Wang F. (2024). LightHouse: A Survey of AGI Hallucination. arXiv preprint arXiv: 2401.06792. |
| [30] | He Z., Zhang B. and Cheng L. (2025). Shakespearean Sparks: The Dance of Hallucination and Creativity in LLMs' Decoding Layers. arXiv preprint arXiv: 2503.02851. |
| [31] | Treffinger D. J. (1998). Creativity, Creative Thinking, and Critical Thinking: In Search of Definitions. Gifted and Talented International. |
| [32] | Couger J. D., Higgins L. F. and McIntyre S. C. (1993). (Un)Structured Creativity in Information Systems Organizations. MIS Q.. |
| [33] | Torrance E. P. (1977). Creativity in the Classroom; What Research Says to the Teacher. |
| [34] | Mednick S. A. (1962). The associative basis of the creative process. Psychological review. |
| [35] | Khatena J. and Torrance E. P. (1973). Thinking creatively with sounds and words: Normstechnical manual. Res. ed.) Bensenville, IL: Scholastic Testing Service. |
| [36] | Gardner H. (2011). Creating minds: An anatomy of creativity seen through the lives of Freud, Einstein, Picasso, Stravinsky, Eliot, Graham, and Ghandi. (Civitas books). |
| [37] | Kaufman J. C. and Beghetto R. A. (2009). Beyond Big and Little: The Four C Model of Creativity. Review of General Psychology. |
| [38] | Olson J. A., Nahas J., Chmoulevitch D., et al. (2021). Naming unrelated words predicts creativity. Proceedings of the National Academy of Sciences. |
| [39] | Marko M., Michalko D. and Rievcansky I. (2018). Remote associates test: An empirical proof of concept. Behavior Research Methods. |
| [40] | Gianotti L. R. R., Mohr C., Pizzagalli D. A., et al. (2001). Associative processing and paranormal belief. Psychiatry and Clinical Neurosciences. |
| [41] | Amabile T. M. (1982). Social psychology of creativity: A consensual assessment technique. Journal of Personality and Social Psychology. |
| [42] | Baer J. (2015). The Importance of Domain-Specific Expertise in Creativity. Roeper Review. |
| [43] | Niu W. and Sternberg R. J. (2001). CULTURAL INFLUENCES ON ARTISTIC CREATIVITY AND ITS EVALUATION. International Journal of Psychology. |
| [44] | Stevenson C., Smal I., Baas M., et al. (2022). Putting GPT-3's Creativity to the Alternative Uses Test. |
| [45] | Summers-Stay D., Voss C. R. and Lukin S. M. (2023). Brainstorm, then Select: a Generative Language Model Improves Its Creativity Score. The AAAI-23 Workshop on Creative AI Across Modalities. |
| [46] | Cropley D. (2023). Is artificial intelligence more creative than humans? : ChatGPT and the Divergent Association Task. Learning Letters. |
| [47] | Góes L. F., Volpe M., Sawicki P., et al. (2023). Pushing gpt’s creativity to its limits: Alternative uses and torrance tests. |
| [48] | Guzik E. E., Byrge C. and Gilde C. (2023). The originality of machines: AI takes the Torrance Test. Journal of Creativity. |
| [49] | Wenger E. and Kenett Y. N. (2026). Large language models are homogeneously creative. PNAS nexus 5:pgag042. DOI:10.1093/pnasnexus/pgag042 |
| [50] | Wang H., Zou J., Mozer M., et al. (2024). Can AI be as creative as humans?. arXiv preprint arXiv: 2401.01623. |
| [51] | Chuang Y. S., Xie Y., Luo H., et al. (2024). DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models. International Conference on Learning Representations. |
| [52] | Pressing J. (2007). Improvisation: Methods and models. Physical Theatres: A Critical Reader 66−78. |
| [53] | Tian Y., Ravichander A., Qin L., et al. (2024). MacGyver: Are Large Language Models Creative Problem Solvers?. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) 5303−24. |
| [54] | Colin T. R., Belpaeme T., Cangelosi A., et al. (2016). Hierarchical reinforcement learning as creative problem solving. Robotics and Autonomous Systems. |
| [55] | Yuan S., Lyu X., Wang S., et al. (2025). FlexAC: Towards Flexible Control of Associative Reasoning in Multimodal Large Language Models. Advances in Neural Information Processing Systems. |
| [56] | Beaty R., Smeekens B., Silvia P., et al. (2013). A First Look at the Role of Domain-General Cognitive and Creative Abilities in Jazz Improvisation. Psychomusicology: Music, Mind, & Brain. |
| [57] | Hubert K., Awa K. and Zabelina D. (2023). Artificial intelligence is more creative than humans: A cognitive science perspective on the current state of generative language models. |
| [58] | Koivisto M. and Grassini S. (2023). Best humans still outperform artificial intelligence in a creative divergent thinking task. Scientific reports. |
| [59] | Weber S., Kordyaka B., Palombo R., et al. (2023). Is a Fool With a (n AI) Tool Still a Fool? An Empirical Study of the Creative Quality of Human–AI Collaboration. ACIS 2023 Proceedings. |
| [60] | Rick S. R., Giacomelli G., Wen H., et al. (2023). Supermind Ideator: Exploring generative AI to support creative problem-solving. ArXiv preprint. |
| [61] | Zhang C. (2023). User-controlled knowledge fusion in large language models: Balancing creativity and hallucination. arXiv preprint arXiv: 2307.16139. |
| [62] | Banerjee M., Wangsajaya N. Y., Alsagoff S. A. R., et al. (2025). Does Less Hallucination Mean Less Creativity? An Empirical Investigation in LLMs. arXiv preprint arXiv: 2512.11509. |
| [63] | Argese A., Lisena P. and Troncy R. (2026). Hallucination or Creativity: How to Evaluate AI-Generated Scientific Stories?. arXiv preprint arXiv: 2602.02290. |
| [64] | Yuan S., Qu Z., Kangen A. Y., et al. (2025). Can Hallucinations Help? Boosting LLMs for Drug Discovery. arXiv preprint arXiv: 2501.13824. |
| [65] | Silvia P. J., Winterstein B., Willse J. T., et al. (2008). Assessing creativity with divergent thinking tasks: exploring the reliability and validity of new subjective scoring methods. Psychology of Aesthetics, Creativity, and the Arts. |
| [66] | Chan C. M., Chen W., Su Y., et al. (2024). Chateval: Towards better llm-based evaluators through multi-agent debate. International conference on learning representations 2024:9079−93. |
| [67] | Zhang L., Zhang M., Wang W. L., et al. (2025). Simulation as Reality? The Effectiveness of LLM-Generated Data in Open-ended Question Assessment. arXiv preprint arXiv: 2502.06371. |
| [68] | Karpowicz M. and others (2025). On the Fundamental Impossibility of Hallucination Control in Large Language Models. arXiv preprint arXiv: 2506.06382. |
| [69] | Qiu Z. and Hu R. (2025). Deep Associations, High Creativity: A Simple yet Effective Metric for Evaluating Large Language Models. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing 10870−83. |
| [70] | Hu Y., Liu B., Kasai J., et al. (2023). TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering. Proceedings of the IEEE/CVF International Conference on Computer Vision 20406−17. |
| [71] | Lim Y., Choi H. and Shim H. (2025). Evaluating Image Hallucination in Text-to-Image Generation with Question-Answering. Proceedings of the AAAI Conference on Artificial Intelligence 39:26290−8. DOI:10.1609/aaai.v39i25.34827 |
| [72] | Huang K., Duan C., Sun K., et al. (2025). T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-Image Generation. IEEE Transactions on Pattern Analysis and Machine Intelligence 47:3563−79. DOI:10.1109/TPAMI.2025.3531907 |
| [73] | Lin Z., Pathak D., Li B., et al. (2024). Evaluating Text-to-Visual Generation with Image-to-Text Generation. European Conference on Computer Vision. |
| [74] | Liu Y., Cun X., Liu X., et al. (2024). EvalCrafter: Benchmarking and Evaluating Large Video Generation Models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 22139−49. |
| [75] | Huang Z., He Y., Yu J., et al. (2024). VBench: Comprehensive Benchmark Suite for Video Generative Models. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition 21807−18. |
| [76] | Zheng D., Huang Z., Liu H., et al. (2025). VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness. arXiv preprint arXiv: 2503.21755. |
| [77] | Liu J., Liu G., Liang J., et al. (2025). Improving Video Generation with Human Feedback. Advances in Neural Information Processing Systems. |
| Jiang X., Liu Y., Shen Y., et al. (2026). Extracting creativity from hallucination: rethinking large language models. AI Plus 1:100008. https://doi.org/10.59717/ipj.aiplus.2026.100008 |
To request copyright permission to republish or share portions of our works, please visit Copyright Clearance Center's (CCC) Marketplace website at marketplace.copyright.com.
General framework of the survey.
Cases of relations between hallucination and creativity.
The divergent phase stimulates the hallucination of LLMs, fostering creative thought.
The convergent phase refines hallucinations into valuable creative contributions.