Application of Comics Generation Techniques in the Visualization of Vietnamese Folk Tales

Online First: 25/08/2026

Các tác giả

Email tác giả liên hệ:

tanlm@hcmute.edu.vn

DOI:

https://doi.org/10.54644/jte.2026.2148

Từ khóa:

Story-Iter, Diffusion, Generative AI, Comics generation, Folk comics

Tóm tắt

In recent years, the integration of Artificial Intelligence into tasks requiring creative cognition has become increasingly prevalent. Although the nature of machine creativity remains a subject of debate, it is undeniable that these models have exerted a positive influence, ranging from the reduction of lead times, significant cost optimizations for enterprises to versatile support in day-to-day activities. This study focuses on the generation of Vietnamese folk comics utilizing the Story-Iter architecture, which possesses an inherent advantage in maintaining contextual consistency via historical image references. Furthermore, we propose an enhancement involving the selective curation of reference image sets rather than incorporating the entire dataset or a fixed sequence of recent frames, thereby mitigating the integration of noise during the generation of novel imagery. This strategy effectively minimizes the inclusion of noisy data during the synthesis of new images. Experimental results performed on our self-built small-sized story dataset demonstrate that the approach yields efficacy when synthesized with the original Story-Iter framework, while simultaneously highlighting persistent challenges.

Tải xuống: 0

Dữ liệu tải xuống chưa có sẵn.

Tiểu sử của Tác giả

Minh Tan Le, Ho Chi Minh City University of Technology and Engineering, Vietnam

Minh Tan Le received his Bachelor’s degree in Information Technology from Ho Chi Minh City University of Technology and Engineering in 2019, and completed his Master’s degree in Computer Science in 2021. He is currently a Lecturer at the Faculty of Information Technology, Ho Chi Minh City University of Technology and Engineering, Vietnam. His academic and research interests center on Generative Artificial Intelligence and Computer Vision, with a focus on developing innovative solutions that push the boundaries of intelligent computing.

Email: tanlm@hcmute.edu.vn. ORCID:  https://orcid.org/0009-0004-4912-6795.

Dinh Tan Loc Nguyen, Ho Chi Minh City University of Technology and Engineering, Vietnam

Dinh Tan Loc Nguyen has been studying Data Engineer at Ho Chi Minh City University of Technology and Engineering since 2023. His academic and research interests center on Generative Artificial Intelligence.

Email: 23133041@student,hcmute.edu.vn. ORCID:  https://orcid.org/0009-0007-8294-9007

Tài liệu tham khảo

H. King, “Exclusive: Gen AI music app Suno comes out of stealth,” Axios, Dec. 20, 2023. [Online]. Available: https://www.axios.com/2023/12/20/suno-gen-ai-music-microsoft. [Accessed: Mar. 25, 2026].

OpenAI, “DALL·E 3 is now available in ChatGPT Plus and Enterprise,” OpenAI, Oct. 19, 2023. [Online]. Available: https://openai.com/index/dall-e-3-is-now-available-in-chatgpt-plus-and-enterprise/. [Accessed: Mar. 25, 2026].

I. Goodfellow et al., “Generative adversarial networks,” Communications of the ACM, vol. 63, no. 11, pp. 139–144, Nov. 2020, doi: 10.1145/3422622. DOI: https://doi.org/10.1145/3422622

Y. Li et al., “StoryGAN: A Sequential Conditional GAN for Story Visualization,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 6329–6338. DOI: https://doi.org/10.1109/CVPR.2019.00649

A. Maharana, D. Hannan, and M. Bansal, “Improving Generation and Evaluation of Visual Stories via Semantic Consistency,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Jun. 2021, pp. 2432–2443, doi: 10.18653/v1/2021.naacl-main.194. DOI: https://doi.org/10.18653/v1/2021.naacl-main.194

A. Maharana and M. Bansal, “Integrating visuospatial, linguistic, and commonsense structure into story visualization,” in Proc. 2021 Conf. Empirical Methods Nat. Lang. Process. (EMNLP), Nov. 2021, pp. 6772–6786, doi: 10.18653/v1/2021.emnlp-main.543. DOI: https://doi.org/10.18653/v1/2021.emnlp-main.543

Y. Zhou, D. Zhou, M. M. Cheng, J. Feng, and Q. Hou, “StoryDiffusion: Consistent Self-Attention for Long-Range Image and Video Generation,” in Advances in Neural Information Processing Systems, vol. 37, 2024, pp. 110315–110340, doi: 10.52202/079017-3501. DOI: https://doi.org/10.52202/079017-3501

H. He et al., “DreamStory: Open-domain story visualization by LLM-guided multi-subject consistent diffusion,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 47, no. 12, pp. 11874–11891, 2025, doi: 10.1109/TPAMI.2025.3600149. DOI: https://doi.org/10.1109/TPAMI.2025.3600149

J. Mao et al., “Story-Iter: A Training-free Iterative Paradigm for Long Story Visualization,” Jan. 2026. [Online]. Available: https://openreview.net/forum?id=puBVb9vTah. [Accessed: Mar. 25, 2026].

C. Liu, H. Wu, Y. Zhong, X. Zhang, Y. Wang, and W. Xie, “Intelligent Grimm-open-ended visual storytelling via latent diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 6190–6200. DOI: https://doi.org/10.1109/CVPR52733.2024.00592

X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “Sigmoid loss for language image pre-training,” in Proceedings of the IEEE/CVF international conference on computer vision, 2023, pp. 11975–11986. DOI: https://doi.org/10.1109/ICCV51070.2023.01100

A. Radford et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning, 2021, pp. 8748–8763.

Minh Tan Le et al., “VNFairyTales,” GitHub, 2026. [Online]. Available: https://github.com/RoggerTan/VNFairyTales. [Accessed: Mar. 25, 2026].

X. Zhai, B. Mustafa, A. Kolesnikov, and L. Beyer, “google/siglip-base-patch16-224,” Hugging Face, 2023. [Online]. Available: https://huggingface.co/google/siglip-base-patch16-224. [Accessed: Mar. 25, 2026].

J. Cheng et al., “Theatergen: Character management with LLM for consistent multi-turn image generation,” 2024, arXiv:2404.18919. [Online]. Available: https://arxiv.org/abs/2404.18919.

J. Deng, J. Guo, N. Xue, and S. Zafeiriou, “Arcface: Additive angular margin loss for deep face recognition,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2019, pp. 4690–4699. DOI: https://doi.org/10.1109/CVPR.2019.00482

Y. Zheng et al., “General facial representation learning in a visual-linguistic manner,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 18697–18709. DOI: https://doi.org/10.1109/CVPR52688.2022.01814

Tải xuống

Đã Xuất bản

2026-08-25

Cách trích dẫn

[1]
M. T. Le và D. T. L. Nguyen, “Application of Comics Generation Techniques in the Visualization of Vietnamese Folk Tales: Online First: 25/08/2026”, JTE, tháng 8 2026.

Số

Chuyên mục

Bài báo khoa học

Categories