Not-for-profit DH Training for 11 Days: Over 270 students explore the intersection of AI and Humanities at the 4th Corpus and Digital Humanities Summer College
The Fourth Corpus and Digital Humanities Summer College was held at Suiyuan Campus, Nanjing Normal University from July 25 to August 4, 2026. The 11-days event was hosted by the School of Chinese Language and Literature and the Academy for Studies in Humanities and Social Sciences at Nanjing Normal University (NNU). The Faculty of Arts and Humanities at the University of Macau (UM), the School of Humanities and Social Science at the Hong Kong University of Science and Technology (HKUST), the School of Humanities and Social Science and the School of General Education at Beijing Normal-Hong Kong Baptist University (BNBU), and the College of Information Management at Nanjing Agricultural University (NAU). Department of Minority Language Experiment at the Institute of Ethnology and Anthropology, Chinese Academy of Social Sciences is also a main sponsor. By offering free trainings, the summer college sought to cultivate a new generation of interdisciplinary talent in digital humanities from around the world.

Opening Ceremony: A Century of Intellectual Heritage Meets the Voice of a New Era
The opening ceremony marked the beginning of this summer college in the morning of July 26 at Yifang Lecture Hall, Suiyuan Campus. It was chaired by Jun Tian, Party Secretary of the School of Chinese Language and Literature at Nanjing Normal University. More than 40 scholars and experts from China and abroad joined in, with 130 in-person participants and 140 online ones. The summer school scheduled 20 keynote lectures, four themed symposiums (including 15 presentations), two workshops on machine learning and corpus establishment, and a trip to Changzhou. We split all participants to three separate courses: corpus construction, statistical analysis of language, and the application of large language models.

In his opening speech, Guodong Ben, the vice deputy Party Secretary of Nanjing Normal University, recalled two milestones from a century ago: the renowned educator Heqin Chen’s pioneering quantitative analysis of Chinese characters, later compiled as A Vocabulary of Characters in Modern Vernacular Chinese, and the history of summer college from 1920 by Xingzhi Tao, Heqin Chen and their colleagues. The reopening summer school of corpus linguistics a century later inherited and carried forward the scholarly spirit of these pioneers, but also responded to the new demands placed on linguistic research in the AI age.
Ping Li, Dean of the School of Humanities and Social Science at HKUST, noted that digital technologies are engaged in all stages of research, including literature retrieval, corpus process, and testing hypotheses. Humanities scholars, he argued, should be neither opponents of AI nor subordinate to technology, instead, they should provide value-based leadership in the AI era—embracing technology with openness while guiding it with humanistic insight. Haipeng Guo, Dean of the School of General Education at BNBU, noted that the integration of corpus research and digital humanities offers both a new window for linguistic and cultural exchange and a new methodology for examining the transmission of human civilization. Cameron Campbell, Chair Professor in the Division of Social Science at HKUST, encouraged early-career researchers to remain grounded in humanities and social science research while engage new technologies. Dongbo Wang, Deputy Dean of the College of Information Management at Nanjing Agricultural University, added that big data and AI are driving the humanities towards a new paradigm, one that is data-driven and digitally empowered, and that the convergence of corpus construction, quantitative analysis, and large language models is central in this transformative age.

At the opening ceremony, the co-organising institutions also presented a series of recent research achievements. Yulin Yuan from University of Macau introduced the “Chinese Room” project; Michael Yan Hon Chung from HKUST presented the modelling OCR techniques for ethnic minority languages; BNBU announced the setup for an interdisciplinary Digital Humanities research program; Bin Li from Nanjing Normal University unveiled a ranking list for ancient Chinese characters; and Dongbo Wang from Nanjing Agricultural University launched version 3.0 of the “Xunzi” multimodal large language model for premodern Chinese texts. The guests then jointly illuminated a holographic launch sphere to declare the opening of Fourth Summer College. The event was reported by Jiangsu Broadcasting Corporation, Jiangsu City Channel, Xinhua News, Nanjing Daily, and other media channels.

Focus on Digital Scholarships: 4 Symposiums and 40 Research Presentations
The summer college invited more than 40 scholars from the United States, the United Kingdom, Belgium, Chinese Mainland, Hong Kong (China) and Macao (China) to deliver lectures and join discussions both online and in person. Industry specialists from Baidu, Alibaba, ByteDance, and other companies also contribute their achievements in a symposium. The program comprised more than 40 academic presentations, four themed symposiums, two workshops, one private meeting, and three parallel courses covering database programming for corpus construction, statistical analysis of language and linguistics, and the application of large language models.
The first industry symposium was held at Yifang Lecture Hall on the afternoon of July 26. Specialists from Baidu Paddle, Alibaba, ByteDance, Gulian (Beijing) Media Tech Co., Ltd., a subsidiary of Zhonghua Book Company (Gulian), Jiangsu Sunyu Information Technology Co., Ltd. (Sunyu), and Nanjing Toushi Tech (Toushi) shared perspectives on collaboration among industry, universities, and research institutions. Fang Qian, who leads education ecosystem operations at Paddle, argued that advances in AI depend not only on algorithms and computing power, but also on the deep integration of domain expertise and industry knowledge. Haijun Li, an algorithm specialist with Alibaba’s Token Hub division, examined the practical deployment of large models from a data-centered perspective. Mingyue Zhang, project manager of ByteDance’s non-profit ancient-text initiative, introduced “Shidian Guji”, a digital platform where users can freely access approximately 70,000 premodern Chinese texts. Its app has been downloaded more than 1.6 million times, while the “Collating Ancient Texts with AI” initiative has attracted around 60,000 participants and facilitated the collation of approximately 2.1 billion Chinese characters. Ruixin Su, Director of the Ancient Texts Laboratory of Gulian, explained how the editing and publication of premodern works is moving beyond digital text storage towards intelligent, knowledge-centered systems. Fengfeng Sun, founder of Toushi, demonstrated applications of the “Qiyun Reborn” AI painting model in reinterpreting traditional culture. Xinhua News afterwards reported the event, and received more than 500,000 views online.

On 27th July, the Summer College convened its second symposium on standardizing Language and Writing. Scholars from institutions including the Institute of Applied Linguistics at the Ministry of Education and the China Electronics Standardization Institute discussed standards for language and writing in the age of AI, as well as the intelligent development of language resources. Hui Li of the Institute of Applied Linguistics noted that, in an era of human–machine symbiosis, language norms are expanding from designed for human use to machine use. Yang Tao of the China Electronics Standardization Institute and Huidan Liu of the Institute of Software at the Chinese Academy of Sciences highlighted the lack of phonetic and semantic attribute data for characters in the extension blocks of the CJK encoding standard, calling for a character-resource platform on a large scale that integrates pronunciation, graphic form, and semantic meaning. Professor Bin Li from Nanjing Normal University emphasized the need of internationalization of Chinese-language digitization, and to collaborate with scholars worldwide on ISO language and writing standards and the construction of AI corpora. Xinhua News reported on this symposium and received more than one million views.

On 1st August, the Future of Linguistics symposium examined the position and prospects of the discipline in the era of large language model. Professor Wei Yuan from National University of Defense Technology reminded us to be careful of LLM in value judgement and the importance of humanist views. Dr Yunfei Long, Senior Lecturer at Queen Mary University of London, shared his publication on the lack reliability of machine learning in sensory knowledge and that research must move beyond the constraints of a purely linguistic modality. During the panel, Professor Rou Song from Beijing Language and Culture University pointed out that large language models “can speak without being able to explain the reasonale”: they are good at capturing co-occurrence patterns but struggle to account for deeper rationales. Professor Xuri Tang of Huazhong University of Science and Technology said that large language models are only tools, he also advocated closer collaboration between linguistics and science and engineering. These speakers shared the idea that LLM should not be seen as a threat but as an opportunity, however, linguists should embrace scientific methods while maintaining its commitment to foundational humanities research.
At the future of Digital Humanities symposium on 3rd August, Simon Mahony, Professor Emeritus at University College London, and Edward Vanhoutte, Editor-in-Chief of Digital Scholarship in the Humanities, spoke respectively on reimagining digital humanities in the age of generative AI and the paradox of scope in scholarly publishing. Dr Zhu Cuiping, chief-in-editor of Gulian Publishing Company, introduced us the history and development of DH in China.
Training Courses and Cultural Trip: Applying Knowledge in Daily Practise
Three parallel training streams covered corpus construction, statistical analysis of language and linguistics, and the application of large language models. Class A used the free, open-source combination of MySQL and PHP, taking the Complete Tang Poems corpus as an example to introduce corpus construction, character encoding, and quantitative analysis. Class B focused on exploring statistical methods using SPSS. Class C explored core large-model technologies through the “Xunzi” large language model for premodern Chinese texts.
On the afternoon of 28th July, students visited The Oriental Metropolitan Museum in Nanjing. Nearly 100 participants toured the museum in groups led by Dr. Gaoqi Rao, Dr. Haonan Ma, Professor. Bin Li, Dr. Minxuan Feng, and other instructors. As they encountered the artefacts and architecture of the Six Dynasties, they drew on ideas from the summer school to reflect on such topics as the digital capture of cultural objects, the structured organization of premodern texts, and immersive, interactive exhibition design.
Closing Ceremony: Taking Technology as our Wings, the Humanities as our Anchor
The summer college concluded at Yifang Lecture Hall on the afternoon of 4th August. Qi Su, Associate Professor at the School of Foreign Languages, Peking University, delivered a keynote entitled “The Synergistic Development of AI and the Humanities: From Tool Use to Theoretical Guidance”. She concluded that the integration of AI and the humanities must go beyond the instrumental use of technology: humanities scholars should evolve from users of tools into builders of models.

Professor Xianyi Sha, Dean of the School of Chinese Language and Literature at Nanjing Normal University, reflected on the expansion of the Summer College in both scale and reach—from an initiative first hosted by Nanjing Normal University, to a partnership between two universities, and now to a program co-organized by five. Its curriculum and formats had also become more substantial and diverse, with the introduction of the industry–academia–research forum and the Forum on Standards for Language and Writing. From Heqin Chen’s corpus-based character-frequency research a century ago to the program’s recognition as a national innovation case today, this intellectual tradition has continued without interruption.
Professor Emeritus Huang Chu-Ren of the Hong Kong Polytechnic University argued that humanities research cannot be detached from its time: big data and large language models are not external threats, but part of the world in which we live, and scholars should actively incorporate technology into a framework of humanistic concern. Research Professor Congjun Long, the Institute of Ethnology and Anthropology at the Chinese Academy of Social Sciences, described the humanities as undergoing a paradigm shift in which corpus resources and AI reinforce one another, creating an urgent need for interdisciplinary researchers with both humanistic literacy and technical expertise. Professor Xiaobing Fang of the School of Liberal Arts at Nanjing University added that the rise of digital humanities is broadening and renewing approaches to language planning, with language itself becoming an important component of new quality productive forces.
The closing ceremony rewarded certificates to instructors, teaching assistants, volunteers, and all participants.

It also rewarded three outstanding participants, one from each course. In the database-programming class, Gang Zhang of Shanghai Normal University developed a system for tracing phonological change in Guangyun, transforming a rhyme dictionary containing tens of thousands of characters into a searchable, visual database. In the statistical-methods class, Huanhui Ye of Tianjin University constructed a statistical model for hotel selection by analysing textual reviews. In the large-language-model programming class, Yanzhe Liu of Nanjing University received wide acclaim for a system that recognizes and synthesizes Nanjing dialect speech.

Three Years of Sustained Effort and Achievement
The Summer College has maintained four defining principles: grounding in Chinese culture, empowerment through technology, case-based teaching, and open sharing. Since its launch in January 2024, it has been held successfully in the last three years. In June 2026, it was selected as a case in the “National Innovation Program for Language-Technology Empowerment in Key Fields”, organized by the National Language Commission and the Department of Language Information Management at the Ministry of Education, China. This year’s summer college attracted approximately 2,000 applications from nearly 400 universities across almost 20 countries and regions.
From Heqin Chen’s hand-drawn character-frequency tables in 1920 to the program’s selection as a national innovation case in 2026, a century-old scholarly lineage continues to flourish. As Professor Bin Li put it, the core value of this year’s summer college was to “bring good questions with good methods”. The closing ceremony marked not an end, but a new beginning. It is hoped that more early-career researchers will take this program as a point of departure, continue to cultivate the expansive field of digital humanities, and enable the humanities and technology to flourish through deeper integration.
