Việt Điển: Turning Digitized Hán–Nôm Collections into a Researchable Corpus

Digitization makes historical materials accessible. But access is only the beginning.
For scholars working with Hán–Nôm materials, a digitized manuscript or printed book may still require the same fundamental process as its physical counterpart: locating a text, opening its pages, reading through them, and identifying relevant characters, phrases, names, or passages one page at a time. What becomes possible when hundreds of digitized texts can instead be searched together as a corpus?
Việt Điển (越典), now integrated into the Hán–Nôm Research Hub on the Digitizing Vietnam platform, explores this possibility precisely.
Created by Albert Errickson, Việt Điển is a searchable digital library designed specifically for working with Hán–Nôm texts. It brings digitized page images together with machine-readable text, allowing researchers to move between the historical document and its OCR transcription rather than treating the two as separate resources. The growing corpus currently encompasses hundreds of texts and tens of thousands of digitized pages, drawing substantially on materials made available through Digitizing Vietnam, Columbia University Libraries, the Vietnamese Nôm Preservation Foundation, and other digital repositories.
From digital library to research corpus
The significance of Việt Điển lies not simply in putting more historical books online, but in changing how those books can be explored.
A researcher interested in a concept such as 國家 (quốc gia, state/nation), for example, no longer needs to know beforehand which individual work is likely to contain it. Việt Điển can search OCR across the corpus and identify pages on which the characters occur. Its search system also supports more complex queries: researchers can look for multiple terms on the same page, alternatives, exclusions, character patterns, or words occurring within a specified distance of one another. Searches automatically account for simplified and traditional character forms.
This makes it possible to begin research not only with a particular book, author, or title, but with a word, phrase, name, concept, or textual pattern.
The platform's Concordance function extends this approach further. Rather than simply identifying pages containing a term, it places occurrences alongside their surrounding textual context. Reading these occurrences together can reveal recurring expressions, collocations, shifts in usage, and patterns that would be difficult to notice while reading texts individually.
For Hán–Nôm studies, this opens possibilities for research across texts: tracing terminology through historical works, comparing how a concept appears in different genres, examining recurring formulations, or identifying materials that merit closer philological investigation.
Keeping the original page at the center
At the same time, Việt Điển does not treat OCR as a replacement for the historical document.
Search results lead researchers back to individual pages, where the source image can be read alongside its machine transcription. This relationship is particularly important for Hán–Nôm materials, where historical character forms, woodblock printing, manuscript variation, page condition, and the complexities of chữ Nôm can all present challenges for automated recognition.
Việt Điển therefore explicitly distinguishes between machine-generated OCR and texts that have been manually checked. Its Corrected Texts collection identifies transcriptions that have been read against the original scan character by character. For the larger machine-generated corpus, the platform cautions researchers that OCR may contain errors and encourages users to verify results against the original images.
This principle is fundamental to the approach of the Hán–Nôm Research Hub: computational tools can help researchers discover patterns and navigate collections at a scale that would otherwise be difficult, but scholarly interpretation still depends on returning to sources, examining textual context, and evaluating evidence critically.
Connecting Hán–Nôm sources with AI research
Việt Điển also provides an experiment in how historical collections might interact with a new generation of AI-assisted research.
Through its public API and Model Context Protocol (MCP) interface, researchers can connect compatible AI assistants to the Việt Điển corpus. The AI can then search the collection, locate occurrences of terms, compare their distribution across texts, retrieve OCR from particular pages, and direct the researcher back to the corresponding source images.
A researcher might ask:
Which texts discuss 科舉, the civil service examination system?
Or:
Where does 皇越 appear across the corpus, and in what contexts?
Or instead of knowing precisely what characters to search for, a researcher might begin with a question in ordinary language and use the corpus to identify relevant texts and passages for further investigation.
The important point is not to ask AI to replace reading or interpretation. Rather, the corpus gives AI a structured pathway back to historical evidence. Việt Điển instructs users to verify AI-assisted findings against page images and to cite the source rather than the assistant itself. It also cautions that a failed OCR search cannot establish that something is absent from the historical record: a character may have been recognized incorrectly, or a relevant work may not yet have been digitized.
In this model, AI becomes an interface for discovery and navigation, while the historical document remains the evidentiary foundation.
Building an interconnected Hán–Nôm research environment
The integration of Việt Điển into the Hán–Nôm Research Hub is part of Digitizing Vietnam's broader effort to build an interconnected digital environment for the study of Vietnam's premodern textual heritage.
The Hub brings together digital archives and manuscript catalogs with tools for OCR, multi-dictionary lookup, corpus searching, date conversion, reference resources, and platforms for learning chữ Nôm. Instead of requiring researchers and students to approach digitized collections, dictionaries, transcription tools, and computational methods as isolated resources, the Hub aims to connect them within a common research workflow.
Việt Điển adds an important layer to this ecosystem: corpus-level exploration.
A digitized page can be read. An OCR transcription can be searched. A corpus allows hundreds of texts to be investigated in relation to one another.
The distinction points toward a larger ambition for Digitizing Vietnam. Preserving cultural heritage digitally should not mean creating a static endpoint where historical documents are simply stored online. Digital collections can become foundations for new questions, new methods, and new forms of collaboration—while maintaining clear pathways back to the original materials on which scholarship depends.
Việt Điển demonstrates one way that transformation can take place: from page to text, from text to corpus, and from corpus back to the historical source.
Explore Việt Điển within the Hán–Nôm Research Hub on the Digitizing Vietnam Platform: