Modeling Text Editions as Linked Data (edition2LD)
Duration: January 1, 2023–December 31,12, 2023
Editions as Linked Data
Scholarly text editions pose a major challenge in terms of secondary, cross-disciplinary analysis and their reusability due to a variety of factors. These challenges encompass both the compilation of content and the analysis and preservation of the collected data. Among the challenges are:
- differences in the temporal and geographical focus of the research content
- various languages and language levels
- various writing systems
- Presentation in various combinations of editorial text, translations, facsimiles, and more.
- Digital data is available in various system architectures and data models
- Variability of the collected data during the edition’s lifetime “hot data” (Only the long-term archiving of the data after project completion “cold data” guarantees the immutability of the data.)
Making research data accessible over the long term is therefore an essential element of scientific research.
The project
The “edition2LD” project is working on a solution for data curation that makes heterogeneous research data accessible in the long term and across the boundaries listed above. This solution must be flexible enough to accommodate the heterogeneity of the data and, at the same time, stable enough for a long-term perspective. To achieve this, the edition2LD project follows the paradigm of Linked Data (LD) and the integration of data into the Semantic Web.
Use Case as “Best Practice”
To develop the workflow for Linked Data (LD) modeling, the project has chosen as its use case the editions of the HAdW research project “Documents on the History of Religion and Law of Pre-modern Nepal” (Nepal-FS) at the HAdW as its use case. The research project produces digital editions of Nepali texts (in Devanagari script) including facsimiles, English translations (in the Latin script), commentaries, and index entries (people, places, technical terms).
Publications are available at nepalica.hadw-hw.de and are assigned a DOI on the Heidelberg University Library website (see, e.g., https://digi.hadw-bw.de/view/dna_0001_0005).
Approach
The workflow for LD modeling should be generic enough to allow for future transferability to other projects. At the same time, it should be capable of—when triggered repeatedly—converting data in batches to RDF, thereby addressing the major challenge posed by the ever-changing “hot data”. When developing the automated mapping processes, it is therefore immensely important to minimize the need for manual post-processing—ideally to the point where it needs to be performed only once.
The project focuses on modeling the information units “text,” “English translation,” named entities (names of people and places) and technical terms, as well as metadata. For modeling named entities, the project can draw on Nepal-FS registry entries, some of which already contain further references to authority data repositories and encyclopedic resources.
The vocabularies, ontologies, and repositories used for modeling are those already established as standards: RDFS, SKOS, Gemeinsame Normdatei GND, VIAF, DBpedia, GeoNames, FOAF, etc. In addition, LD modeling includes links to instances of two ontologies for historical personal and place names in Nepal (NepalPeople and NepalPlaces, see Tittel 2022*), which are developed based on research data from the Nepal-FS.
The data sources are:
- the Nepal-FS database
- files containing additional information
- Information integrated into the pipeline via web crawling from registry and glossary entries
- NepalPeople and NepalPlaces ontologies: Python scripts synchronize the modeled data with entries from NepalPeople and NepalPlaces and, where applicable, integrate links to their instances.
As of September 2023, the modeling of named entities and terms is 90% complete; the modeling of the “English translation,” “Nepali Edition,” and the metadata is in progress.
The project “Language-Data-Based Modeling of Knowledge Networks in Medieval Romance Speaking Europe (ALMA)” (internal link), which was launched on August 1, 2022, as an inter-academic project of the HAdW, BAdW, and AdW Mainz within the Academies Program. Since ALMA produces text editions (in this case, of medieval legal and medical texts), this dataset is also well-suited for an edition2LD approach.
Publications
Svoboda-Baas, Dieta/Tittel, Sabine: Text+Plus, #04: Modeling Text Editions as Linked Data (edition2LD), in: Text+ Blog, Dec. 18, 2023, https://textplus.hypotheses.org/8723.
*Tittel, Sabine. "Towards an Ontology for Toponyms in Nepalese Historical Documents," in: Proceedings of the Workshop on Resources and Technologies for Indigenous, Endangered, and Lesser-Resourced Languages in Eurasia within the 13th Language Resources and Evaluation Conference, Marseille, June 2022, 2022, Marseille (European Language Resource Association - ELRA), pp. 7–16.

