Newspaper articles have long been a valuable source for understanding historical events (Meroño-Peñuela et al., 2015). For researchers of Dutch history in particular, Delpher newspaper dataset can offer important clues about how people, places, and events were described in the past. This is especially true for colonial history, where newspaper coverage often helps explain the roles of military personnel, civil servants, explorers, and other actors involved in expeditions and administrative activities.
These articles are not only useful for reconstructing events themselves. They also can help us understand how cultural heritage objects ended up in museum collections today. In many cases, the context in which an object was collected, transferred, or displaced can be traced by looking at historical newspaper references to the people and events connected to it (Drieënhuizen, 2022). For this reason, exploring historical events through newspaper articles—and identifying who appears in connection with those events—is both important and insightful.
In my Research-in-Residence project, together with colleagues at the KB (Rick Schouten, Celonie Rozema, Mirjam Cuper and Steven Claeyssens), we are exploring how to automatically find newspaper articles that mention specific historical events and how to build a network of people associated with those events. This will help us map both the well-known figures and the less visible actors who appear only occasionally in the sources. In other words, we want to uncover the “long tail” of historical actors: the people who are not yet widely recognized, but who nevertheless played a role in these events.
Traditionally, domain experts have approached such information retrieval tasks in a snowballing method. They begin with a historical event—often with a known or plausible event title—and retrieve a few relevant articles. From those articles, they identify names of people, places, or related events. These new clues are then used to search for more articles, which leads to more names, more context, and gradually a larger picture of the event network. While effective, this manual process is time-consuming and difficult to scale.
Our project takes a computational approach to this problem. Instead of relying only on manual searching, we aim to systematically collect and process articles linked to historical events and extract the people mentioned in them. The work is divided into three main steps:
- Building a corpus of newspaper articles relevant to selected historical events.
- Running named entity recognition (NER) to identify persons, places, and historical events in those articles.
- Disambiguating the extracted entities so that people, places, and events can be reliably grouped together.
Together, these steps will allow us to create a more structured view of how historical events are represented in newspapers and how people are connected across different articles.
Challenges and adapted solutions
1. Harvesting the corpus
One of the first challenges is deciding how to identify which articles are truly relevant to a historical event. Events are often referred to in different ways, with spelling variation, alternative naming, and incomplete contextual descriptions. This makes retrieval of relevant articles difficult, especially when the event is known only from a vague title or a partial reference.
Previous work often addresses this problem through iterative search (Blin et al., 2025) or broader context matching to retrieve relevant articles (Voskarides et al., 2021). In our project, we will use a more systematic retrieval workflow that begins with a list of already known historical events, together with each event’s description and date. We then compare the event title lexically with article text and restrict retrieval to a publication window of plus or minus ten years around the event date. This helps us keep the search focused on articles that are both temporally plausible and textually related to the event.
2. Running named entity recognition
Another challenge is extracting useful entities from historical newspaper text. Historical newspapers are noisy: OCR errors, old spelling, inconsistent punctuation, and long sentence structures can all make entity recognition difficult. Standard NER models often struggle when applied directly to this kind of material.
Many studies address this by fine-tuning models on historical text (Manjavacas Arevalo & Fonteyn, 2022) or by using domain-specific annotation workflows (Arnoult et al., 2021). In our project, we will use GLiNER (Zaratiana et al., 2024), because it offers a flexible and multilingual way to identify entities across different categories while being suitable for adaptation to noisy historical text. We expect this to support robust recognition of persons, places, and event mentions in newspaper articles.
3. Disambiguation
After entity extraction, a third challenge appears: not every extracted named entity refers to a unique person, place, or event. Historical entities are often ambiguous. A person may appear under different spellings or name forms, such as “J. van Heutsz,” “Joan van Heutsz,” or “Van Heutsz”; a place may have different names over different time periods, such as “Batavia” and “Jakarta,” or colonial and postcolonial variants of the same location; and an event may be mentioned in slightly different ways across articles, such as “the Central New Guinea Expedition,” “the 1921–1922 expedition to New Guinea,” or simply “the New Guinea expedition.”
Researchers usually handle this through a combination of heuristics, authority files, knowledge bases, or string matching techniques. In our project, we will use a tailored approach for each entity type:
- Historical events: we will combine heuristic filtering and string similarity, using a time window of approximately twenty years around the event date and a similarity threshold to determine relevance.
- Places: we will use GeoNames mapping, following an approach similar to DezzyMatch (Hosseini et al., 2020), which has shown good results in earlier work.
- Persons: we will use BurgerLinker, since it has already shown promising performance for person disambiguation in historical contexts.
This combination of methods should help us connect extracted entities to the correct historical referents and make the resulting network more reliable.
The ultimate goal of this project is not only to retrieve articles, but to uncover patterns that are otherwise hard to see. By linking events, people, and places across large newspaper collections, we can better understand how historical knowledge is distributed in the archive. We can also identify less visible historical actors, helping to fill in gaps in established narratives.
This is especially valuable for colonial history, where the archive is rich but uneven, and where important actors may appear only briefly or indirectly. By combining computational methods with historical expertise, we hope to create a workflow that is both scalable and trustworthy. In that sense, the project is not just about automation—it is about enabling deeper historical interpretation at a much larger scale.
References
Meroño-Peñuela, A., Ashkpour, A., van Erp, M., Mandemakers, K., Breure, L., Scharnhorst, A., Schlobach, S., & van Harmelen, F. (2015). Semantic technologies for historical research: A survey. Semantic Web, 6(6), 539–564. https://doi.org/10.3233/SW-140158
Drieënhuizen, C. A. (2022). Objects of belonging and displacement; Artefacts and European migrants from colonial Indonesia in colonial and post-colonial times. Wacana, 23(3), 635-653. Article 6. https://doi.org/10.17510/wacana.v23i3.1005
Zaratiana, U., Tomeh, N., Holat, P., & Charnois, T. (2024). GLiNER: Generalist Model for Named Entity Recognition using Bidirectional Transformer. In K. Duh, H. Gomez, & S. Bethard (Eds), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers) (pp. 5364–5376). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.naacl-long.300
Manjavacas Arevalo, E., & Fonteyn, L. (2022). Non-Parametric Word Sense Disambiguation for Historical Languages. In M. Hämäläinen, K. Alnajjar, N. Partanen, & J. Rueter (Eds), Proceedings of the 2nd International Workshop on Natural Language Processing for Digital Humanities (pp. 123–134). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.nlp4dh-1.16
Arnoult, S. I., Petram, L., & Vossen, P. (2021). Batavia asked for advice. Pretrained language models for Named Entity Recognition in historical texts. In S. Degaetano-Ortlieb, A. Kazantseva, N. Reiter, & S. Szpakowicz (Eds), Proceedings of the 5th Joint SIGHUM Workshop on Computational Linguistics for Cultural Heritage, Social Sciences, Humanities and Literature (pp. 21–30). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.latechclfl-1.3
Hosseini, K., Nanni, F., & Coll Ardanuy, M. (2020). DeezyMatch: A Flexible Deep Learning Approach to Fuzzy String Matching. In Q. Liu & D. Schlangen (Eds), Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations (pp. 62–69). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.emnlp-demos.9
Blin, I., Tiddi, I., Trijp, R. van, & Teije, A. ten. (2025). ChronoGrapher: Event-Centric Knowledge Graph Construction via Informed Graph Traversal. Semantic Web, 16(5), 22104968251377247. https://doi.org/10.1177/22104968251377247
Voskarides, N., Meij, E., Sauer, S., & de Rijke, M. (2021). News Article Retrieval in Context for Event-centric Narrative Creation. Proceedings of the 2021 ACM SIGIR International Conference on Theory of Information Retrieval, ICTIR ’21, 103–112. https://doi.org/10.1145/3471158.3472247
Image by: Philipp Schmitt & AT&T Laboratories Cambridge / https://betterimagesofai.org / https://creativecommons.org/licenses/by/4.0/