Speakers

More speakers will still be added.


Tomáš Foltýn, National Library of Czech Republic
Czech Republic

Tomáš Foltýn is the graduate of the Faculty of Arts of the University of Pardubice, where he completed his master studies in the field of cultural history in 2008. Since 2007 he worked in the Digitization Department of the National Library of Czech Republic, first as Project Manager, then as the Head of the Metadata Creation and Management Department. In January 2013, he was appointed as Collections´ Management Division Director. On May 1st 2021, he became the General Director of the National Library of the Czech Republic.

Tomáš Foltýn was involved in various national and international research projects, is an expert guarantor of the Ministry of Culture of the Czech Republic VISK 7 Funding Mechanism, the regular member of the Central Library Council of the Czech Republic, CENL Executive Committee Member etc. From the professional perspective, he is interested especially in the area of the effective collections´ management and long-term preservation of the modern collections including the digital content. During last years, he started to be involved also in various activities connected with digital humanities, R&D projects and cultural diplomacy.



Sanna Haukkala, The National Library of Finland
Finland

Sanna Haukkala is the Information Specialist at The Legal Deposit Services of The National Library of Finland. Sanna has an extensive work history with legal deposits, starting from printed legal deposits, music recordings and the ephemera collection since 2013 and she has worked mostly with the content of The Finnish Web Archive since 2017. She has a MA from Cultural Studies (Sociology of Art) with additional studies from Information Science and Interactive Media, and Museology.

Presentation: When the National Library Preserves the internet, what is being preserved, and why, for the future?
Since 2006 The National Library of Finland has had a legal right to preserve Finnish content from internet and from 2008 the legal obligation to harvest Finnish internet and content intended for Finnish audiences. The Legal Deposit Services has preserved printed Finnish publications in cooperation with the publishing industry for many centuries even before this. The National Library of Finland is also one of the founding members of International Internet Preservation Consortium (IIPC), that was founded in 2003 to gather and share knowledge between institutions that are doing web archiving.

The National Library of Finland have guidelines how online content have to be selected and preserved, to be ensured to be available to the future generations. The selected materials must first be able to be harvested, then preserved in a way they can be provided to users at the legal deposit workstations, and finally also ensure the digital preservation of the materials. These guidelines steer the work to the direction that The National Library of Finland can to provide a representative and diverse sample of Finnish online materials, but do we know who are the future users and how they’d want to utilize our materials?

Besides web archiving, preserving online materials includes also e-legal deposits, that are collected as e-books, e-journals, music recordings, sheet music and such. But while our mandate has been quite the same for over twenty years, the internet has changed drastically. What is the reason why we are still doing the same work while trying to keep up with changing online content, technical development, and massive changes in the publishing industry? This presentation will share some of the challenges considering preserving online materials and questions that are not anymore whispers between the library bookshelves.




Neil Jefferies, Open Preservation Foundation/Oxford university and Megan Gooch, Bodleian Libraries/Oxford university
United Kingdom

Neil Jefferies is Executive Director of the Open Preservation Foundation, a Director of Data Futures GmbH and a Digital Innovation Specialist at the Bodleian Libraries, University of Oxford. He is a co-creator of the International Image Interoperability Framework and the Oxford Common File Layout, Community Manager for the SWORD protocol and a member of the Bit List Council for the Digital Preservation Coalition. His research interests include knowledge and information models, APIs for enhancing access to digital resources, digital preservation and long-terms access, and the mechanisms of digital scholarship. He teaches on various topics on Oxford’s MSc in Digital Scholarship and the Digital Humanities Summer School.

Megan Gooch is the Head of the Centre for Digital Scholarship at the Bodleian Libraries, University of Oxford. She has worked in museum and heritage roles as a curator and producer. Her library work and research focuses on research data management, digitisation, digital skills and AI.

Presentation: The possibilities and pitfalls of working with big tech: the Oxford experience
In this paper we’ll talk about why and how the Bodleian Libraries, University of Oxford, has worked with several big tech companies in recent years. In short, it has enabled us to digitise and make available our collections at a scale unachievable through other workflows and funding models. But what do libraries gain, or lose, from working in partnership with vast multinational companies? We reflect on both the technical affordances of our work – what partnerships have allowed in terms of research and digitisation, but also the cultural challenges of innovation, working with AI, and staff, reader and public critiques of such collaborations.



Norman Meuschke, GippLab / University of Göttingen
Germany

Dr. Norman Meuschke’s main research interests are methods for semantic similarity analysis and their application for information retrieval. Beyond his core research, he is interested in applied data science and knowledge management challenges and the application of blockchain technology to tackle these challenges.

His research spans the fields of:

  • Information Retrieval (esp. for text, images, and mathematical content)
  • Natural Language Processing
  • Plagiarism Detection
  • Citation and Link Analysis
  • Blockchain Technology
  • Information Visualization


David Minor, UC San Diego Library
USA

David Minor works at the University of California, San Diego, where he is the Director of the Research Data Curation Program in the UC San Diego Library. In this role he helps define and lead work needed for the contemporary and long-term management digital resources. His position includes significant interaction with stakeholders on the UC San Diego campus, throughout the UC System, and national initiatives. His program also includes management of Chronopolis, a national-scale digital preservation network.



Jan Rychtář, Trinera s.r.o.
Czech Republic

Jan Rychtář has more than fifteen years of experience developing software for libraries and has been involved for many years in the development of Kramerius, the Czech digital library platform.

He is the co-founder and director of Trinera, where he leads a team building systems primarily for libraries, archives and other memory institutions, with a strong focus on automation, access to digital collections and the practical use of artificial intelligence.

Presentation: Liberio: A Digital Library Built for the AI Era
Liberio is a new digital library built from the ground up on modern technologies for processing, discovery and presentation of digital collections.

Rather than adding AI capabilities to an existing architecture, Liberio takes today’s technological capabilities as a starting point for the design of the entire digital library. Semantic and multimodal search, computer vision, speech recognition, machine translation, generative AI and automated content enrichment are not separate add-ons, but are embedded throughout the way content is processed, structured and presented to users.

Documents are not simply ingested and indexed. Liberio extracts text and structure, identifies articles, images, people, places and topics, processes speech and music, and connects the resulting information so that it can be used directly for search, exploration and interaction.

Books, newspapers, maps, images, sheet music, audio, video and 3D objects are brought together within one environment, while each medium has tools and interfaces adapted to its specific characteristics.

One of Liberio’s core principles is to give curators direct control over the whole digital library without dependence on technical staff.

The presentation will first demonstrate the full range of Liberio’s capabilities across different document and media types. It will then show examples from production instances already in active use, including how this approach has simplified operation and content management, opened up new ways of working with collections, and enhanced the experience offered to users.

Finally, the presentation will reflect on how the current pace of technological change may affect libraries in the years ahead, and what it may require from them to operate, develop and continuously adapt their digital libraries.



Alicia Wise, CLOCKSS
USA

Alicia Wise is Executive Director of CLOCKSS, where research libraries and academic publishers come together to ensure the long-term preservation of the scholarly record. She has been active in increasing access to research information in roles within the academic, library and publishing communities.

Presentation: Who will protect the eBooks?
How many books are published each year, and where are they being preserved for future generations? In 2026, these deceptively simple questions remain surprisingly difficult to answer and particularly for eBooks.

The preservation challenge for books is not simply a smaller version of the challenge for scholarly journals. eBooks involve different communities of stakeholders, distribution models, formats, identifiers, metadata, platforms, and rights arrangements. Much of the infrastructure and cooperation that has developed to preserve eJournals therefore does not translate directly to books.

This presentation will examine what makes eBooks particularly challenging to preserve, and what is already being done to reduce those risks. We will look at emerging projects, services and models for long-term preservation and access, and consider the roles that libraries, infrastructure providers, publishers, standards organisations, and other stakeholders can play.
The scale and international nature of the challenge make cooperation essential. Europe is particularly well placed to help shape a more coordinated approach, connecting national and institutional initiatives with global infrastructure and expertise.

During this session we will explore what a more global and resilient preservation environment for books could look like, and what practical steps we can take now to build it.





Petr Žabička, Moravian Library, Filip Bím, Moravian Library and Jan Rychtář, Trinera s.r.o.
Czech Republic

Petr Žabička is an expert in library automation with experience in digitisation, digital libraries, and machine learning. As an associate director at the Moravian Library, he is responsible for research and development projects. Currently, his activities focus on implementing machine learning technologies to enhance access to digitised documents. He has been involved in the PERO project, which aimed to improve the accuracy of digitised texts through the application of machine learning algorithms to optical character recognition (OCR). Previously, he led projects related to map digitisation, online access to digitised maps, and the development of the Czech library portal Knihovny.cz.

Filip Bím specializes in machine learning, OCR and computer vision, with a focus on the analytical applications of artificial neural networks. As Head of the Digital Document Management Department at the Moravian Library, he leads projects that integrate machine learning into library workflows. His team develops tools that apply modern AI methods to digitization, data processing, and the accessibility of digital collections. His work aims to make these methods practical for libraries and cultural heritage institutions.

Jan Rychtář has more than fifteen years of experience developing software for libraries and has been involved for many years in the development of Kramerius, the Czech digital library platform.

He is the co-founder and director of Trinera, where he leads a team building systems primarily for libraries, archives and other memory institutions, with a strong focus on automation, access to digital collections and the practical use of artificial intelligence.

Presentation: From Scan to Archival Package: Automating the Digitisation Workflow
Digitisation involves many processing steps beyond scanning itself. Many of these are traditionally performed manually, especially image cropping, identifying document boundaries, and creating structural metadata.

Using a large collection of broadside ballads as a case study, we present an automated workflow covering the entire process from raw scans to a validated archival package.

Broadside ballads represent a particularly challenging type of material: they are usually very short, often only four or eight unnumbered pages, and may also be bound together in composite volumes. For efficient digitisation, large batches are therefore scanned sequentially without identifying individual documents during scanning.

The subsequent processing is largely automated. Scans containing two facing pages are detected and split into individual page images, which are then cropped and deskewed. Large sequences consisting of thousands or tens of thousands of pages are subsequently analysed and divided into individual documents. The workflow reconstructs their internal structure, detects potentially missing, duplicated, or incorrectly assigned pages, and can also preserve or identify relationships between documents that belong together.

The workflow combines AI-based methods with conventional algorithmic processing to automate image processing, OCR, bibliographic metadata processing, structural metadata generation, and quality control. Existing bibliographic records are retrieved from the library catalogue, transformed into the required target formats, and enriched with information derived during digitisation. The system also detects inconsistencies between the digitised objects and catalogue data, supports the registration and assignment of national identifiers, and automatically assembles the resulting data into archival packages compliant with Czech digitisation standards, including their final validation.

The case study demonstrates how individual automation components can be connected into a single processing pipeline in which human intervention is concentrated primarily on exceptional or ambiguous cases. The resulting workflow enables tens of thousands of documents to be processed efficiently and provides a reusable model that can be adapted to other types of digitised collections.



Anastasia Zhukova, GippLab / University of Göttingen
Germany

Anastasia Zhukova has completed her Engineering degree (Diploma in Information Technology) at Moscow Aviation Institute (National Research University) and proceeded with her Master’s studies at the University of Konstanz with a focus on Machine Learning and Natural Language Processing. She compiled her master’s thesis on the topic of Automated Identification of Framing by Word Choice and Labeling to Reveal Media Bias in News Articles.

Her research interests focus on two projects: identifying media bias and developing an AI assistant for plant operations. The first project is interdisciplinary research that applies cross-document coreference resolution to resolve mentions with high lexical diversity to automatically identify media bias by word choice and labeling. The second project is an industry project focusing on the domain adaptation of language models, information retrieval and information extraction in the low resource setup.

Anastasia’s prime interest lies in the areas of:

  • Applied Natural Language Processing
  • Cross-document coreference resolution
  • Domain adaptation of language models
  • NLP for the low-resource languages, e.g., German
  • Information visualization