More speakers will still be added.
![]() ![]() ![]() |
Petr Žabička, Moravian Library, Filip Bím, Moravian Library and Jan Rychtář, Trinera s.r.o. Czech Republic Petr Žabička is an expert in library automation with experience in digitisation, digital libraries, and machine learning. As an associate director at the Moravian Library, he is responsible for research and development projects. Currently, his activities focus on implementing machine learning technologies to enhance access to digitised documents. He has been involved in the PERO project, which aimed to improve the accuracy of digitised texts through the application of machine learning algorithms to optical character recognition (OCR). Previously, he led projects related to map digitisation, online access to digitised maps, and the development of the Czech library portal Knihovny.cz. Filip Bím specializes in machine learning, OCR and computer vision, with a focus on the analytical applications of artificial neural networks. As Head of the Digital Document Management Department at the Moravian Library, he leads projects that integrate machine learning into library workflows. His team develops tools that apply modern AI methods to digitization, data processing, and the accessibility of digital collections. His work aims to make these methods practical for libraries and cultural heritage institutions. Jan Rychtář has more than fifteen years of experience developing software for libraries and has been involved for many years in the development of Kramerius, the Czech digital library platform. He is the co-founder and director of Trinera, where he leads a team building systems primarily for libraries, archives and other memory institutions, with a strong focus on automation, access to digital collections and the practical use of artificial intelligence. Presentation: From Scan to Archival Package: Automating the Digitisation Workflow Using a large collection of broadside ballads as a case study, we present an automated workflow covering the entire process from raw scans to a validated archival package. Broadside ballads represent a particularly challenging type of material: they are usually very short, often only four or eight unnumbered pages, and may also be bound together in composite volumes. For efficient digitisation, large batches are therefore scanned sequentially without identifying individual documents during scanning. The subsequent processing is largely automated. Scans containing two facing pages are detected and split into individual page images, which are then cropped and deskewed. Large sequences consisting of thousands or tens of thousands of pages are subsequently analysed and divided into individual documents. The workflow reconstructs their internal structure, detects potentially missing, duplicated, or incorrectly assigned pages, and can also preserve or identify relationships between documents that belong together. The workflow combines AI-based methods with conventional algorithmic processing to automate image processing, OCR, bibliographic metadata processing, structural metadata generation, and quality control. Existing bibliographic records are retrieved from the library catalogue, transformed into the required target formats, and enriched with information derived during digitisation. The system also detects inconsistencies between the digitised objects and catalogue data, supports the registration and assignment of national identifiers, and automatically assembles the resulting data into archival packages compliant with Czech digitisation standards, including their final validation. The case study demonstrates how individual automation components can be connected into a single processing pipeline in which human intervention is concentrated primarily on exceptional or ambiguous cases. The resulting workflow enables tens of thousands of documents to be processed efficiently and provides a reusable model that can be adapted to other types of digitised collections. |









