Show simple item record

 
dc.contributor.author Martinc, Matej
dc.contributor.author Dimitrievska, Mihaela
dc.date.accessioned 2026-09-13T10:01:22Z
dc.date.available 2026-09-13T10:01:22Z
dc.date.issued 2026-08-07
dc.identifier.uri http://hdl.handle.net/11356/2317
dc.description This entry contains the Ilustrirani Slovenec Multimodal Document Understanding Dataset, a benchmark dataset designed for training and evaluating vision-language models on historical Slovenian newspaper pages. The dataset is derived from Ilustrirani Slovenec, a historical illustrated weekly supplement of the newspaper Slovenec, published between 1924 and 1932 and available through the Digital Library of Slovenia (dLib.si): https://www.dlib.si/details/URN:NBN:SI:spr-XWQWZFUW. A total of 410 newspaper issues were collected in PDF format and split into individual page images, resulting in 2,566 newspaper pages. Visual regions on the pages were automatically detected using Segment Anything Model 3 (SAM3; https://ai.meta.com/research/sam3/), using prompts targeting common visual elements found in historical newspapers, including photographs, portraits, illustrations, sketches, cartoons, maps, diagrams, advertisements, and posters. The resulting detections were post-processed to remove unsuitable, overlapping, and nested regions. In total, the detection pipeline identified 14,805 image regions. The detected image regions and the corresponding newspaper pages were subsequently processed using Gemini 3 Pro Preview (https://ai.google.dev/gemini-api/docs/models/gemini-3-pro-preview). For each page, the model generated structured annotations containing image-region bounding boxes, corresponding image captions where available, and page-level Slovenian text excluding image captions. The original Slovenian text was preserved without translation, normalization, or spelling correction. The generated annotations were transformed into six task-specific multimodal instruction-following datasets, each distributed with predefined training and test files: dataset1_train.jsonl and dataset1_test.jsonl contain the image region to caption task, where a bounding box corresponding to an image region is provided and the model must predict the associated image caption. dataset2_train.jsonl and dataset2_test.jsonl contain the caption to image region task, where an image caption is provided and the model must identify the corresponding image region by predicting its bounding-box coordinates. dataset3_train.jsonl and dataset3_test.jsonl contain the image region detection task, where the model must identify all image regions present on a newspaper page. dataset4_train.jsonl and dataset4_test.jsonl contain the image region detection and captioning task, where image regions must be detected and the corresponding caption generated for each detected region. dataset5_train.jsonl and dataset5_test.jsonl contain the page-level text extraction task, where all textual content on a newspaper page, including headlines, article text, and other textual elements, must be extracted while excluding image captions. dataset6_train.jsonl and dataset6_test.jsonl contain the full-page understanding task, requiring extraction of the complete structured content of a newspaper page, including page text, image-region locations, and corresponding image captions. Datasets 1 and 2 are constructed from individual image-caption pairs, while Datasets 3–6 are constructed at the page level. Of the 2,566 extracted newspaper pages, 2,556 were included in the final benchmark; 10 pages did not yield valid final annotations during the automatic processing pipeline. The benchmark is distributed with predefined training and test splits, comprising 2,300 training pages and 256 test pages. The same page-level partition is applied consistently across all six datasets so that no newspaper page appears in both the training and test partitions. Manual Quality Assessment To verify the quality of the automatically generated annotations, a subset of the test data was manually evaluated. Ten examples were randomly sampled from each of the six benchmark tasks, resulting in 60 evaluated benchmark instances. Each example was independently assessed by two annotators using two task-specific binary quality criteria, resulting in a total of 120 binary quality assessments. The two annotators agreed on 118 of the 120 assessments, corresponding to an observed agreement of 98.33%. Inter-annotator agreement was high, with Cohen's kappa of 0.8483 and Gwet's AC1 of 0.9813. Agreement was 100% for all evaluated criteria except joint localization and caption generation in Dataset 4 and structural hierarchy in Dataset 6, which both achieved 90% agreement. All manually reviewed examples that passed both task-specific quality criteria according to both annotators were additionally aggregated into manually_verified_minibench.jsonl. This file contains 53 manually verified examples across the six benchmark tasks and can be used as an additional high-confidence evaluation benchmark. Accessing the Corresponding Newspaper Pages The original Ilustrirani Slovenec PDF files and page images are not redistributed as part of this resource. The corresponding newspaper issues are available through the Digital Library of Slovenia (dLib.si): https://www.dlib.si/details/URN:NBN:SI:spr-XWQWZFUW. The Python script download_and_split_ilustrirani_slovenec.py is included for downloading the corresponding PDF issues and converting them into individual page images. The generated page images correspond to the image identifiers referenced in the six datasets and in the manually verified mini-benchmark.
dc.language.iso slv
dc.publisher Jožef Stefan Institute
dc.rights Creative Commons - Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)
dc.rights.uri https://creativecommons.org/licenses/by-nc/4.0/
dc.rights.label PUB
dc.subject historical newspapers
dc.subject document understanding
dc.subject multimodal dataset vision-language models
dc.subject OCR
dc.subject page layout analysis
dc.subject visual grounding
dc.title Multimodal document understanding dataset from "Ilustrirani Slovenec"
dc.type corpus
metashare.ResourceInfo#ContentInfo.mediaType text
has.files yes
branding CLARIN.SI data & tools
contact.person Mihaela Dimitrievska mihaeladimitrievska2702@gmail.com Jožef Stefan Institute
sponsor Public Agency for Scientific Research and Innovation of the Republic of Slovenia GC-0002 Large Language Models for Digital Humanities (LLM4DH) nationalFunds
size.info 2556 pages
size.info 39834 texts
files.count 1
files.size 7940104


 Files in this item

This item is
Publicly Available
and licensed under:
Creative Commons - Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)
Distributed under Creative Commons Attribution Required Noncommercial
Icon
Name
Ilustrirani_Slovenec_Dataset.zip
Size
7.57 MB
Format
application/zip
Description
ZIP archive containing the six train/test benchmark datasets, the manually verified mini-benchmark, and the Python script for downloading and splitting the Ilustrirani Slovenec PDF issues into individual page images.
MD5
87918b6eb9559a06013592577dc02bfc
 Download file  Preview
 File Preview  
  • Ilustrirani_Slovenec_Dataset
    • dataset5_train.jsonl-1 B
    • dataset6_test.jsonl-1 B
    • dataset3_train.jsonl-1 B
    • dataset1_train.jsonl-1 B
    • dataset5_test.jsonl-1 B
    • dataset6_train.jsonl-1 B
    • manually_verified_minibench.jsonl-1 B
    • download_and_split_ilustrirani_slovenec.py-1 B
    • dataset4_test.jsonl-1 B
    • 00README.md-1 B
    • dataset3_test.jsonl-1 B
    • dataset4_train.jsonl-1 B
    • dataset2_test.jsonl-1 B
    • dataset2_train.jsonl-1 B
    • dataset1_test.jsonl-1 B

Show simple item record