| dc.contributor.author | Martinc, Matej |
| dc.contributor.author | Dimitrievska, Mihaela |
| dc.date.accessioned | 2026-09-13T10:01:22Z |
| dc.date.available | 2026-09-13T10:01:22Z |
| dc.date.issued | 2026-08-07 |
| dc.identifier.uri | http://hdl.handle.net/11356/2317 |
| dc.description | This entry contains the Ilustrirani Slovenec Multimodal Document Understanding Dataset, a benchmark dataset designed for training and evaluating vision-language models on historical Slovenian newspaper pages. The dataset is derived from Ilustrirani Slovenec, a historical illustrated weekly supplement of the newspaper Slovenec, published between 1924 and 1932 and available through the Digital Library of Slovenia (dLib.si): https://www.dlib.si/details/URN:NBN:SI:spr-XWQWZFUW. A total of 410 newspaper issues were collected in PDF format and split into individual page images, resulting in 2,566 newspaper pages. Visual regions on the pages were automatically detected using Segment Anything Model 3 (SAM3; https://ai.meta.com/research/sam3/), using prompts targeting common visual elements found in historical newspapers, including photographs, portraits, illustrations, sketches, cartoons, maps, diagrams, advertisements, and posters. The resulting detections were post-processed to remove unsuitable, overlapping, and nested regions. In total, the detection pipeline identified 14,805 image regions. The detected image regions and the corresponding newspaper pages were subsequently processed using Gemini 3 Pro Preview (https://ai.google.dev/gemini-api/docs/models/gemini-3-pro-preview). For each page, the model generated structured annotations containing image-region bounding boxes, corresponding image captions where available, and page-level Slovenian text excluding image captions. The original Slovenian text was preserved without translation, normalization, or spelling correction. The generated annotations were transformed into six task-specific multimodal instruction-following datasets, each distributed with predefined training and test files: dataset1_train.jsonl and dataset1_test.jsonl contain the image region to caption task, where a bounding box corresponding to an image region is provided and the model must predict the associated image caption. dataset2_train.jsonl and dataset2_test.jsonl contain the caption to image region task, where an image caption is provided and the model must identify the corresponding image region by predicting its bounding-box coordinates. dataset3_train.jsonl and dataset3_test.jsonl contain the image region detection task, where the model must identify all image regions present on a newspaper page. dataset4_train.jsonl and dataset4_test.jsonl contain the image region detection and captioning task, where image regions must be detected and the corresponding caption generated for each detected region. dataset5_train.jsonl and dataset5_test.jsonl contain the page-level text extraction task, where all textual content on a newspaper page, including headlines, article text, and other textual elements, must be extracted while excluding image captions. dataset6_train.jsonl and dataset6_test.jsonl contain the full-page understanding task, requiring extraction of the complete structured content of a newspaper page, including page text, image-region locations, and corresponding image captions. Datasets 1 and 2 are constructed from individual image-caption pairs, while Datasets 3–6 are constructed at the page level. Of the 2,566 extracted newspaper pages, 2,556 were included in the final benchmark; 10 pages did not yield valid final annotations during the automatic processing pipeline. The benchmark is distributed with predefined training and test splits, comprising 2,300 training pages and 256 test pages. The same page-level partition is applied consistently across all six datasets so that no newspaper page appears in both the training and test partitions. Manual Quality Assessment To verify the quality of the automatically generated annotations, a subset of the test data was manually evaluated. Ten examples were randomly sampled from each of the six benchmark tasks, resulting in 60 evaluated benchmark instances. Each example was independently assessed by two annotators using two task-specific binary quality criteria, resulting in a total of 120 binary quality assessments. The two annotators agreed on 118 of the 120 assessments, corresponding to an observed agreement of 98.33%. Inter-annotator agreement was high, with Cohen's kappa of 0.8483 and Gwet's AC1 of 0.9813. Agreement was 100% for all evaluated criteria except joint localization and caption generation in Dataset 4 and structural hierarchy in Dataset 6, which both achieved 90% agreement. All manually reviewed examples that passed both task-specific quality criteria according to both annotators were additionally aggregated into manually_verified_minibench.jsonl. This file contains 53 manually verified examples across the six benchmark tasks and can be used as an additional high-confidence evaluation benchmark. Accessing the Corresponding Newspaper Pages The original Ilustrirani Slovenec PDF files and page images are not redistributed as part of this resource. The corresponding newspaper issues are available through the Digital Library of Slovenia (dLib.si): https://www.dlib.si/details/URN:NBN:SI:spr-XWQWZFUW. The Python script download_and_split_ilustrirani_slovenec.py is included for downloading the corresponding PDF issues and converting them into individual page images. The generated page images correspond to the image identifiers referenced in the six datasets and in the manually verified mini-benchmark. |
| dc.language.iso | slv |
| dc.publisher | Jožef Stefan Institute |
| dc.rights | Creative Commons - Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) |
| dc.rights.uri | https://creativecommons.org/licenses/by-nc/4.0/ |
| dc.rights.label | PUB |
| dc.subject | historical newspapers |
| dc.subject | document understanding |
| dc.subject | multimodal dataset vision-language models |
| dc.subject | OCR |
| dc.subject | page layout analysis |
| dc.subject | visual grounding |
| dc.title | Multimodal document understanding dataset from "Ilustrirani Slovenec" |
| dc.type | corpus |
| metashare.ResourceInfo#ContentInfo.mediaType | text |
| has.files | yes |
| branding | CLARIN.SI data & tools |
| contact.person | Mihaela Dimitrievska mihaeladimitrievska2702@gmail.com Jožef Stefan Institute |
| sponsor | Public Agency for Scientific Research and Innovation of the Republic of Slovenia GC-0002 Large Language Models for Digital Humanities (LLM4DH) nationalFunds |
| size.info | 2556 pages |
| size.info | 39834 texts |
| files.count | 1 |
| files.size | 7940104 |
Datoteke v tem vnosu
To je vnos
Creative Commons - Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)
Publicly Available
z licenco:Creative Commons - Attribution-NonCommercial 4.0 International (CC BY-NC 4.0)
- Ime
- Ilustrirani_Slovenec_Dataset.zip
- Velikost
- 7.57 MB
- Format
- application/zip
- Opis
- ZIP archive containing the six train/test benchmark datasets, the manually verified mini-benchmark, and the Python script for downloading and splitting the Ilustrirani Slovenec PDF issues into individual page images.
- MD5
- 87918b6eb9559a06013592577dc02bfc
- Ilustrirani_Slovenec_Dataset
- dataset5_train.jsonl-1 B
- dataset6_test.jsonl-1 B
- dataset3_train.jsonl-1 B
- dataset1_train.jsonl-1 B
- dataset5_test.jsonl-1 B
- dataset6_train.jsonl-1 B
- manually_verified_minibench.jsonl-1 B
- download_and_split_ilustrirani_slovenec.py-1 B
- dataset4_test.jsonl-1 B
- 00README.md-1 B
- dataset3_test.jsonl-1 B
- dataset4_train.jsonl-1 B
- dataset2_test.jsonl-1 B
- dataset2_train.jsonl-1 B
- dataset1_test.jsonl-1 B