Published February 6, 2021 | Version 1.0
Dataset Open

Mask Dataset for: Unconstrained Text Detection in Manga: a New Dataset and Baseline

  • 1. Universidad de Buenos Aires
  • 2. Universidad Nacional de Luján

Description

This is the dataset used in out paper "Unconstrained Text Detection in Manga: a New Dataset and Baseline". It contains 450 images with the text segmentation of images from Manga109 dataset (need to request access to this dataset in order to view original manga image).

Pre-processed version of the images is how they were saved straight out of GIMP. These were later processed before using for training.

Post-processed version of the images is after automatically removing small connected components and filling small holes. They are also slightly bigger in width/height in order to be multiples of 8. 

The text is split in 2 colors: black and pink. Text in black represents text we consider easy to recognize, which is mostly when inside a speech bubble. Text in pink represents text we consider harder to detect, such as text in covers, sound effects or text outside speech bubbles.

Further details can be found in our paper. 

Files

post-processed.zip

Files (19.2 MB)

Name Size Download all
md5:6555af6b2aa6d80d87fe7a4ce0b886a7
11.3 MB Preview Download
md5:8d878735631bfd5fa98398ae5203a53d
7.9 MB Preview Download

Additional details

Related works

Is referenced by
Conference paper: 10.1007/978-3-030-67070-2_38 (DOI)