Published May 18, 2026 | Version v1
Other Open

HapCap: Estimating Haptic Adjectives with Vision-Language Model for Dataset Annotation

  • 1. ROR icon University of Tsukuba
  • 2. Cluster Metaverse Lab

Description

Haptic datasets capturing human haptic perceptual ratings underpin haptic texture rendering and cross-modal understanding. However, such datasets are scarce and hard to scale, because annotating objects with human perceptual ratings requires participants to physically handle each item. On the other hand, human haptic perception is closely tied to an object's visual appearance and linguistic description, suggesting that tactile attributes may be inferable without physical contact. We propose HapCap, a haptic-adjective annotation method using Vision-Language Model (VLM) trained on large-scale image-text corpora. We feed each object's photograph and name from the dataset into the model to estimate adjective ratings. To investigate how well VLMs can reproduce human haptic judgments, we benchmark three open-weight models on a haptic dataset with human perceptual ratings. All three models stayed within roughly one point of human ratings on the 1--5 scale on average. They broadly captured human haptic perception. These preliminary results position VLMs as a promising alternative that can substantially reduce the cost of haptic dataset annotation.

Files

Eguchi_EuroHaptics26_wip (2).pdf

Files (499.0 kB)

Name Size Download all
md5:404d48e82b9a0b604cf04aeb9ead8dc8
499.0 kB Preview Download