HapCap: Estimating Haptic Adjectives with Vision-Language Model for Dataset Annotation
Authors/Creators
Description
Haptic datasets capturing human haptic perceptual ratings underpin haptic texture rendering and cross-modal understanding. However, such datasets are scarce and hard to scale, because annotating objects with human perceptual ratings requires participants to physically handle each item. On the other hand, human haptic perception is closely tied to an object's visual appearance and linguistic description, suggesting that tactile attributes may be inferable without physical contact. We propose HapCap, a haptic-adjective annotation method using Vision-Language Model (VLM) trained on large-scale image-text corpora. We feed each object's photograph and name from the dataset into the model to estimate adjective ratings. To investigate how well VLMs can reproduce human haptic judgments, we benchmark three open-weight models on a haptic dataset with human perceptual ratings. All three models stayed within roughly one point of human ratings on the 1--5 scale on average. They broadly captured human haptic perception. These preliminary results position VLMs as a promising alternative that can substantially reduce the cost of haptic dataset annotation.
Files
Eguchi_EuroHaptics26_wip (2).pdf
Files
(499.0 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:404d48e82b9a0b604cf04aeb9ead8dc8
|
499.0 kB | Preview Download |