Published December 12, 2022 | Version v1

Semantic Similarity of IT Support Tickets

Description

Collection of 300 support tickets manually labeled for semantic similarity, obtained from a IT support company in the Florianópolis (Brazil) region. Each ticket is represented by an unstructured text field, which is typed by the user that opened the call. The labeling process was performed in 2022 by three IT support professionals. The corpus contains tickets in many languages, mainly English, German, Portuguese and Spanish.

All Personal Identifiable Information (PII) and sensitive information were removed (substituted by a tag indicating the original content, for instance: the sentence "this text was written by Leonardo" is converted to "this text was written by [NAME]"). The removal was performed in three steps: first, the automated machine learning-based tool AWS Comprehend PII Removal was used; then, a sequence of custom regular expressions was applied; last, the entire corpus was manually verified.

Files

group_1.csv

Files (102.0 kB)

Name Size Download all
md5:f6f1d1dda7b336260667d8c2d6f05d38
25.4 kB Preview Download
md5:8aa91247734fb89b89911531c3769ad5
4.4 kB Preview Download
md5:320b4e121968e7f03a79f4bbe99f7b63
30.7 kB Preview Download
md5:4316ee220f545188106ef07604f078ee
4.4 kB Preview Download
md5:bb8ccc04e33cb4bab19307f9651b3689
32.7 kB Preview Download
md5:5204ee018ef7f6e93a7f6f2aed8b336d
4.5 kB Preview Download