Published May 21, 2026 | Version v1

Schwartz Value Annotations on Cross-Population Work-Goal Responses (Indian, Asian, IJH, Ultra)

Authors/Creators

Description

Annotated corpus of short free-text work-goal responses from four populations,
labeled with Schwartz value categories. Companion data for the paper "Lexical
Register, Not Cultural Proximity: Cross-Population Generalization of Schwartz
Value Classifiers" (under review; author names withheld for double-blind
review and will be updated here upon publication).

## Overview

2,699 annotated instances spanning four populations that vary along
**cultural** (geography, religion, language community) and **organizational**
(workplace, religious-educational, lay-survey) axes:

| Label  | Population description                                                                    |   N    
|--------|--------------------------------------------------------------|------
| Indian | Employees from India (largest sample)                                         | 1161 
| Ultra  | Female educators from the Israeli ultra-Orthodox community    |  721 
| IJH    | Employees of an Israeli Jewish humanitarian nonprofit                |  415 
| Asian  | Asian-American survey respondents                                            |  402 

The IJH and Ultra populations responded in Hebrew; their responses were
translated to English via Google Translate with manual quality control prior
to release.

## Elicitation

Each respondent first completed an adapted Schwartz Values Survey (Schwartz,
1992) and then wrote up to three short sentences describing their work goals.
The work-goal sentences (median 3 words; range 2–5 words across populations)
are what is annotated and distributed here.

## Annotation

A team of four to six researchers trained in Schwartz value theory developed a
detailed annotation guide. Each goal was independently coded by two
annotators using the 12-category coarse Schwartz scheme (SD, ST, HE, AC, PO,
FA, SE, TR, CO, HU, BE, UN), and a curator resolved disagreements.
Inter-annotator agreement on the dominant value ranges from 67% (Asian) to
88% (Indian); 75% (IJH) and 83% (Ultra).

## Files

- `merged.csv` — full corpus (2,699 rows). Columns: `Dataset`, `Text`,
  `Annotated Value`.

## Ethics

All four studies were approved by the institutional ethics review board of
the lead author's institution (committee details withheld for anonymous
review). Participation was voluntary, anonymous, and obtained with informed
consent. A two-annotator + curator review pass functioned as a screen for
personally-identifying information; no such cases were flagged. Population
labels in the released data are anonymized (e.g., the host organization for
the IJH population is referred to only by the IJH label).

## Recommended use

Cross-population evaluation of value-detection classifiers, lexical-register
diversity studies, and replication of the experiments in the companion paper.
Practitioners deploying value classifiers cross-culturally should note the
register-driven mirror-image error patterns documented in the companion
paper.

## Citation

Cite both this dataset (via its DOI) and the companion paper.

Files

dataset.csv

Files (106.0 kB)

Name Size Download all
md5:89337ecfc29adbc23e3ce0b09e1cbe5f
106.0 kB Preview Download

Additional details

Additional titles

Alternative title (English)
Data for: Lexical Register, Not Cultural Proximity: Cross-Population Generalization of Schwartz Value Classifiers