Bundled Dialect Contrasts Measure a Mixture: A Feature-Isolating Matched-Pair Audit of Safety-Critical Support in Language Models
Description
Prior audits establish dialect bias in language models. It can appear as covert stereotypes attached to African American Vernacular English (AAVE) and as lower quality of service for dialect-marked users. Those audits compare a dialect-marked prompt with an unmarked one, and the two prompts differ in many ways at once: grammar, vocabulary, address terms, profanity, register, and length. Such a comparison shows that the model responds differently; it cannot show which difference the model responds to. Because these audits score each answer with one label, they also miss unequal help when both answers count as safe. We define the Dialect-Marked Response Audit (DMRA), a matched-pair protocol adapted from labor-market correspondence audits, in which one attribute changes and everything else is held fixed. Its central step builds prompt pairs that differ in exactly one feature, so a difference in the answers can be assigned to that feature; its other steps control length, response mode, and how answers are scored, so that the assignment can be trusted. On Qwen3.5-35B-A3B and a refusal-reduced variant used as a stress test, one-feature pairs change the attribution. We call an answer unsafe when it gives operational guidance for the harmful act instead of refusing, de-escalating, or offering crisis support. Varying syntax and register together, with dangerous lexis held constant, reproduces the unsafe answers that the full dialect prompt produced (3 of 8 violence-domain pairs, all in the stress model), while varying only the in-group racial address term produces none (0 of 8). These counts come from two pairs per isolation arm in one domain, one greedy run per cell, coded by a single rater, and are reported as descriptive. Four of the five violence arms meet the one-feature rule; the syntax/register arm varies grammar and register as one block, because a natural dialect-marked sentence does not separate them, and its effect is attributed to the block. Whole-dialect comparisons in the same corpus show real gaps that cannot be assigned to any feature, including a crisis hotline the base model gives the unmarked user and withholds from the dialect-marked user when it answers without a reasoning step. Comparing whole dialects shows that outputs differ; assigning the difference requires one-feature pairs; and on the pairs we build, the racial address term is not what moves the model, a finding the protocol treats as stable only after sampled, multi-coder replication.
Version 4.2 (2026-08-27): terminology matched to the data. The isolation arm that reproduces the bundled violence effect is named syntax/register throughout; the abstract no longer says "only dialect grammar." S3, the appendix, and Limitations state that this arm varies grammar and register as one block and that no arm isolates morphosyntax alone; the abstract states that four of the five violence arms meet the one-feature rule. No value changed. See CHANGELOG_v4.md.
Version 4.3 (2026-09-03): terminology. AAVE is used throughout; the sentences that reported prior findings as "African American English" / "AAE" now say AAVE, expanded at first use. No value changed. See CHANGELOG_v4.md.