Dissociating Prompt Quality from Response Compliance in Automated Prompt-Engineering Assessment: evaluator, results and figure code
Description
MATLAB implementation, per-item numeric results and figure code for the study "Dissociating Prompt Quality from Response Compliance in Automated Prompt-Engineering Assessment".
A four-agent evaluator running on a locally hosted Qwen 2.5-7B-Instruct judge scores each prompt as an artifact before any response exists, then scores the resulting response twice on identical items: once against criteria extracted from that prompt, and once against a fixed rubric that never sees it. Verifiable constraints are checked with the IFEval reference implementation rather than model-generated labels. The whole pipeline runs on consumer CPU hardware with no external API calls.
Contents: src/ the MATLAB agents and run scripts; scripts/ the IFEval verifier and the figure code; data/ a SHA-256 manifest of all 4,275 benchmark prompts and the filtered IFEval subset; results/ per-item results for the main run (n = 500) and the ablation (n = 120), the construct-validation output and the IFEval verification output.
Corpus not redistributed. No prompt or response text from LMSYS-Chat-1M or WildChat-1M is included; the LMSYS agreement prohibits transfer to third parties and WildChat-1M carries flow-down obligations under the AI2 ImpACT licence. Prompts appear only as SHA-256 hashes, so an independently rebuilt corpus can be verified as identical to the one used here. The IFEval subset is included in full under Apache-2.0.
Files
shahoismael/prompt-quality-vs-response-compliance-v1.0.0.zip
Files
(364.5 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:1608fa100232f19b51df66c494560bd2
|
364.5 kB | Preview Download |
Additional details
Related works
- Is supplement to
- Software: https://github.com/shahoismael/prompt-quality-vs-response-compliance/tree/v1.0.0 (URL)