Deception Archive: Standardized Verbal Deception Datasets
Authors/Creators
- 1. Magna Græcia University of Catanzaro, Italy
- 2. Tilburg University, The Netherlands
Description
The Deception Archive is a curated repository of verbal deception datasets identified through a systematic review of the computational verbal deception literature. The Archive was developed through a structured workflow involving dataset identification, retrieval, manual inspection, metadata curation, and data standardization. The current release contains 42 standardized datasets organized according to a common data structure, while preserving the original textual content and veracity annotations and providing additional standardized metadata describing dataset characteristics. The Archive is designed to facilitate dataset discovery, exploration, cross-study comparison, and reuse in computational research on verbal deception.
Files
Dataset_Deception Archive.zip
Additional details
References
- Skalicky, S., Duran, N. D., & Crossley, S. A. (2020). Please, Please, Just Tell Me: The Linguistic Features of Humorous Deception. Dialogue & Discourse, 11(2), 128–149. https://doi.org/10.5210/dad.2020.205
- Salvetti, F., Lowe, J. B., & Martin, J. H. (2016). A tangled web: The faint signals of deception in text—Boulder Lies and Truth Corpus (BLT-C). In Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC 2016) (pp. 3510–3517). European Language Resources Association (ELRA).
- Soldner, F., Pérez-Rosas, V., & Mihalcea, R. (2019). Box of lies: Multimodal deception detection in dialogues. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers) (pp. 1768–1777). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1175
- Verhoeven, B., & Daelemans, W. (2014). CLiPS Stylometry Investigation (CSI) corpus: A Dutch corpus for the detection of age, gender, personality, sentiment and deception in text. In Proceedings of the Ninth International Conference on Language Resources and Evaluation
- Hirschberg, J., Benus, S., Brenier, J. M., Enos, F., Friedman, S., Gilman, S., Girand, C., Graciarena, M., Kathol, A., Michaelis, L., Pellom, B., Shriberg, E., & Stolcke, A. (2005). Distinguishing deceptive from non-deceptive speech. In Proceedings of Interspeech 2005 (pp. 1833–1836). International Speech Communication Association (ISCA).
- Pérez-Rosas, V., & Mihalcea, R. (2014). Cross-cultural deception detection. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers) (pp. 440–445). Association for Computational Linguistics. https://doi.org/10.3115/v1/P14-2072
- Fornaciari, T., & Poesio, M. (2014). Identifying fake Amazon reviews as learning from crowds. In Proceedings of the 14th Conference of the European Chapter of the Association for Computational Linguistics (pp. 279–287). Association for Computational Linguistics. https://doi.org/10.3115/v1/E14-1030
- Fornaciari, T., Cagnina, L., Rosso, P., Hernández-Farías, D. I., Patti, V., Pardo, F. M. E., Ruini, M., & Banerjee, S. (2020). Fake opinion detection: How similar are crowdsourced datasets to real data? Language Resources and Evaluation, 54(4), 1019–1058. https://doi.org/10.1007/s10579-020-09486-5
- Ott, M., Choi, Y., Cardie, C., & Hancock, J. T. (2011). Finding deceptive opinion spam by any stretch of the imagination. In Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies (pp. 309–319). Association for Computational Linguistics.
- Capuozzo, P., Lauriola, I., Strapparava, C., Aiolli, F., & Sartori, G. (2020). DecOp: A multilingual and multi-domain corpus for detecting deception in typed text. In Proceedings of the Twelfth Language Resources and Evaluation Conference (pp. 1423–1430). European Language Resources Association (ELRA).
- Fornaciari, T., & Poesio, M. (2012). DeCour: A corpus of deceptive statements in Italian courts. In Proceedings of the Eighth International Conference on Language Resources and Evaluation (LREC'12) (pp. 1585–1590). European Language Resources Association (ELRA).
- Velutharambath, A., Wührl, A., & Klinger, R. (2024). Can factual statements be deceptive? The DeFaBel corpus of belief-based deception. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 2708–2723). ELRA and the International Committee on Computational Linguistics (ICCL).
- Sap, M., Jafarpour, A., Choi, Y., Smith, N. A., Pennebaker, J. W., & Horvitz, E. (2022). Quantifying the narrative flow of imagined versus autobiographical stories. Proceedings of the National Academy of Sciences, 119(45), Article e2211715119. https://doi.org/10.1073/pnas.2211715119
- Kleinberg, B., & Verschuere, B. (2021). How humans impair automated deception detection performance. Acta Psychologica, 213, Article 103250. https://doi.org/10.1016/j.actpsy.2020.103250
- Li, J., Ott, M., Cardie, C., & Hovy, E. (2014). Towards a general rule for identifying deceptive opinion spam. In Proceedings of the 52nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 1566–1576). Association for Computational Linguistics. https://doi.org/10.3115/v1/P14-1147
- Ibraheem, S., Zhou, G., & DeNero, J. (2022). Putting the con in context: Identifying deceptive actors in the game of Mafia. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 158–168). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.naacl-main.13
- de Ruiter, B., & Kachergis, G. (2018). The Mafiascum dataset: A large text corpus for deception detection. arXiv. https://doi.org/10.48550/arXiv.1811.07851
- Lloyd, E. P., Deska, J. C., Hugenberg, K., McConnell, A. R., Humphrey, B. T., & Kunstman, J. W. (2019). Miami University deception detection database. Behavior Research Methods, 51(1), 429–439. https://doi.org/10.3758/s13428-018-1061-4
- Monaro, M., Maldera, S., Scarpazza, C., Sartori, G., & Navarin, N. (2022). Detecting deception through facial expressions in a dataset of videotaped interviews: A comparison between human judges and machine learning models. Computers in Human Behavior, 127, Article 107063. https://doi.org/10.1016/j.chb.2021.107063
- Ott, M., Cardie, C., & Hancock, J. T. (2013). Negative deceptive opinion spam. In Proceedings of the 2013 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 497–501). Association for Computational Linguistics.
- Pérez-Rosas, V., & Mihalcea, R. (2015). Experiments in open domain deception detection. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (pp. 1120–1125). Association for Computational Linguistics. https://doi.org/10.18653/v1/D15-1130
- Pérez-Rosas, V., Abouelenien, M., Mihalcea, R., & Burzo, M. (2015). Deception detection using real-life trial data. In Proceedings of the 2015 ACM on International Conference on Multimodal Interaction (ICMI '15) (pp. 59–66). Association for Computing Machinery. https://doi.org/10.1145/2818346.2820758
- Banerjee, R., Feng, S., Kang, J.S., Choi, Y.: Keystroke patterns as prosody in digi- tal writings: A case study with deceptive reviews and essays. In: Proceedings of the 2014 conference on empirical methods in natural language processing (EMNLP). pp. 1469–1473 (2014)
- Soldner, F., Kleinberg, B., & Johnson, S. D. (2022). Confounds and overestimations in fake review detection: Experimentally controlling for product-ownership and data-origin. PLOS ONE, 17(12), e0277869. https://doi.org/10.1371/journal.pone.0277869
- Pérez-Rosas, V., Abouelenien, M., Mihalcea, R., Xiao, Y., Linton, C. J., & Burzo, M. (2015). Verbal and nonverbal clues for real-life deception detection. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing (pp. 2336–2346). Association for Computational Linguistics. https://doi.org/10.18653/v1/D15-1274
- Hsiao, S.-W., & Sun, C.-Y. (2023). LoRA-like calibration for multimodal deception detection using ATSFace data. arXiv. https://doi.org/10.48550/arXiv.2309.01383
- Sarzynska-Wawer, J., Pawlak, A., Szymanowska, J., Hanusz, K., & Wawer, A. (2023). Truth or lie: Exploring the language of deception. PLOS ONE, 18(2), e0281179. https://doi.org/10.1371/journal.pone.0281179
- Kleinberg, B., van der Toolen, Y., Vrij, A., Arntz, A., & Verschuere, B. (2018). Automated verbal credibility assessment of intentions: The model statement technique and predictive modeling. Applied Cognitive Psychology, 32(3), 354–366. https://doi.org/10.1002/acp.3407
- Hayat, U., Saeed, A., Vardag, M. H. K., Ullah, M. F., & Iqbal, N. (2022). Roman Urdu fake reviews detection using stacked LSTM architecture. SN Computer Science, 3(6), Article 470. https://doi.org/10.1007/s42979-022-01385-6
- Boumber, D. A., Qachfar, F. Z., & Verma, R. (2024). Domain-agnostic adapter architecture for deception detection: Extensive evaluations with the DIFrauD benchmark. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) (pp. 5260–5274). ELRA and the International Committee on Computational Linguistics (ICCL).
- Loconte, R., & Kleinberg, B. (2025). Examining embedded lies through computational text analysis. Scientific Reports, 15(1), Article 26482. https://doi.org/10.1038/s41598-025-11327-w
- Kim, S., Lee, S., Park, D., & Kang, J. (2017). Constructing and evaluating a novel crowdsourcing-based paraphrased opinion spam dataset. In Proceedings of the 26th International Conference on World Wide Web (WWW '17) (pp. 827–836). International World Wide Web Conferences Steering Committee. https://doi.org/10.1145/3038912.3052607
- Spyridis, Y., Younes, J.-P., Deeb, H., & Argyriou, V. (2024). Empowering prior to court legal analysis: A transparent and accessible dataset for defensive statement classification and interpretation. arXiv. https://doi.org/10.48550/arXiv.2405.10702
- Cormack, G. V., & Lynam, T. R. (2005). TREC 2005 Spam Track overview. In Proceedings of the Fourteenth Text REtrieval Conference (TREC 2005). National Institute of Standards and Technology (NIST).
- Peskov, D., Cheng, B., Elgohary, A., Barrow, J., Danescu-Niculescu-Mizil, C., & Boyd-Graber, J. (2020). It takes two to lie: One to lie, and one to listen. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 3811–3854). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.349
- Baldivas, R. I. A., Sreenivasan, N., Kang, S. Y., Miller, A. M.-L., Chacko, M., Krishnan, S., Ayala, C., Ayala, E., & Kim, D. (2025). LegalEye: Multimodal court deception detection across multiple languages. Behavioral Sciences, 15(12), Article 1707. https://doi.org/10.3390/bs15121707
- Himdi, H., & Alhayan, F. (2026). Optimized ensemble stacking approaches to detect Arabic phishing email. Journal of Engineering Research, 14(1), 776–788. https://doi.org/10.1016/j.jer.2025.08.013
- Patra, C., Giri, D., Kundu, B., Maitra, T., & Wazid, M. (2025). Rhetorical Structure Theory-based machine intelligence-driven deceptive phishing attack detection scheme. Journal of Information Security and Applications, 94, Article 104184. https://doi.org/10.1016/j.jisa.2025.104184
- Sharma, N., Gogineni, V., Budda, R. M., Gupta, K., Jiwani, N., & Gupta, U. (2025). Gradient boosting decision trees for real-time phishing attack prevention in cybersecurity. In 2025 10th International Conference on Smart Structures and Systems (ICSSS). IEEE. https://doi.org/10.1109/ICSSS66939.2025.11346380
- Mendes, P., Maia, E., & Praça, I. (2025). MeAJOR corpus: A multi-source dataset for phishing email detection. arXiv. https://doi.org/10.48550/arXiv.2507.17978
- Velutharambath, A., Sassenberg, K., & Klinger, R. (2026). What if deception cannot be detected? A cross-linguistic study on the limits of deception detection from text. Computational Linguistics. https://doi.org/10.1162/COLI.a.614