Published September 14, 2023 | Version 1.001

Report on Transformers interpretability for Natural Language Processing: A case study on Technical Debt classification

Authors/Creators

  • 1. University of Oslo, Norway

Description

Transformer models have significantly advanced the field of natural language processing (NLP), achieving exceptional results in various tasks. However, these models are often seen as "black boxes", providing limited insight into the factors influencing their predictions. It has become crucial to develop and utilise methods for interpreting and explaining these models to uncover their complex inner workings. This report discusses the latest techniques and tools that aid in a more profound understanding of transformer models within NLP. Additionally, it explores a vital industrial use case: Technical Debt (TD) classification. In this context, the report leverages transformer model interpretability tools and Retrieval Augmented Generation (RAG) to analyse and understand the characteristics of text in Github issues, distinguishing between TD and non-TD.

This report thoroughly outlines an approach to improve the transparency and reproducibility of machine learning models, with a special emphasis on TD classification. It integrates the RAG approach and exploits feature attribution techniques, presenting a route to create AI systems that are not only high-performing but also demonstrably trustworthy and comprehensible. Through a detailed examination of word patterns in TD classification and the innovative use of the RAG approach, the research highlights a strong dedication to promoting transparency and responsibility in AI systems, potentially ushering in a new phase in machine learning research that focuses on clarity and dependability.

Files

All_requirements.txt

Files (758.3 MB)

Name Size
md5:21d51136c5c73a3ad38cb2ad1bddae20
25.0 kB Preview Download
md5:83badcaddbc43fcd65077f10346d918d
42.7 MB Preview Download
md5:4e5b54dcefe0c5f075b1d99b694d43a6
715.6 MB Preview Download