TOKENIZATION-BASED TRUNCATING AND PADDING METHOD TO SOLVE LONG OPCODE SEQUENCE PROBLEM IN LSTM FOR IOT MALWARE DETECTION
Description
Detecting malicious software in Internet of Things (IoT) datasets and environments remain a significant challenge for researchers striving to secure IoT networks. As the massive interconnection of Internet devices causes, most previous research has explored deep learning approaches, particularly the LSTM model, in detecting malware within operation codes (Opcodes) for ARM-based IoT applications, owing to its strong classification capabilities. Despite the advantages of using opcode features for detecting malicious software, the length of the opcode sequence poses a major challenge for deep learning methods such as the LSTM algorithm. Long opcode sequences can lead to information loss and increased computational burden, which may result in the vanishing gradient problem in the LSTM algorithm. To address this issue. This paper proposes a tokenization-based truncating and padding method to solve this problem. It shortens the length of the opcode sequence while keeping the classification performance the same. This approach extracts a subset of opcode sequences based on uniqueness, significantly reducing the length and size of the sequence while preserving the quality of the dataset. The method was evaluated on three IoT datasets from the Linux system and one dataset from the Windows system. The findings show that, for all datasets, the suggested approach performs better than current approaches in terms of accuracy and time.
Files
2Vol104No11.pdf
Files
(2.2 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:444ae6fb23ec2fa582182b90d11fedde
|
2.2 MB | Preview Download |