Manual Corpora Development for Generative Pre-trained Transformers (GPT) & Evaluation of GPT Model Learning Capability
Authors/Creators
Description
This report describes the process of manual auditing and refining of a set of books to develop a Corpora that is suitable for GPT2 (Medium Size) training. The preprocessing criteria involved removing peripheral text, sanitizing all references to authors and commercial products, and modifying first-person references. The GPT2 model was trained on the "The Power of Now.txt" corpus using a sample ratio of 70-30 for training and validation. The training and validation losses converged at the last epoch, indicating that the model learned accurately. The model's accuracy and generated content relevance were evaluated using cosine similarity and semantic similarity tests. The publication also includes the perplexity score and details of the irrelevant content used in the evaluation.