There is a newer version of the record available.

Published August 7, 2021 | Version v1

Manual Corpora Development for Generative Pre-trained Transformers (GPT) & Evaluation of GPT Model Learning Capability

Description

This report describes the process of manual auditing and refining of a set of books to develop a Corpora that is suitable for GPT2 (Medium Size) training. The preprocessing criteria involved removing peripheral text, sanitizing all references to authors and commercial products, and modifying first-person references. The GPT2 model was trained on the "The Power of Now.txt" corpus using a sample ratio of 70-30 for training and validation. The training and validation losses converged at the last epoch, indicating that the model learned accurately. The model's accuracy and generated content relevance were evaluated using cosine similarity and semantic similarity tests. The publication also includes the perplexity score and details of the irrelevant content used in the evaluation.

Files

White Paper - Introduction to GPT Generative Pretrained Transformers Training and Evaluation using a Manual Corpora.pdf