Published November 5, 2021 | Version v1

Evaluating Natural Language Descriptions Generated in a Workspace-Based Architecture

  • 1. Queen Mary University of London

Description

This paper concerns the evaluation of a workspace architecture for generating natural language descriptions, including methods for evaluating both its output and its own self-evaluation. Herein are details of preliminary results from evaluation of an early iteration of the architecture operating in the domain of weather. The domain is not typically seen as creative, but provides a simple testbed for the architecture and evaluation methodology. The program does not yet match humans in terms of fluency of language, factual correctness, and how completely the input is described, but human judges did find the program’s output easier to read than human generated texts. Planned improvements to the program also described in the paper will incorporate self-monitoring and better self-evaluation with the aim of producing descriptions that are more fluently written and more accurate.

Files

ICCC_2021_paper_97.pdf

Files (253.0 kB)

Name Size
md5:7d8ee0c70acddf23485b280c39a868d1
253.0 kB Preview Download

Additional details

Funding

European Commission
EMBEDDIA - Cross-Lingual Embeddings for Less-Represented Languages in European News Media 825153