Published June 1, 2020 | Version v1

LingConc: A Tool for Linguistics, Learning, and Digital Humanities with Top-down Access

  • 1. Department of Computer Science, National Tsing Hua University

Description

This paper presents LingConc, a corpus reading and analysis system for a user-provided corpus, supporting flexible queries with wild-card capability, displaying search results top-down according to key-phrase frequency, extracting keywords, lexical bundles and linking Wikipedia information. At run time, the system compiles a corpus from user submitted pdf files or URLs, and automatically retrieves documents summary, including the word list, keyword list, collocation list and lexical bundles sorted by frequency. Clicking on a word or phrase demands to display sentences containing the target word, along with the document names. With a click on a document name, users can read the content in an interactive environment, retrieve relevant information (e.g., word definitions, collocations, grammar patterns and Wikipedia information) on demand (also see Linggle Booster) (Chen, et al., 2019). Additionally, the system also includes the functions of Linggle (Boisson et al., 2013), an existing linguistic search engine on 1 trillion words in public Web pages (https://linggle.com) that retrieves lexical bundles in response to a given query. The query might contain keywords, wildcards, wild parts of speech (PoS), and thus users can compose a query fit their own search purposes, displaying the results top-down according to key-phrase frequency. The supported functions are presented in Table 1. One click on a ngram result triggers pop-ups of examples tailed by the source.

Files

LingConc_DH_poster.pdf

Files (1.1 MB)

Name Size Download all
md5:1ace05b52e7c8b6f8845cb5402497bf4
1.1 MB Preview Download