Published November 12, 2018
| Version v1
Dataset
Open
OpenNLP tokenization model for Picard
Description
OpenNLP tokenization model for Picard, trained on the Restaure corpus.
The apostrophes must be standardized in the input file: l’bas -> l'bas
To tokenize a file: <input-file.txt opennlp TokenizerME pcd-token.bin
Files
Files
(40.6 kB)
| Name | Size | Download all |
|---|---|---|
|
md5:33fca5aab973e787ef3ff6bbc0c2fb57
|
40.6 kB | Download |