Published December 3, 2021 | Version v1.2

BioVAE: a pre-trained latent variable language model for biomedical text mining

Description

We release BioVAE, the first large-scale pre-trained latent variable language model for the biomedical domain, which uses the OPTIMUS framework to train on large volumes of biomedical text.

This version contains the pre-trained models for text mining tasks such as named entity recognition or relation extraction, and text generation task.

Explanation of each file: (lt32: latent_size = 32, beta05: beta=0.5)

  • pm-full-lt32-beta00
  • pm-full-lt32-beta05
  • pm-full-lt768-beta00
  • pm-full-lt768-beta05
  • pm-full-generation

Files

Files (5.2 GB)

Name Size
md5:75ebc5026a5e8850265799e128d3e063
3.6 GB Download
md5:72a862847b94cfca44176e960a5ce64a
408.2 MB Download
md5:3a553eb9131242104100e58fa7ac6737
408.2 MB Download
md5:d4e2febce9c9d8a75aad9a5dfaefc48f
412.4 MB Download
md5:40c2f83b6ebe04ede1efd234625981ae
412.5 MB Download

Additional details