There is a newer version of the record available.

Published May 18, 2026 | Version v3

Efficient globin production during terminal erythropoiesis depends on the cooperative action of TENT5C poly(A) polymerase and LARP4B

  • 1. ROR icon International Institute of Molecular and Cell Biology
  • 2. Laboratory of RNA Biology – ERA Chairs Group, International Institute of Molecular and Cell Biology, Warsaw, Poland
  • 3. Clinical Research Center, Medical University of Białystok, Białystok, Poland
  • 4. Laboratory of Iron Homeostasis, International Institute of Molecular and Cell Biology, Warsaw, Poland

Description



This repository contains supplementary data accompanying related publication.

 Our preprint can be found here: https://doi.org/10.1101/2024.11.14.623596

Abstract

Red blood cell development is a unique process in which reduced transcriptome and proteome complexity enable extensive hemoglobin production. Here, we describe the cooperative roles of cytoplasmic poly(A) polymerase TENT5C and the poly(A) tail-protecting LARP4B RNA-binding protein in ensuring proper hemoglobin synthesis. TENT5C catalytic mutant knock-in mice exhibit microcytic hypochromic anemia similar to the constitutive knockout. Through poly(A) tail extension, TENT5C counteracts the gradual deadenylation of globin mRNA during erythropoiesis. In the late stages, TENT5C dysfunction results in globin poly(A) tail shortening and a pronounced reduction in mRNA levels in reticulocytes. Proteomic experiments revealed a transient but specific association of TENT5C with LARP4B. Consistent with this interaction, LARP4B depletion resulted in reduced globin mRNA abundance and shortened poly(A) tails, which proves a novel physiological role for this RNA binding protein. Furthermore, we show that TENT5C is a highly unstable protein whose stability is partially dependent on CNOT4, a deadenylase-associated E3 ubiquitin ligase.

Table of contents

  1. Tables
  2. Raw data
  3. Figures
    • Graphical abstract.pdf - a visual summary of the study
    • Fig.1.pdf - TENT5C catalytic activity is required for normal erythropoiesis.
    • Fig.2.pdf - TENT5C inactivity induces splenic stress erythropoiesis.
    • Fig.3.pdf - TENT5C counteracts globin mRNA degradation during late erythropoiesis
    • Fig.4.pdf - Globin poly(A) dynamics during erythropoiesis. 
    • Fig.5.pdf - LARP4B as a novel interactor determining stability of TENT5C
    • Fig.6.pdf - Functional association of LARP4A/4B with globin mRNAs stability regulation
    • Fig.7.pdf - CNOT4 drives TENT5C instability
  4. Supplementary figures
    • Fig.S1.pdf - TENT5C-dependent anemic phenotype is not related to iron pathways
    • Fig.S2.pdf - HBA, HBB, PCBP1, PCBP2 and TENT5C expression ex vivo and evaluation of putative TENT5C substrates suggested by Yang et al.
    • Fig.S3.pdf - Sequence alignment of mouse hemoglobin transcripts. Reference according to Gencode VM32, generated using MAFFT (v7.511)
    • Fig.S4.pdf - TurboID supplementary information
    • Fig.S5.pdf - Supplementary data on shRNA silencing
    • Fig.S6.pdf - Globin expression pattern during LARP4A, LARP4B and/or TENT5C shRNA silencing
    • Fig.S7.pdf - TENT5C degradation pathways and supplementary data on CNOT4 silencing
    • Fig.S8_gating ABC.pdf - Gating strategy used in experiments (part 1)
    • Fig.S8_gating DE.pdf - Gating strategy used in experiments (part 2)
  5. Additional resources (raw data underlying figures etc.)

 

Technical info

Detailed description of provided files:

  • Supplementary_table_1_Differential_adenylation_E14.5.xlsx - Excel file with differential adenylation results (WT vs Tent5c KO, E14.5 FLEB). Statistical significance was assessed using the two-sided Wilcoxon signed-rank test (α = 0.05). Table contains the following columns:
    • ensembl_transcript - transcript identifier in Ensembl format  
    • p.value - statistical significance calculated using a two-sided Wilcoxon rank-sum test with alpha = 0.05 
    • stats_code - a quality indicator showing whether read coverage in both conditions was adequate to support reliable statistical inference  
    • cohen_d - effect size (an auxiliary metric that helps discern transcripts with differences in poly-A tail length between conditions, even when statistical significance may be driven primarily by high read counts) 
    • *_counts - number of mapped reads in WT and mutant samples, respectively
    • *_polya_gm_mean - geometric mean of poly(A) tail length in each condition, respectively
    • length_diff - a measure of the change in poly(A) tail length between WT and mutant samples [nt]
    • fold_change - the magnitude of length_diff
    • padj - adjusted p-value (FDR-corrected) controlling for multiple testing significance
    • effect_size - descriptive measure of the magnitude of the change
    • significance - categorical label (e.g., FDR<0.05 / NotSig) based on padj threshold indicating whether the observed difference is statistically significant
    • ensembl_gene_id - gene identifier corresponding to the transcript in Ensembl format  
    • transcript_biotype - classification of transcript type (protein_coding, lncRNA, pseudogene, etc.).
    • description - short functional summary of each gene 
    • gene_name - gene symbol
  • Supplementary_table_2_Differential_adenylation_BasoE_OrthoE_PolyE_Retc.xlsx - Excel file with differential adenylation results (each sheet contains one  comparison for BasoE, PolyE, OrthoE, Retc for WT vs Tent5c catalytic mutant, respectively). Statistical significance was assessed using the two-sided Wilcoxon signed-rank test (α = 0.05). Each of the subtable (sheet) contains the following columns:
    • ensembl_transcript - transcript identifier in Ensembl format  
    • p.value - statistical significance calculated using a two-sided Wilcoxon rank-sum test with alpha = 0.05 
    • stats_code - a quality indicator showing whether read coverage in both conditions was adequate to support reliable statistical inference  
    • cohen_d - effect size (an auxiliary metric that helps discern transcripts with differences in poly-A tail length between conditions, even when statistical significance may be driven primarily by high read counts) 
    • *_counts - number of mapped reads in WT and mutant samples, respectively
    • *_polya_gm_mean - geometric mean of poly(A) tail length in each condition, respectively
    • length_diff - a measure of the change in poly(A) tail length between WT and mutant samples [nt]
    • fold_change - the magnitude of length_diff
    • padj - adjusted p-value (FDR-corrected) controlling for multiple testing significance
    • effect_size - descriptive measure of the magnitude of the change
    • significance - categorical label (e.g., FDR<0.05 / NotSig) based on padj threshold indicating whether the observed difference is statistically significant
    • ensembl_gene_id - gene identifier corresponding to the transcript in Ensembl format  
    • transcript_biotype - classification of transcript type (protein_coding, lncRNA, pseudogene, etc.).
    • description - short functional summary of each gene 
    • gene_name - gene symbol
  • Supplementary_table_3_TurboID_assay.xlsx - Excel file with TurboID analysis summaries. Contains the following sheets:
    • all proteins - all identified proteins in the TurboID assay. The table reports results for each identified protein, including basic annotation (gene name, protein ID, description, organism, sequence length, and coverage), quality and identification metrics (protein existence, probability, top peptide probability, number of peptides), and detailed quantitative measurements across replicates. Quantitative data include spectral counts, unique spectral counts, total intensities, unique intensities, MaxLFQ intensities, and relative intensities for both Tent5c and WT samples, enabling comparison of protein enrichment and reproducibility across biological replicates.
    • replicates overlap - overlap between 3 replicates. Table summarizes the quantitative and qualitative data from the TurboID proteomics experiment for each identified protein. It includes protein annotation (gene, protein ID, sequence length, coverage, organism, description), identification confidence metrics (protein existence, probabilities, peptide counts), and detailed quantitative measurements across all biological replicates. Quantitative data encompass spectral counts, unique counts, total and unique intensities, MaxLFQ intensities, and relative abundance and specificity metrics for Tent5c and WT samples.
    • GO terms - statistically significant GO terms for overlapping proteins with average specificity > 0. Contains following columns:
      • significant – indicates whether the GO term is statistically significant (TRUE/FALSE)
      • p_value – raw p-value for the enrichment of the term
      • term_size – total number of genes associated with the GO term in the reference set
      • query_size – number of genes in the query set tested for enrichment
      • intersection_size – number of genes shared between the query set and the GO term
      • precision – proportion of query genes in the intersection relative to the query set (intersection_size/query_size)
      • recall – proportion of term genes captured by the intersection (intersection_size/term_size)
      • term_id – GO identifier for the term (e.g., GO:0043604)
      • source – ontology source (e.g., GO:BP for Biological Process)
      • term_name – descriptive name of the GO term
      • effective_domain_size – total number of genes in the background/reference used for enrichment calculation
      • source_order – internal ordering or index of the GO term within the source database
    • differential abundance - proteins with estimated changes and associated statistics. Contains the following columns:
      • gene – gene symbol of the protein
      • protein_id – database identifier 
      • comparison – experimental contrast tested 
      • missingness – type of missing data: MNAR, MAR, or complete
      • diff – estimated difference in abundance between groups
      • CI_2.5 – lower bound of 95% confidence interval for diff
      • CI_97.5 – upper bound of 95% confidence interval for diff
      • avg_abundance – average abundance across all samples
      • t_statistic – t-statistic for the comparison
      • pval – raw p-value
      • adj_pval – multiple testing–corrected p-value
      • B – log-odds of differential abundance 
      • n_obs – number of observations used per protein
  • Supplementary_table_4_Differential_adenylation_Tent5c_Larp4ab_silencing.xlsx - Excel file with differential adenylation results (each sheet contains one  comparison for control shRNA vs given silencing with either one or multiple shRNA, respectively). Statistical significance was assessed using the two-sided Wilcoxon signed-rank test (α = 0.05). Each of the subtable (sheet) contains the following columns:
    • ensembl_transcript - transcript identifier in Ensembl format  
    • p.value - statistical significance calculated using a two-sided Wilcoxon rank-sum test with alpha = 0.05 
    • stats_code - a quality indicator showing whether read coverage in both conditions was adequate to support reliable statistical inference  
    • cohen_d - effect size (an auxiliary metric that helps discern transcripts with differences in poly-A tail length between conditions, even when statistical significance may be driven primarily by high read counts) 
    • *_counts - number of mapped reads in control and silenced samples, respectively
    • *_polya_gm_mean - geometric mean of poly(A) tail length in each condition, respectively
    • length_diff - a measure of the change in poly(A) tail length between control and silenced samples [nt]
    • fold_change - the magnitude of length_diff
    • padj - adjusted p-value (FDR-corrected) controlling for multiple testing significance
    • effect_size - descriptive measure of the magnitude of the change
    • significance - categorical label (e.g., FDR<0.05 / NotSig) based on padj threshold indicating whether the observed difference is statistically significant
    • ensembl_gene_id - gene identifier corresponding to the transcript in Ensembl format  
    • transcript_biotype - classification of transcript type (protein_coding, lncRNA, pseudogene, etc.).
    • description - short functional summary of each gene 
    • gene_name - gene symbol
  • Supplementary_table_5_Degronopedia_analysis.xlsx - Excel file with Degronopedia analysis overview:
    • degronopedia_overview - Overview of used input and method details used in degronopedia analysis
    • degron_data - Identified degron motifs
    • degron_conservation - Degron conservation
    • consurf_evo_conservation - Evolutionary conservation of TENT5C aa sequence scored by ConSurf (The table shows the residue variety in % for each position in the query sequence, each column shows the % for that amino acid found in position in the MSA).
  • Supplementary_table_6_Key_resources.xlsx - Key resource Excel file with 2 sheets:
    • Resources - table listing resources and reagents used in this study:
      • antibodies,
      • bacterial and viral strains
      • chemicals, peptides & proteins
      • critical commercial assays
      • experimental model cell lines
      • experimental model organisms
      • oligonucleotides
      • recombinant DNA
      • software and algorithms.
    • Sequenced_samples - metadata of all samples sequenced in this study, containing following informations: 
      • experiment - the experimental context for which the sample was sequenced
      • sample_alias - the revised sample identifier following review (e.g., Larp4 updated to Larp4a; Larp5 updated to Larp4b)
      • sample_title - sample ID in European Nucleotide Archive (ENA)
      • sequencing_type - sequencing protocol (either cDNA or DRS)
      • kit - sequencing chemistry used to prepare the library
      • reads - count of reads produced
      • Guppy - version of Gupy basecaller used
      • Nanopolish - version of Nanopolish polya used
      • Dorado - version of Dorado basecaller used
      • reference - version of reference sequence used
      • organism - organism the sample was derived from
      • accession - sample accession number in European Nucleotide Archive (ENA)
      • project - project accession number in European Nucleotide Archive (ENA)
      • sample_description - a description of the sample contents and the preparation methodology
  • Supplementary_Resource_1.tar.gz - TurboID assay MS/MS raw data, each file in the archive corresponds to a biological replicate of either WT (control) or Tent5c-EGFP.
  • Supplementary_Resource_2.tar.gz - Nanopolish polya predictions for FLEB E14.5 data (corresponding to Supplementary table 1). Contains .tsv files one per each sequenced sample, with naming convention corresponding to sample aliases in Supplementary table 6. Each of the .tsv file within the archive contains the following columns:
    • readname - unique read ID from the fast5 file
    • contig - contig/sequence in the reference that the read aligns to
    • position - position on the reference contig at which the read-alignment starts
    • leader_start - index of the raw (pico-amp) sample at which segmentation algorithm declares the leader starts
    • adapter_start - index of the raw (pico-amp) sample at which segmentation algorithm declares the sequencing adapter region starts
    • polya_start - index of the raw (pico-amp) sample at which segmentation algorithm declares the poly(A) region of the RNA starts
    • transcript_start - index of the raw (pico-amp) sample at which segmentation algorithm declares the transcript body region of the RNA starts
    • read_rate - number of nucleotides moving through the pore per second
    • polya_length - estimated poly(A) tail length [nt]
    • qc_tag - quality tag assigned by nanopolish polya function
  • Supplementary_Resource_3.tar.gz - Dorado polya predictions for sorted BasoE, OrthoE, PolyE and Retc data (corresponding to Supplementary table 2). Contains .tsv files one per each sequenced sample, with naming convention corresponding to sample aliases in Supplementary table 6. Each of the .tsv file within the archive contains the following columns:
    • read_id - unique read ID from the pod5 file
    • reference - contig/sequence in the reference that the read aligns to
    • ref_start - position on the reference contig at which the read-alignment starts
    • ref_end - position on the reference contig at which the read-alignment ends
    • mapq - alignment mapping quality
    • pt - estimated poly(A) tail length [nt]
    • sequence - basecalled sequence of read
  • Supplementary_Resource_4.tar.gz - Dorado polya predictions for silencing experiment data (corresponding to Supplementary table 4). Contains .tsv files one per each sequenced sample, with naming convention corresponding to sample aliases in Supplementary table 6. Each of the .tsv file within the archive contains the following columns:
    • read_id - unique read ID from the pod5 file
    • reference - contig/sequence in the reference that the read aligns to
    • ref_start - position on the reference contig at which the read-alignment starts
    • ref_end - position on the reference contig at which the read-alignment ends
    • mapq - alignment mapping quality
    • pt - estimated poly(A) tail length [nt]
    • sequence - basecalled sequence of read
  • Cytometry_raw.tar - raw data from cytometry experiments
    • archive containing subfolders corresponding to each experiment. Within each subfolder, fcs files are stored.
  • ELISA.tar - raw data from ELISA experiments
    • xlsx files with the reports from Assayfit Pro 1.41 and Magellan measurements
  • Procyte_Dx.tar - raw data from blood parameters assessment
    • Results from ProCyte Dx hematology analysis of WT vs TENT5Ccat mice in xlsx format
  • RT_qPCR.tar - raw dara from RT-qPCR experiments
    • raw and merged results in xlsx format 
  • WB.tar - raw data from western blotting

Files

Fig.1.pdf

Files (6.1 GB)

Name Size
md5:7cf3cb8aca46ca51bc45714c69ec2a5e
3.5 GB Download
md5:7bf7188330441bb4680ad001c9da5958
189.4 kB Download
md5:edebec4db957085d02bd5c671adbce0a
291.1 kB Preview Download
md5:e59ef9f017edab8ff563514b7d74aa1a
311.4 kB Preview Download
md5:8cfdecb7285827d51acca276f538235d
220.2 kB Preview Download
md5:fd37147b123e65279bac66b5d4189925
1.6 MB Preview Download
md5:89d32abc9beab173fd390c3cb4e6ab19
341.8 kB Preview Download
md5:6b7761cf8ed77c97d58350e4ee06e0a1
226.9 kB Preview Download
md5:2276e60c27901fbefaae230c5aa61582
986.4 kB Preview Download
md5:40a32caf6eedd298d62fe88a2981da48
657.3 kB Preview Download
md5:f5ba0bafbf59b1df1168c7a334418724
250.6 kB Preview Download
md5:e616f8e32bb4334b4f680f98d4402e76
47.0 kB Preview Download
md5:261472ec8650b34bc5c96149e406efff
281.1 kB Preview Download
md5:f001e0b08bfdf7cdf8d66e09efcd558e
831.7 kB Preview Download
md5:1f7012b588f0ff6b75e31b4fd3df538b
196.9 kB Preview Download
md5:b26b17906d3fe74baaee564f19bb97da
158.8 kB Preview Download
md5:810bc705eff69d288140f42ad0c5f089
6.1 MB Preview Download
md5:c73073e5ff19e6d2b2e5cd6bd321aa0c
2.7 MB Preview Download
md5:4b70efeed803600a40a2fc7ba458c3c4
7.3 MB Preview Download
md5:a0cbd0714faf44c2ccef8b37e2ee3dd6
41.0 kB Download
md5:e6078a745f4300ea7ecf190b0be86c07
153.6 kB Download
md5:4d9d614b91c4e4cda539ca69886e1bfe
1.9 GB Download
md5:4eb90bbdb4d368a0229e9b4772aa76cb
32.4 MB Download
md5:bd5a69fc84cc14df16cd59053357f2f4
78.4 MB Download
md5:8311d3d8b44135be0c3498143c7ccd05
27.2 MB Download
md5:8e5d33481ff7fb2f51ad4adbc3cefceb
1.6 MB Preview Download
md5:705b3e5b1413e1ae4714d237ca21ff2b
2.2 MB Download
md5:604200e67fde46e2a9cce247bda006f8
2.2 MB Download
md5:8e95e957b2a9947475e82b1a16dd5bfc
1.2 MB Download
md5:650c93220c82a66466eb35adfdc117e1
2.7 MB Download
md5:79ab431d539c572635ca293d549f62d7
65.6 kB Download
md5:3b4b909dc8f40de270f476e4655a692a
58.3 kB Download
md5:f721410d91758326cfff06742171da25
109.8 kB Download
md5:b2229b8bc21fba74ecae55ce27991e7d
515.0 MB Download

Additional details

Related works

Is supplement to
Preprint: 10.1101/2024.11.14.623596 (DOI)