Hudson: A computational pipeline for spatial analysis of multiplexed images
Authors/Creators
Description
Processing and analysis of multiplexed whole tissue section images at sub-cellular resolution remains a challenge in the spatial multi-omics field. Previously, we developed PySeq2500 which repurposed Illumina HiSeq 2500 Sequencing Systems as a spatial assay-agnostic instrument with 4 color fluorescence imaging, independent dual flowcell fluidics, and temperature control. To easily and reproducibly analyze images from multi-omic spatial experiments generated from PySeq2500s, we created a modular, open-source, computational pipeline, called Hudson, to preprocess raw images, segment cells, build a graph network representation of tissue sections, measure features, identify cell types and neighborhoods, and perform spatial analyses and statistics using the Snakemake workflow management system. Modular steps in the Hudson pipeline are configurable with a human readable and writable YAML configuration file. The initial image preprocessing steps are specific to PySeq2500 instruments and include illumination correction, stitching, and channel/cycle registration. Additionally, spectral overlap of fluorescent signals are removed using PICASSOnn. Alternative preprocessing steps specific to different image acquisition systems could be substituted in or omitted entirely. Polished whole tissue section images are then segmented into single cells and a graph network representation of the tissue section is computed from the Delaunay triangulation of cell centroids. Then features such as morphological characteristics, marker mean intensities, or marker counts are measured for each cell. Features are clustered to identify cell types from reference single cell data and cell neighborhoods are found by local recurrent cell compositions. Lastly spatial statistics and patterns are computed and identified across markers, cells, and neighborhoods. All computations in Hudson are lazily performed for memory efficiency and data objects are saved in compressed chunked formats to scale the pipeline from a workstation to a cluster or cloud environment. Interactive reports are also generated for each step to explore the data in depth. Hudson enables users to process whole tissue multi-omic data starting from raw images all the way to conducting spatial analyses in an automated, configurable, and reproducible manner.
Files
KRP_AGBT_2023.pdf
Files
(61.4 MB)
| Name | Size | Download all |
|---|---|---|
|
md5:7acb67e9a7a90129efb8357a80c25cdb
|
61.4 MB | Preview Download |
Additional details
Software
- Repository URL
- https://github.com/nygctech/hudson
- Programming language
- Python , Snakemake
- Development Status
- Active