Published February 6, 2023 | Version v1

Hudson: A computational pipeline for spatial analysis of multiplexed images

Description

Processing and analysis of multiplexed whole tissue section images at sub-cellular resolution remains a challenge in the spatial multi-omics field. Previously, we developed PySeq2500 which repurposed Illumina HiSeq 2500 Sequencing Systems as a spatial assay-agnostic instrument with 4 color fluorescence imaging, independent dual flowcell fluidics, and temperature control. To easily and reproducibly analyze images from multi-omic spatial experiments generated from PySeq2500s, we created a modular, open-source, computational pipeline, called Hudson, to preprocess raw images, segment cells, build a graph network representation of tissue sections, measure features, identify cell types and neighborhoods, and perform spatial analyses and statistics using the Snakemake workflow management system. Modular steps in the Hudson pipeline are configurable with a human readable and writable YAML configuration file. The initial image preprocessing steps are specific to PySeq2500 instruments and include illumination correction, stitching, and channel/cycle registration. Additionally, spectral overlap of fluorescent signals are removed using PICASSOnn. Alternative preprocessing steps specific to different image acquisition systems could be substituted in or omitted entirely. Polished whole tissue section images are then segmented into single cells and a graph network representation of the tissue section is computed from the Delaunay triangulation of cell centroids. Then features such as morphological characteristics, marker mean intensities, or marker counts are measured for each cell. Features are clustered to identify cell types from reference single cell data and cell neighborhoods are found by local recurrent cell compositions. Lastly spatial statistics and patterns are computed and identified across markers, cells, and neighborhoods. All computations in Hudson are lazily performed for memory efficiency and data objects are saved in compressed chunked formats to scale the pipeline from a workstation to a cluster or cloud environment. Interactive reports are also generated for each step to explore the data in depth. Hudson enables users to process whole tissue multi-omic data starting from raw images all the way to conducting spatial analyses in an automated, configurable, and reproducible manner.

Files

KRP_AGBT_2023.pdf

Files (61.4 MB)

Name Size Download all
md5:7acb67e9a7a90129efb8357a80c25cdb
61.4 MB Preview Download

Additional details

Software

Repository URL
https://github.com/nygctech/hudson
Programming language
Python , Snakemake
Development Status
Active