Inferring spatial transcriptomics markers from whole slide images to characterize metastasis-related spatial heterogeneity of colorectal tumors: A pilot study

Michael Fatemi; Eric Feng; Cyril Sharma; Zarif Azher; Tarushii Goel; Ojas Ramwala; Scott M Palisoul; Rachael E Barney; Laurent Perreard; Fred W Kolling; Lucas A Salas; Brock C Christensen; Gregory J Tsongalis; Louis J Vaickus; Joshua J Levy

doi:10.1016/j.jpi.2023.100308

Inferring spatial transcriptomics markers from whole slide images to characterize metastasis-related spatial heterogeneity of colorectal tumors: A pilot study

J Pathol Inform. 2023 Mar 29:14:100308. doi: 10.1016/j.jpi.2023.100308. eCollection 2023.

Authors

Michael Fatemi¹, Eric Feng², Cyril Sharma³, Zarif Azher², Tarushii Goel⁴, Ojas Ramwala⁵, Scott M Palisoul⁶, Rachael E Barney⁶, Laurent Perreard⁷, Fred W Kolling⁷, Lucas A Salas^{8

9

10}, Brock C Christensen^{8

9

11}, Gregory J Tsongalis⁶, Louis J Vaickus⁶, Joshua J Levy^{6

8

12

13}

Affiliations

¹ Department of Computer Science, University of Virginia, Charlottesville, VA, USA.
² Thomas Jefferson High School for Science and Technology, Alexandria, VA, USA.
³ Department of Computer Science, Purdue University, West Lafayette, IN, USA.
⁴ Department of Computer Science, Massachusetts Institute of Technology, Cambridge, MA, USA.
⁵ Department of Biomedical Informatics and Medical Education, University of Washington, Seattle, WA, USA.
⁶ Emerging Diagnostic and Investigative Technologies, Department of Pathology and Laboratory Medicine, Dartmouth Health, Lebanon, NH, USA.
⁷ Dartmouth Cancer Center, Lebanon, NH, USA.
⁸ Department of Epidemiology, Dartmouth College Geisel School of Medicine, Hanover, NH, USA.
⁹ Department of Molecular and Systems Biology, Dartmouth College Geisel School of Medicine, Hanover, NH, USA.
¹⁰ Integrative Neuroscience at Dartmouth (IND) graduate program, Dartmouth College Geisel School of Medicine, Hanover, NH, USA.
¹¹ Department of Community and Family Medicine, Dartmouth College Geisel School of Medicine, Hanover, NH, USA.
¹² Department of Dermatology, Dartmouth Health, Lebanon, NH, USA.
¹³ Program in Quantitative Biomedical Sciences, Dartmouth College Geisel School of Medicine, Hanover, NH, USA.

Abstract

Over 150 000 Americans are diagnosed with colorectal cancer (CRC) every year, and annually over 50 000 individuals will die from CRC, necessitating improvements in screening, prognostication, disease management, and therapeutic options. Tumor metastasis is the primary factor related to the risk of recurrence and mortality. Yet, screening for nodal and distant metastasis is costly, and invasive and incomplete resection may hamper adequate assessment. Signatures of the tumor-immune microenvironment (TIME) at the primary site can provide valuable insights into the aggressiveness of the tumor and the effectiveness of various treatment options. Spatially resolved transcriptomics technologies offer an unprecedented characterization of TIME through high multiplexing, yet their scope is constrained by cost. Meanwhile, it has long been suspected that histological, cytological, and macroarchitectural tissue characteristics correlate well with molecular information (e.g., gene expression). Thus, a method for predicting transcriptomics data through inference of RNA patterns from whole slide images (WSI) is a key step in studying metastasis at scale. In this work, we collected tissue from 4 stage-III (pT3) matched colorectal cancer patients for spatial transcriptomics profiling. The Visium spatial transcriptomics (ST) assay was used to measure transcript abundance for 17 943 genes at up to 5000 55-micron (i.e., 1-10 cells) spots per patient sampled in a honeycomb pattern, co-registered with hematoxylin and eosin (H&E) stained WSI. The Visium ST assay can measure expression at these spots through tissue permeabilization of mRNAs, which are captured through spatially (i.e., x-y positional coordinates) barcoded, gene specific oligo probes. WSI subimages were extracted around each co-registered Visium spot and were used to predict the expression at these spots using machine learning models. We prototyped and compared several convolutional, transformer, and graph convolutional neural networks to predict spatial RNA patterns at the Visium spots under the hypothesis that the transformer- and graph-based approaches better capture relevant spatial tissue architecture. We further analyzed the model's ability to recapitulate spatial autocorrelation statistics using SPARK and SpatialDE. Overall, the results indicate that the transformer- and graph-based approaches were unable to outperform the convolutional neural network architecture, though they exhibited optimal performance for relevant disease-associated genes. Initial findings suggest that different neural networks that operate on different scales are relevant for capturing distinct disease pathways (e.g., epithelial to mesenchymal transition). We add further evidence that deep learning models can accurately predict gene expression in whole slide images and comment on understudied factors which may increase its external applicability (e.g., tissue context). Our preliminary work will motivate further investigation of inference for molecular patterns from whole slide images as metastasis predictors and in other applications.

Keywords: Colorectal cancer; Deep learning; Graph neural network; Histomorphology; Spatial transcriptomics; Transformers.

Abstract

Grants and funding