Skip to contents

TAPE uses an autoencoder whose latent layer encodes cell-type fractions and whose decoder weights collapse into a signature matrix. Its distinguishing feature is an adaptive stage that refines the trained model directly on the target bulk data, which lets it adjust to a tissue or condition the reference does not perfectly match.

Chunks below are not evaluated during build. Epochs are reduced for demonstration; use the function defaults for real analyses.

Data

library(quasar)
library(QuasarDeconData)

sc_ref <- load_cov_imm("train")
pbulk  <- load_cov_pbulk(1)

bulk  <- pbulk$bulk_expression_profiles   # genes x samples
truth <- pbulk$ground_truth_proportions

Simulate

tape_sim <- tape_simulate(
  sc_data      = sc_ref,
  celltype_col = "cell_type",
  samplenum    = 5000,
  n            = 500,
  sparse       = TRUE,
  random_state = 42
)

Process

tape_proc <- tape_process(
  simudata           = tape_sim,
  real_bulk          = bulk,          # genes x samples
  variance_threshold = 0.98,
  scaler             = "mms"
)

scaler = "mms" is per-sample min-max; "ss" is per-sample standardisation with a population standard deviation, matching sklearn rather than base R’s sd().

Train

tape_model <- tape_train(
  train_x    = tape_proc$train_x,
  train_y    = tape_proc$train_y,
  batch_size = 128,
  epochs     = 128,
  seed       = 0
)

The loss combines an L1 term on the latent proportions against the known fractions with an L1 reconstruction term, so the decoder learns a signature matrix at the same time as the encoder learns to deconvolve.

Predict

tape_res <- tape_predict(
  model      = tape_model,
  test_x     = tape_proc$test_x,
  genename   = tape_proc$genename,
  celltypes  = tape_proc$celltypes,
  samplename = tape_proc$samplename,
  adaptive   = TRUE,
  mode       = "overall"
)

head(tape_res$pred)
head(tape_res$sigm)

Adaptive modes

The two modes differ substantially in cost and in what they return.

"overall" adapts one model jointly across all target samples and returns a single signature matrix — one genes × cell-types profile for the whole dataset.

"high-resolution" adapts every sample with its own copy of the model, returning a per-sample signature matrix for each cell type. Far more informative, and far more expensive. chunk_size controls how many samples are adapted concurrently as an ensemble of independent models; raise it to use more GPU memory and less wall time.

tape_hr <- tape_predict(
  model      = tape_model,
  test_x     = tape_proc$test_x,
  genename   = tape_proc$genename,
  celltypes  = tape_proc$celltypes,
  samplename = tape_proc$samplename,
  adaptive   = TRUE,
  mode       = "high-resolution",
  chunk_size = 16L
)

names(tape_hr$sigm)
dim(tape_hr$sigm[[1]])

Setting adaptive = FALSE skips refinement entirely and returns proportions only, with sigm as NULL. This is much faster and worth using when the reference already matches the target well.

Wrapper

tape_res <- tape(
  sc_data      = sc_ref,
  real_bulk    = bulk,
  celltype_col = "cell_type",
  samplenum    = 5000,
  n            = 500,
  epochs       = 128,
  adaptive     = TRUE,
  mode         = "overall"
)

Evaluate

quasar_prop_metrics(tape_res$pred, truth)$cell_type_rmse

Reference

Chen, Y., Wang, Y., Chen, Y., Cheng, Y., Wei, Y., Li, Y., Wang, J., Wei, Y., Chan, T.-F., & Li, Y. (2022). Deep autoencoder for interpretable tissue-adaptive deconvolution and cell-type-specific gene analysis. Nature Communications, 13(1), 6735.

If you use this implementation, cite both the QUASAR paper and Chen et al.