scRNA-seq and CITE-seq

PBMC scRNA-seq

scRNA-seq data from peripheral blood mononuclear cells (PBMCs) was generated on the 10x Genomics 3' scRNA-seq Platform (v3.1). For data collection and processing details, see the Cohorts and Experimental Methods sections.

Below, we provide labeled and annotated PBMC scRNA-seq data from our NDMM cohort and healthy controls from the Sound Life cohort. More information about the Sound Life cohort is available in the Dynamics of Immune Health and Age website.

All .h5ad files for this project contain sample and subject metadata, in addition to cell type labels and QC metrics. Click the header below for descriptions of these metadata:

Each file contains sample-level metadata, as well as cell-level cell type labels and QC metrics. The following values are stored in the .obs section of these .h5ad files as descriptions of observations:

Sample Identifiers
cohort.cohortGuid: A Globally Unique Identifier (GUID) of the Cohort the subject enrolled in for our study subject.subjectGuid: A GUID for the Subject
subject.subjectID: The Subject ID used in the figures and text of our article
sample.sampleKitGuid: A GUID for the Sample Kit, representing all material collected at a visit
specimen.specimenGuid: A GUID for the specific aliquot used for the experiment

Subject Metadata
subject.biologicalSex: The biological sex of the Subject
subject.birthYear: The Birth Year of the Subject
subject.ageAtFirstDraw:The Age of the Subject at their first on-study sample collection
subject.race:The self-reported Race of the Subject
subject.ethnicity:The self-reported Ethnicity of the subject
subject.cmv:The CMV Status of the subject, as determined by an HCMV assay (Negative or Positive)
subject.treatmentResponse: The treatment response recorded at End of Induction therapy
subject.fluVaccineResponse: The flu vaccine response category assigned to the subject based on post-ASCT flu vaccine timepoints (Responder or Non-Responder)

Sample Metadata
sample.visitName: The name of the study visit (i.e. time point)
sample.visitLabel: The abbreviated label for the study visit used in manuscript figures
sample.drawYear: The year of the study visit (e.g. 2021)
sample.subjectAgeAtDraw: The age of the Subject in years at the time of sample collection
sample.daysSinceFirstVisit: Number of days since the subject's initial visit

Process Identifiers
batch_id: A GUID for the batch of samples processed together (e.g. B039)
pool_id: A GUID for the pool of samples combined for Cell Hashing (e.g. B039-P1)
chip_id: A GUID for the 10x Genomics chip the cells were loaded into (e.g. B039-P1C2)
well_id: A GUID for the 10x Genomics well the cells were loaded into within the chip (e.g. B039-P1C2W4)
*barcodes: A GUID for the individual cell
original_barcodes: The original, sequence-based barcode generated by 10x Genomics Cell Ranger software
cell_name: A quasi-unique, memorable cell identifier generated using an adjective-adjective-animal structure

*used as the primary cell index in our .h5ad files

Cell QC Metrics
n_reads: Number of reads assigned to the cell barcode
n_umis: Number of Unique Molecular Identifiers (unique molecules) detected
n_genes: Number of genes with at least 1 UMI detected
total_counts_mito: Total number of reads that were assigned to mitochondrial genes
pct_counts_mito: Percent of reads that were assigned to mitochondrial genes
doublet_score: Doublet score assigned by Scrublet for doublet detection

Cell Labeling Results
AIFI_L1: Final broad class cell type label (9 types)
AIFI_L2: Final mid resolution cell type label (29 types)
AIFI_L3: Final high resolution cell type label (71 types)
AIFI_L2_label: Abbreviated mid resolution cell type label used in figures
AIFI_L3_label: Abbreviated high resolution cell type label used in figures

Visit Group .h5ad files

We are providing our scRNA-seq data in AnnData (.h5ad) format. For more details about AnnData, see the AnnData Documentation Page.

To reduce download size, normalized data are not provided. Normalization and log transformation can be performed using the scanpy package with:

scanpy.pp.normalize_total(adata, target_sum = 1e4)
scanpy.pp.log1p(adata)

Each file provided below contains a subset of the full > 5.6 million cell dataset. Sample counts, cell counts, and approximate file sizes are below:

File Name N Subjects N Samples N Cells File Size
MM_VRd_Treatment_Visits.h5ad17711,105,5879.8 GB
MM_VRd_Flu_Vaccine_Visits.h5ad15811,194,80911 GB
MM_VRd_COVID_Vaccine_Visits.h5ad2573,7210.8 GB
Healthy_Flu_Vaccine_Year1_Visits.h5ad331021,695,86027 GB
Healthy_Flu_Vaccine_Year2_Visits.h5ad31941,407,39323 GB
NDMM PBMC scRNA-seq .h5ad files
File NameDescriptionDownload Link
Healthy_Flu_Vaccine_Year1_Visits.h5ad Healthy control vaccine visit PBMC scRNA-seq
Healthy_Flu_Vaccine_Year2_Visits.h5ad Healthy control vaccine visit PBMC scRNA-seq
NDMM_VRd_COVID_Vaccine_Visits.h5ad NDMM COVID vaccine visit PBMC scRNA-seq
NDMM_VRd_Flu_Vaccine_Visits.h5ad NDMM Flu vaccine visit PBMC scRNA-seq
NDMM_VRd_Treatment_Visits.h5ad NDMM treatment visit PBMC scRNA-seq

BMMC CITE-seq

CITE-seq data, which includes both transcriptional and cell-surface epitope measurement, from bone marrow mononuclear cells (BMMCs) was generated on the 10x Genomics 3' scRNA-seq platform (v3.1).

BMMCs were stained with a custom panel of 59 oligo-conjugated antibodies (Antibody-Derived Tags; ADTs). Click the header below to view a list of the antibodies utilized for ADT staining:

All antibodies were obtained from BioLegend as TotalSeq anti-human ADTs. Catalog number and TotalSeq ID columns refer to BioLegend accessions.

TargetCloneTiterTotalSeq IDCatalog NumberRRID
CD1aHI1491.0A0402300133AB_2783146
CD1cL1610.2A0160331539AB_2734326
CD2TS1/80.02A0367309229AB_2783172
CD3UCHT10.075A0034300475AB_2734246
CD4RPA-T40.1A0072300563AB_2734247
CD7CD7-6B70.1A0066343123AB_2734345
CD8SK10.02A0046344751AB_2734351
CD10HI10a0.75A0062312231AB_2734286
CD11bICRF440.05A0161301353AB_2734249
CD11cS-HCL-30.1A0053371519AB_2749971
CD13WM150.2A0364301729AB_2783151
CD14M5E20.1A0081301855AB_2734254
CD163G80.1A0083302061AB_2734255
CD19HIB190.1A0050302259AB_2734256
CD202H70.05A0100302359AB_2749984
CD22S-HCL-10.2A0393363514AB_2734404
CD24ML50.5A0180311137AB_2750374
CD25BC960.08A0085302643AB_2734258
CD27O3230.1A0154302847AB_2750000
CD33P67.60.2A0052366629AB_2734409
CD345810.4A0054343537AB_2749972
CD365-2710.02A0407336225AB_2800892
CD38HB-70.01A0410356635AB_2800967
CD39A10.1A0176328233AB_2750005
CD405C30.5A0031334346AB_2749968
CD41HIP80.1A0353303737AB_2783162
CD45RAHI1000.125A0063304157AB_2734267
CD47CC2C60.2A0026323129AB_2734305
CD56 (NCAM)5.1H110.1A0047362557AB_2749970
CD6410.10.1A0162305037AB_2750366
CD66b6/40c0.25A0166392905AB_2750372
CD69FN500.75A0146310947AB_2749997
CD71CY1G40.1A0394334123AB_2800884
CD802D101.0A0005305239AB_2749958
CD84CD84.1.210.75A0872326011AB_2814189
CD86IT2.20.05A0006305443AB_2734273
CD88S5/11.0A1046344321AB_2888875
CD94DX220.2A0867305521AB_2814142
CD117 (c-kit)104D20.75A0061313241AB_2734287
CD1236H60.1A0064306037AB_2749977
CD127 (IL-7Rα)A019D50.125A0390351352AB_2734366
CD141 (Thrombomodulin)M800.75A0163344121AB_2783229
CD163GHI/610.75A0358333635AB_2750343
CD172a (SIRPα)15-4140.5A0408372109AB_2783285
CD192 (CCR2)K036C21.0A0242357229AB_2750501
CD195 (CCR5)J418F10.5A0141359135AB_2749994
CD206 (MMR)15-21.0A0205321143AB_2750010
CD226 (DNAM-1)GHI/610.75A0805333635AB_2750343
CD269 (BCMA)19F22.0A0056357521AB_2749974
CD274 (B7-H1, PD-L1)29E.2A31.0A0007329743AB_2749959
CD279 (PD-1)EH12.2H72.0A0088329955AB_2734322
CD304 (Neuropilin-1)12C20.5A0406354525AB_2783261
CD314 (NKG2D)1D110.75A0165320835AB_2734298
CD319 (CRACC)162.11.0A0830331821AB_2800872
CX3CR1K0124E11.0A0179355709AB_2832698
HLA-DRL2430.2A0159307659AB_2750001
IgDIA6-20.05A0384348243AB_2783238
IgG FcM1310G051.0A0375410725AB_2783329
IgMMHM-880.1A0136314541AB_2749992
Subject Group .h5ad files

We are providing our scRNA-seq data in AnnData (.h5ad) format. For more details about AnnData, see the AnnData Documentation Page.

To reduce download size, normalized data are not provided. Normalization and log transformation can be performed using the scanpy package with:

scanpy.pp.normalize_total(adata, target_sum = 1e4)
scanpy.pp.log1p(adata)

Antibody-dervided tags (ADT) for a custom panel of 59 targets are provided as a table of counts of ADT-derived UMIs for each antibody per cell. These counts are stored in the adata.obsm['adt_counts'] slot.

As for PBMC data, the .h5ad files for this project contain sample and subject metadata. Click the header below for descriptions of these metadata fields:

Each file contains sample-level metadata, as well as cell-level cell type labels and QC metrics. The following values are stored in the .obs section of these .h5ad files as descriptions of observations:

Sample Identifiers
cohort.cohortGuid: A Globally Unique Identifier (GUID) of the Cohort the subject enrolled in for our study subject.subjectGuid: A GUID for the Subject
subject.subjectID: The Subject ID used in the figures and text of our article
sample.sampleKitGuid: A GUID for the Sample Kit, representing all material collected at a visit
specimen.specimenGuid: A GUID for the specific aliquot used for the experiment

Subject Metadata
subject.biologicalSex: The biological sex of the Subject
subject.birthYear: The Birth Year of the Subject
subject.ageAtFirstDraw:The Age of the Subject at their first on-study sample collection
subject.race:The self-reported Race of the Subject
subject.ethnicity:The self-reported Ethnicity of the subject
subject.cmv:The CMV Status of the subject, as determined by an HCMV assay (Negative or Positive)
subject.treatmentResponse: The treatment response recorded at End of Induction therapy
subject.fluVaccineResponse: The flu vaccine response category assigned to the subject based on post-ASCT flu vaccine timepoints (Responder or Non-Responder)

Sample Metadata
sample.visitName: The name of the study visit (i.e. time point)
sample.visitLabel: The abbreviated label for the study visit used in manuscript figures
sample.drawYear: The year of the study visit (e.g. 2021)
sample.subjectAgeAtDraw: The age of the Subject in years at the time of sample collection
sample.daysSinceFirstVisit: Number of days since the subject's initial visit

Process Identifiers
batch_id: A GUID for the batch of samples processed together (e.g. B039)
pool_id: A GUID for the pool of samples combined for Cell Hashing (e.g. B039-P1)
chip_id: A GUID for the 10x Genomics chip the cells were loaded into (e.g. B039-P1C2)
well_id: A GUID for the 10x Genomics well the cells were loaded into within the chip (e.g. B039-P1C2W4)
*barcodes: A GUID for the individual cell
original_barcodes: The original, sequence-based barcode generated by 10x Genomics Cell Ranger software
cell_name: A quasi-unique, memorable cell identifier generated using an adjective-adjective-animal structure

*used as the primary cell index in our .h5ad files

Cell QC Metrics
n_reads: Number of reads assigned to the cell barcode
n_umis: Number of Unique Molecular Identifiers (unique molecules) detected
n_genes: Number of genes with at least 1 UMI detected
total_counts_mito: Total number of reads that were assigned to mitochondrial genes
pct_counts_mito: Percent of reads that were assigned to mitochondrial genes
doublet_score: Doublet score assigned by Scrublet for doublet detection

Cell Labeling Results
BMMC_L1: Final broad class cell type label (7 types)
BMMC_L2: Final mid resolution cell type label (39 types)
BMMC_L3: Final high resolution cell type label (64 types)
BMMC_L2_label: Abbreviated mid resolution cell type label used in figures
BMMC_L3_label: Abbreviated high resolution cell type label used in figures

The .h5ad files provided below are separated into two groups, from NDMM patients or from the 10 healthy control subjects for BMMCs:

File NameN SubjectsN SamplesN CellsFile Size
MM_VRd_BMMC_Visits.h5ad17591,102,66930 GB
Healthy_BMMC_Visits.h5ad101069,2961.2 GB
NDMM BMMC CITE-seq .h5ad files
File NameDescriptionDownload Link
Healthy_BMMC_Visits.h5ad Bone Marrow CITE-seq from Healthy controls
NDMM_VRd_BMMC_Visits.h5ad Bone Marrow CITE-seq from MM patients