Authors: Lucas T. Graybuck and Nicholas Moss
Date: 2026-05-20
Overview
This page provides example Python code for implementing the non-negative least scores (NNLS) method for IFN-α and IFN-γ response scoring based on the Human IFN Response Across Immune Cell Subsets Atlas (HIRISA) resource from the Allen Institute.
This page provides the code that underlies our interactive scoring tool, which may be helpful for exploring these methods:
https://apps.allenimmunology.org/aifi/resources/ifn-response/tools/ifn-score/
A Jupyter Notebook version of this page can be viewed online here:
View notebook: EX-Python_scoring_file_set.ipynb
All code related to this study is available on Github:
aifimmunology/IFN_manuscript
The manuscript cited below describes the HIRISA resource and its generation. If you utilize this data and/or method, please cite:
Moss N, Sakai C, Kaul SN, Graybuck LT, Rachid Zaim S, Angus-Hill ML, et al. Dissecting type I and II interferon impacts on human immune cells in disease by a cell type-specific interferon response atlas. bioRxiv; 2025. doi:10.64898/2025.12.02.691676
Requirements
To utilize this code, you'll need the input data .zip file provided on our Downloads page. You will not need to unzip this file.
To score your own data, you'll need to have human differential gene expression results that include both a Gene column (containing HGNC gene symbols) and a Log2FC column (containing log2-transformed fold change).
The following versions of key packages were used to run this example:
Package Version
numpy 2.3.3
pandas 2.3.2
plotly 6.3.0
python 3.12.11
scipy 1.16.1
Import libraries
from scipy.optimize import nnls # Non-negative least squares computation
import numpy as np # General numeric functions and classes
import os # General file handling
import pandas as pd # General data frame handling
import plotly.express as px # Interactive plotting
import re # Regular expressions
import session_info # Display session information/package versions
import zipfile # .zip file handling
Helper functions
Below, we provide custom helper functions to easily read data, perform scoring, and plot results. Click the section below to view these functions.
Helper functions to read data frames from the .csv files stored in our .zip file
def read_tables_by_suffix(in_zip: str, suffix: str):
df_dict = {}
with zipfile.ZipFile(in_zip, 'r') as z:
file_list = z.namelist()
for f in file_list:
if suffix in f:
cell_type = re.sub('_.+','',os.path.basename(f))
with z.open(f) as zf:
df_dict[cell_type] = pd.read_csv(zf)
if len(df_dict) > 1:
return(df_dict)
else:
df = list(df_dict.values())[0]
return(df)
def read_reference_matrices(in_zip):
mat_dict = read_tables_by_suffix(in_zip, "_matrix.csv")
return(mat_dict)
def read_test_degs(in_zip):
deg_dict = read_tables_by_suffix(in_zip, "_degs.csv")
return(deg_dict)
def read_disease_scores(in_zip):
score_df = read_tables_by_suffix(in_zip, "_scores.csv")
return(score_df)
Helper functions for calculation of IFN scores
def scale(y):
x = y.copy()
x /= np.sqrt(x.pow(2).sum().div(x.count() - 1))
return x
def calculate_ifn_score(
test_df,
gene_column: str,
lfc_column: str,
score_mats,
cell_type,
result_name = None
):
test_df = test_df[[gene_column, lfc_column]]
test_df.columns = ['Gene', 'log2FC']
stim_mat = score_mats[cell_type]
final_mat = pd.merge(stim_mat, test_df, on='Gene', how='left').fillna(0)
final_mat_scaled = scale(final_mat.drop(columns='Gene'))
A = final_mat_scaled[['IFNa', 'IFNg']].values
b = final_mat_scaled['log2FC'].values
result, _ = nnls(A, b)
# Populations that don't have an IFNg receptor
# get IFNg scores set to zero.
ifng_nonresponding = ["gdT", "MAIT", "Memory CD8", "NK", "NK CD56dim", "NK CD56hi", "Plasma"]
if cell_type in ifng_nonresponding:
result[1] = 0
out_df = pd.DataFrame({
'Name': result_name,
'Cell Type': [cell_type],
'IFNa Score': [result[0]],
'IFNg Score': [result[1]],
})
return out_df
Helper function to plot IFN scores relative to reference cohorts
def plot_ifn_scores(
score_df,
disease_df,
cell_type
):
disease_df = disease_df[disease_df['celltype'] == cell_type]
fig = px.scatter(
disease_df,
x = 'avg_IFNa',
y = 'avg_IFNg',
error_x = 'var_IFNa',
error_y = 'var_IFNg',
size = 'var_combined',
color = 'Cohort',
labels = {
'avg_IFNa': 'IFN-α Score',
'avg_IFNg': 'IFN-γ Score'
},
render_mode = 'webgl'
)
score_traces = px.scatter(
score_df,
x = 'IFNa Score',
y = 'IFNg Score',
render_mode = 'webgl'
).update_traces(
marker_size = 12,
marker_symbol = 'star',
marker_color = 'red'
).data
fig.add_traces(score_traces)
fig.update_layout(
height = 500,
width = 700
)
return fig
Read reference data
In this section, we'll read in the reference IFN score matrices, precomputed disease score references, and example tables of DEG results to demonstrate how to use these functions.
If you use your own data, you will need the first two components (IFN matrices and disease score references), but could skip example DEG tables.
# Path to your copy of the hirisa scoring reference data
in_zip = 'hirisa_scoring_data_2026-05-20.zip'
reference_matrices = read_reference_matrices(in_zip)
disease_scores = read_disease_scores(in_zip)
example_degs = read_test_degs(in_zip)
Score data against reference matrices
To score data, we'll need to specify which cell population matches our input data. Scoring matrices were calculated based on specific subpopulations of PBMC cells, or using a whole-PBMC dataset. The following population names can be used to select relevant cell populations for analysis:
Bulk cells
PBMC
Broad classes (L1)
Bcell
Monocyte
NK
Tcell
Finer populations (L2)
B Naive
B Memory
CD14 Monocyte
CD4 Memory
CD4 Naive
CD8 Naive
CD8 Memory
gdT
MAIT
NK CD56dim
NK CD56hi
Plasma
Treg
As an example for scoring, we'll use one of our example data frames stored in the test_degs dictionary. This dictionary contains two-column data frames with Gene and log2FC columns for "Bcell", "Monocyte", "NK", and "Tcell" populations.
example_df = example_degs['Tcell']
example_df.head()
| Gene | log2FC |
| FOS | 1.127489 |
| MT-ND2 | 0.703931 |
| TNFAIP3 | 1.16574 |
| CXCR4 | 0.79754 |
| DUSP1 | 0.712384 |
To perform scoring with NNLS, we need to provide the deg table, specify which columns contain gene symbols (gene_column) and Log2(Fold Change) results (lfc_column), as well as the reference matrices dictionary (score_mats, loaded above), the cell population to use (cell_type), and an optional name for the result (result_name).
example_score = calculate_ifn_score(
example_df,
gene_column = 'Gene',
lfc_column = 'log2FC',
score_mats = reference_matrices,
cell_type = 'Tcell'
)
The result will be a Pandas DataFrame with Name, Cell Type, IFNa Score, and IFNg score columns:
| Name | Cell Type | IFNa Score | IFNg Score |
| Tcell | 0.327087 | 0.043044 |
Visualize scores relative to disease references
The IFNa and IFNg score results are on an arbitrary scale that can make them hard to interpret alone. To provide context for these scores, we precomputed IFNa and IFNg scores from each immune cell population from multiple cohorts of subjects with immune-related diseases compared to healthy controls: Aging, lupus (SLE), rheumatoid arthritis (RA), COVID-19 (Acute and Long Covid), flu, malaria, multiple myeloma (MM), and autoimmune lymphoproliferative syndrome (ALPS).
For each disease, we plot the average scores (center of points), variance of scores (shown as error bars), and the combined variance of IFNa and IFNg scores (size of points) computed from the subject-specific scores in each cohort.
Our computed results can be overlaid on this context, and is shown as a red star.
To do so, we need to provide our score result data frame, the reference disease scores, and the cell population:
plot_ifn_scores(
example_score,
disease_scores,
'Tcell'
)Session Info
This function shows the version of packages currently loaded in this notebook session: