Complete

Psoriatic Arthritis
Biomarker Discovery

Use Case · Biomedical Research  ·  Started April 2026  ·  Completed May 2026

An exhaustive distributed search across ~426 million gene expression pairs to identify 2-gene classifiers that distinguish psoriatic arthritis (PsA) from psoriasis (Ps), because earlier diagnosis means earlier treatment, before irreversible joint damage occurs.

426M
total gene pairs to evaluate
100%
of search space covered to date
0.927
best AUC found · 2-gene LR classifier
71.2%
of top pairs contain KLF5

Search Progress

This search has actively run and completed on the DCP network. What would have taken a single computer 25 years was completed in 3 weeks. Workers evaluated gene pairs in real time and returned results to a central leaderboard. The numbers below reflect the final results.

Search space coverage 100 % complete
426M total pairs  ·  6 Machine Learning classifiers  ·  10-fold CV per pair  ·  leaderboard: top 1,000 retained
Help accelerate future research like this. If you have idle compute and want to run a DCP Worker, connecting to the network directly contributes to meaningful workloads, like surfacing novel PsA biomarkers to inform future treatments. Open dcp.work to launch an in-browser DCP worker that joins the network.

Background

Psoriasis (Ps) is a chronic inflammatory skin condition affecting approximately 2–3% of the global population. In up to 30% of cases, patients progress to psoriatic arthritis (PsA), a systemic disease involving joint inflammation and destruction that significantly reduces quality of life and is associated with cardiovascular and metabolic comorbidities.

Distinguishing PsA from Ps at the molecular level is clinically important: earlier identification of patients at high risk of joint progression could enable preventive intervention before irreversible joint damage occurs. The two conditions share overlapping clinical features, making classification challenging by conventional means.

Gene expression microarray data offers a window into the transcriptomic differences between these phenotypes. The GSE57383 dataset (NCBI Gene Expression Omnibus) contains expression profiles for 66 patients, 30 Ps and 36 PsA, across 29,201 probesets from the Affymetrix Human Genome U133 Plus 2.0 array.

Computational Approach

Exhaustive testing of all 2-gene combinations from 29,201 probesets produces approximately 426 million candidate classifiers. Each candidate requires 10-fold cross-validation. Training and testing a classifier 10 times on the 66-patient cohort, making this problem computationally intractable on a single machine.

Pipeline

Input Data
29,201 probesets · 66 patients · GSE57383
DCP Search
MultiRangeObject · classifier × gene_i × gene_j
Worker Fleet
10-fold CV · up to 5,000 pairs per slice
Leaderboard
Top 1,000 pairs retained
Analysis
Gene frequency · heatmap · federation

Classifiers evaluated

Six scikit-learn classifiers were evaluated across the full pair space: Logistic Regression (LR), L1-regularized Logistic Regression (L1), Bagging Classifier (BA), Random Forest (RF), Support Vector Machine (SVM), and Decision Tree (DT). Workers received a (classifier, gene_i, gene_j_batch) tuple via MultiRangeObject and returned only their top-10 pairs by AUC. All six classifiers ran sequentially to completion.

LR and L1 are particularly suited for the eventual federated learning phase, as their coefficients can be averaged across institutions without sharing patient-level data.

Job architecture

The input set is expressed as a MultiRangeObject over [classifier] × [gene_i] × [gene_j_batch]. Each slice evaluates one classifier against one anchor gene paired with a batch of 5,000 partner genes. A j > i filter ensures each pair is evaluated exactly once, covering the upper triangle of the pair matrix. Workers report per-pair progress via dcp.progress() at 1% intervals.

The job dispatcher maintains a min-heap leaderboard of the top 1,000 pairs and writes to disk every 100 slices. Memory usage at the job dispatcher remains constant regardless of job scale since full result payloads are discarded immediately after the leaderboard insertion step.

Results

The following results reflect 100% of the total search space, evaluated across all six classifiers. The leaderboard retains the top 1,000 pairs by AUC across all runs.

Figure 1 — Final leaderboard · top gene pairs by AUC · all classifiers · 100% coverage
#AUCACCCLFGene 1Gene 2
10.9270.833LRLINC00189KLF5
20.9180.818LRKLF5LINC02218
30.9170.864LRKLF5L3MBTL4
40.9160.833LRSTK4KLF5
50.9140.818LRKLF5ZNF589
60.9130.788LRPDXDC1KLF5
70.9120.894LRKLF5CYP4V2
80.9100.803LRKLF5PDPK1
90.9090.773LRKLF5PRKCB
100.9090.864LRSNX29P2PDXDC1
AUC = area under ROC curve · ACC = accuracy · 10-fold CV · n=66 · all 6 classifiers
Figure 2 — Gene frequency in top 1,000 pairs · all classifiers
KLF5 appears in 712 of the top 1,000 pairs — 71.2% of the entire leaderboard. This dominant, consistent signal across diverse partner genes and all six classifiers strongly implicates KLF5 as a core discriminator between PsA and Ps.
Figure 3 — AUC score distribution · all classifiers · full search space
Distribution of AUC scores across the top 1,000 retained pairs. The leaderboard is tightly clustered between 0.85 and 0.93, reflecting a high-quality signal pool after exhaustive search.
Figure 4 — Classifier performance comparison · average AUC in top 1,000
Average AUC by classifier among top 1,000 pairs. RF leads marginally at 0.875, with all six classifiers within 0.007 of each other — indicating the signal is robust across model families, not an artefact of any single method.

Key Findings

Across the complete search space evaluated with all six classifiers, three genes emerge with striking consistency. Two are entirely uncharacterized in the PsA literature.

KLF5
712 / 1,000 top pairs · best AUC 0.927
Krüppel-like Factor 5 is a transcription factor governing sphingolipid metabolism and skin barrier function. Published research (Lyu et al., 2024) establishes KLF5 as suppressed in psoriatic lesions relative to normal skin. What the literature has not previously addressed is whether KLF5 is differentially expressed between Ps and PsA specifically. Its dominance here across 712 of the top 1,000 pairs and all six classifiers suggests that skin barrier dysfunction severity may track with joint progression, which would be a clinically meaningful and novel finding if validated in an independent cohort.
LINC00189
Top-ranked pair · AUC 0.927 with KLF5
LINC00189 is a long intergenic non-coding RNA with essentially no published functional characterization. It has not appeared in any PsA or psoriasis literature. Its appearance at rank 1 paired with KLF5 (AUC 0.927) and again at rank 2 paired with TREML3P (AUC 0.906) makes it the most unexpected finding in the dataset. lncRNAs are known to act as chromatin remodelers and transcriptional regulators in immune cells, and some have established roles in autoimmune disease, but LINC00189 specifically is uncharted territory.
TREML3P
26 / 1,000 top pairs · rank 2 partner to LINC00189
TREML3P is a pseudogene in the TREM gene cluster on chromosome 6, adjacent to TREM1 and TREM2, which are key regulators of myeloid cell inflammatory responses in joints. Pseudogenes were long considered non-functional, but growing evidence suggests their transcripts act as competing endogenous RNAs that sponge microRNAs and modulate expression of neighboring genes. TREML3P has no published connection to PsA or psoriasis. Its proximity to the TREM locus makes a biological mechanism plausible, but this remains speculative without functional validation.
TMC4
52 / 1,000 top pairs · #2 by frequency
Transmembrane Channel-Like 4 is poorly characterized in inflammatory contexts. It appears as the second most frequent gene by leaderboard count and pairs with NMNAT2 (an NAD+ synthase primarily studied in neurodegeneration) at AUC 0.903. NAD+ depletion has been implicated in inflammatory signaling broadly, but the specific connection to PsA is uncharted. The TMC4 and NMNAT2 combination has no precedent in the psoriasis or arthritis literature.

Additional genes of interest include MMP8 (a neutrophil collagenase involved in joint tissue destruction, biologically coherent as a PsA signal), LCN2 (Lipocalin-2, an acute-phase inflammatory protein elevated in PsA), and B3GALNT2 (a glycosyltransferase with no established role in PsA, appearing in 15 top pairs).

Literature note: Lyu et al. (2024), KLF5 governs sphingolipid metabolism and barrier function of the skin, established KLF5 as a disease-relevant transcription factor in psoriasis relative to normal skin. The present results raise the possibility that KLF5 also discriminates between Ps and PsA, a clinically distinct and more actionable question. The top-ranked pair (LINC00189 + KLF5, AUC 0.927) involves a gene with no prior characterization in this disease context, and should be interpreted cautiously until validated in an independent cohort.

DCP Architecture

This workload is a natural fit for DCP's distributed compute model. The pair search is embarrassingly parallel. Each slice is independent, stateless, and produces a compact result. No inter-worker communication is required, and leaderboard aggregation happens exclusively at the job dispatcher.

Why CPU, not GPU?

The core computation, 10-fold cross-validation of logistic regression on a 66-patient dataset, is an iterative sequential optimization (L-BFGS / liblinear solver) with data dependencies between steps. This maps poorly to SIMD GPU parallelism. CPU workers on DCP are the correct resource, with parallelism expressed at the pair-batch level rather than within each evaluation.

Scale estimate

An intractable problem, made tractable with the combined power of thousands of otherwise-idle machines.

[ total slices ] 6 classifiers × 29,201 genes × (29,201 / 5,000) ≈ 1.02M slices [ single machine ] ~25 years at 12.9 min/slice [ 1,000 workers ] ~222 hours (~9.3 days) [ 10,000 workers ] ~22 hours

Collaborators

This project is conducted in collaboration with clinical and computational researchers across Canadian academic institutions. Data governance and privacy follow applicable institutional data sharing agreements.

Contribute compute. Enable future searches.

This search is complete, but the next phases — triple-gene combinations, additional datasets, and multi-site federated validation — will require significant compute. Every DCP worker that joins the network contributes to research that could surface novel biomarkers for diseases like PsA. Workers receive gene indices and publicly available expression data, not patient records.

RUN WORKER

Next Steps

01
Exhaustive pair search: all 6 classifiers · GSE57383
Complete. The full 426M pair space has been evaluated across LR, L1, RF, BA, SVM, and DT on the GSE57383 Ps vs PsA dataset. Top findings include LINC00189 + KLF5 at AUC 0.927.
✓ done
02
Replicate on GSE57405 and GSE61281
Run the same exhaustive pair search on two independent Ps vs PsA datasets. Findings that appear consistently across all three datasets are substantially more likely to represent stable biological signal rather than dataset-specific noise. LINC00189, KLF5, and TREML3P are the primary genes to watch for cross-dataset stability.
03
Single-gene baseline evaluation
Evaluate all 29,201 individual genes as solo classifiers to establish a baseline AUC floor. Pairs that substantially exceed their constituent genes' solo AUC represent genuine synergistic signal rather than one strong gene carrying the pair.
04
Guided triple-gene search
Using the top genes stable across all three datasets as seeds, exhaustively test 3-gene combinations. Three-gene classifiers are expected to push AUC meaningfully beyond the current 0.927 ceiling while remaining clinically interpretable.
05
Treatment response prediction
Extend the classifier framework to predict which patients respond to which biologic class: TNF inhibitors, IL-17 inhibitors, IL-23 inhibitors, JAK inhibitors, and PDE4 inhibitors. Currently no biomarker guides treatment selection in PsA and the clinical approach is trial and error. Gene expression signatures that predict drug class response could meaningfully shorten the time to effective treatment for patients.
06
Federated multi-site validation
Deploy the search across independent hospital cohorts using DCP's federated learning framework. Each site returns model coefficients, not patient data. Signatures that generalize cross-site are the strongest clinical candidates for further investigation.
07
LINC00189 and TREML3P functional characterization
The top-performing genes include an uncharacterized lncRNA (LINC00189) and a pseudogene adjacent to the TREM myeloid immune receptor locus (TREML3P). If cross-dataset stability is confirmed, these warrant dedicated functional investigation to understand whether they play a causal role or are markers of an underlying regulatory mechanism.

Project Video