Sunday, October 4, 2026

Is Sparano et al. Better than Oncotype DX? New Publication.

From time to time a test is published as "Better than Oncotype Dx," which has been around since about 2004.

Some have even been based on H&E slides, WSI, and AI.

Here's a complex multi-modal tour-de-force approach to a new test.  Here's some pre-publicity from 12/2025 on the project, from Caris.  

(See supplements at bottom of blog, for a tech critique of Sparano AI (via Dawood), and views about the test and Caris financials.)

##

“Better Than Oncotype”? A New Breast Cancer Study Combines AI, Genomics, and Clinical Data

Recent coverage in ASCO AI in Oncology highlights an intriguing result: a new algorithm predicted breast cancer recurrence more accurately than Oncotype DX. The underlying study, published by Sparano and colleagues in npj Breast Cancer, combines digital pathology, gene expression, and clinical information. It offers substantial evidence of improved prognosis prediction—and several questions about what would make that improvement clinically useful. [1,2]



What the investigators built

The model, IICM+, integrates three information streams: clinical factors such as age, menopausal status, tumor size, and grade; molecular features from an expanded 42-gene panel, obtained through whole-transcriptome sequencing; and AI-derived features from digitized H&E slides.

The image analysis captures both small tissue regions and broader slide-level patterns. A multimodal transformer combines these representations with clinical and molecular features to generate a continuous recurrence-risk score, subsequently classified as high or low risk. Thus, the study evaluates an integrated diagnostic approach requiring both imaging and molecular information. [1]

How substantial—and independent—was the validation?

The investigators analyzed 4,429 TAILORx participants with hormone receptor–positive, HER2-negative, node-negative breast cancer. The development cohort contained 2,808 patients, with 202 distant recurrences and median follow-up of 11.7 years. The validation cohort contained 1,621 patients, with 109 distant recurrences and median follow-up of 11.2 years. These event counts matter: thousands of patients provide valuable data, but the number of recurrences ultimately limits statistical precision.

The algorithm underwent five-fold cross-validation during development. After selecting its specifications and tuning parameters, investigators retrained it on the development cohort and locked both the model and its risk threshold before applying it to the holdout cohort.

The separation was meaningful: patients were assigned by enrolling cooperative group. ECOG-ACRIN, Alliance, and Canadian Cancer Trials Group participants supplied development data; SWOG, NRG Oncology, and Cancer Trials Support Unit participants supplied validation data.

This constitutes independent holdout validation within TAILORx. It remains a retrospective analysis of specimens from the same parent trial, conducted by the developing investigators; replication in a separate contemporary population would strengthen confidence in generalizability. [1]

How much better than Oncotype?

In the holdout cohort, the reported C-indices were:

Distant recurrence endpointOncotype DXIICM+
Overall0.5780.735
Within five years0.7220.791
After five years0.5140.710

A C-index measures how well a model ranks patients by risk; it is not the percentage of patients correctly diagnosed. The largest separation appeared in late recurrence. Overall and late comparisons had P values below .001; the early comparison was more modest, at P=.046. [1]

An especially informative comparison appears beneath the headline: the model using clinical information plus expanded molecular features, without imaging, achieved an overall C-index of 0.721, versus 0.735 for IICM+. The authors describe their holdout performance as comparable. The study therefore provides stronger evidence for the integrated model’s advantage over Oncotype alone than for the incremental value of its imaging component. [1]

What should happen next?

Several priorities follow from these results.

First, investigators should validate the locked model in additional populations and laboratory workflows, including assessment of whether predicted absolute risks match observed outcomes. Second, they should establish how much each component adds—and whether the imaging contribution justifies its operational complexity.

Third, the test needs a clearly defined treatment decision. Improved prediction of recurrence does not itself establish who benefits from chemotherapy, extended endocrine therapy, or a CDK4/6 inhibitor. For example, a chemotherapy-treated patient with a high Oncotype score and a low IICM+ risk cannot automatically be assumed to have safely avoided chemotherapy: treatment may have contributed to the favorable outcome.

Treatment-benefit analyses using randomized trial specimens, followed by prospective studies of test-guided care, would help bridge that gap. For reimbursement, the central question will be whether the test changes management in a way that improves outcomes or safely reduces treatment burden.

Is a company positioned to commercialize it?

Caris Life Sciences is already deeply involved. This is a public–private research collaboration built on federally supported TAILORx infrastructure; Caris supplied scientific capabilities and funding, and multiple authors are company employees. [1,3]

There is also a concrete commercial product: Caris launched MI Clarity on May 5, 2026, for postmenopausal patients with HR-positive/HER2-negative, node-negative early breast cancer. Caris describes it as analyzing H&E images and clinical inputs to assess early and late distant recurrence, without requiring genomic sequencing. [4,5]

That difference is consequential. MI Clarity and the paper’s full IICM+ model have different stated inputs, so IICM+’s performance numbers should not automatically be assigned to the marketed assay. Commercial translation is already underway in this research program; the next question is how the exact marketed test, its validation evidence, and its intended treatment decisions align.

See "MI CLARITY" isn't yet listed in the MolDx DEX registry.



References and links

  1. Sparano JA, Lama N, Gray RJ, et al. An Artificial Intelligence (AI) model integrating multiscale foundation model histopathology representations with molecular and clinical features predicts early and late distant recurrence in TAILORx. npj Breast Cancer. Published online August 5, 2026; advance publication. doi:10.1038/s41523-026-01022-y. Full paper.

  2. LaGrego W. Multimodal AI Model May Improve Prediction of Breast Cancer Recurrence Risk. ASCO AI in Oncology. September 28, 2026. Article.

  3. ECOG-ACRIN Cancer Research Group. A new AI-powered test offers more accurate prediction of breast cancer recurrence risk. September 17, 2026. Press release.

  4. Caris Life Sciences. Caris MI Clarity launch announcement. May 5, 2026. Company press release.

  5. Caris Life Sciences. Caris MI Clarity. Product information; accessed October 4, 2026. Product page.

Source copies reviewed

Supplements

:This method uses WSI AI, Dawood et al say that such papers may overestimate their own accuracy.  Discuss Sparano in light of Dawood.

##

Dawood does not invalidate Sparano’s findings, but it provides a strong reason to distinguish the performance of the complete test from the added value of digital imaging. Sparano’s own comparisons make that distinction particularly important.

The papers ask different questions

DawoodSparano
Main questionDoes AI infer specific molecular biomarkers from H&E, or exploit correlated features?Can imaging, molecular data, and clinical factors jointly improve recurrence prediction?
Main concern or findingGood overall accuracy can conceal dependence on grade, other biomarkers, or cohort composition.A combined model outperformed Oncotype DX for recurrence discrimination in held-out TAILORx patients.
What imaging must demonstrateReliable biomarker prediction beyond familiar histologic correlates.Additional prognostic information beyond the available clinical and molecular inputs.

There is a consequential biological distinction. Tumor grade is a legitimate predictor of recurrence. An outcome model can use grade-associated morphology productively without demonstrating that it has discovered a gene-specific visual signature. My interpretation is that Sparano’s prognostic objective is therefore less directly exposed to Dawood’s critique of H&E-to-biomarker inference.

Nevertheless, the question remains: how much additional information does the expensive, sophisticated image analysis supply?

Sparano’s own table supplies the most revealing comparison

The following are results from Sparano’s independent holdout cohort:

ModelOverall recurrence C-indexEarly recurrenceLate recurrence
Oncotype DX alone0.5780.7220.514
Clinical variables + Oncotype DX0.6570.7720.606
Clinical + expanded molecular features, without imaging0.7210.7730.698
Full IICM+, including imaging0.7350.7910.710
Numerical increment over clinical + expanded molecular model+0.014+0.018+0.012

The full model’s advantage over Oncotype alone is substantial. Its numerical advantage over the strongest non-imaging model is much smaller. The authors explicitly describe the latter models as having comparable holdout performance.

These differences do not establish that imaging is useless; nor should the C-index increments be treated as a mathematical decomposition of each modality’s contribution. They do show why “the multimodal test beats Oncotype” is much better supported than “digital imaging accounts for the improvement.”

Source: Sparano, Table 1 and Discussion.

Independent validation helps—but does not resolve Dawood’s concern

Sparano deserves credit for locking the model and risk threshold before testing them in 1,621 patients assigned to a separate holdout cohort by enrolling cooperative group. That provides meaningful protection against simply fitting and evaluating the algorithm on the same patients.

Dawood’s point goes beyond overfitting. A model can perform well in independent patients if the same underlying correlations persist. Their ER example is especially relevant: overall AUROC increased from 0.87 in development cross-validation to 0.90 in an external cohort, yet further analysis indicated that a stronger relationship between grade and ER helped explain the apparent success. The sophisticated model provided limited additional information beyond grade.

Thus, a holdout test answers whether a prediction carries into new patients under those conditions. It does not, by itself, establish what information the model uses or how much novelty resides in its imaging component.

Source: Dawood, Discussion.

What would a Dawood-informed evaluation of Sparano require?

I would prioritize three analyses:

  • Isolate imaging’s incremental contribution. Compare the locked full model with an appropriately optimized clinical-plus-molecular model in additional cohorts, using paired statistical comparisons, calibration, and decision-relevant measures.
  • Test within clinically similar groups. Does imaging improve prediction among patients with the same grade, similar tumor size, menopausal status, molecular risk, and treatment? An aggregate result can conceal weak performance in the actual subgroup facing a difficult decision.
  • Separate prognosis from treatment benefit. Test whether image-derived information helps identify differential benefit from treatment. Better prediction of recurrence under treatments already received does not establish which treatment should have been chosen.

One further distinction: showing that the full IICM+ score remains significant after adjustment for grade and Oncotype does not isolate an independent imaging effect. The expanded molecular features could contribute to that residual association.

For your readers, the fair characterization is: Sparano provides credible evidence for improved multimodal prognosis prediction, while leaving a more modest and incompletely resolved case for the incremental contribution of digital imaging. Dawood explains why that second question deserves its own evidence.

References

  • Dawood M, Branson K, Tejpar S, Rajpoot N, Minhas FUA. Confounding factors and biases abound when predicting molecular biomarkers from histological images. Nature Biomedical Engineering. 2026. doi:10.1038/s41551-026-01616-8.
  • Sparano JA, Lama N, Gray RJ, et al. An Artificial Intelligence (AI) model integrating multiscale foundation model histopathology representations with molecular and clinical features predicts early and late distant recurrence in TAILORx. npj Breast Cancer. 2026. doi:10.1038/s41523-026-01022-y.
##
##
Do you see much in the financial press about MI Clarity (or similar Sparano testing) boosting the bottom line at Caris?
##

Some—but I’m not finding a substantial financial-news narrative that MI Clarity, or the full Sparano IICM+ model, has materially increased Caris’s shareholder value. MI Clarity appears as one growth opportunity among several. The strongest evidence of serious investor attention is an earnings-call question specifically about its reimbursement.

SourceWhat it saysSignificance
Investing.com, May 5Covered the MI Clarity launch alongside Caris’s financial position and analyst ratings.Primarily launch coverage; it does not establish a Clarity-driven valuation increase.
Caris Q2 earnings call, August 5Management discussed MI Clarity V2; an analyst asked about hiring, coding, and coverage.Concrete evidence that financial analysts are monitoring commercialization.
Hudson Labs, September 27Listed MI Clarity V2 among potential catalysts and asked what new launches would contribute to revenue and margins.A prospective opportunity, with its financial contribution still an open question.

These sources support investor awareness, rather than a demonstrated product-specific boost to shareholder value. investing.com

The reimbursement exchange is particularly relevant to your work. On the August 5 call, Jack Meehan of Nephron Research asked whether MI Clarity V2 needed its own code and coverage, noting hiring to support the launch. CFO Luke Power said Caris was evaluating which LCD might apply, approaching commercial payers directly, and anticipated that Medicare payment could require an updated LCD. Those were management’s statements at that time, rather than confirmation of coverage. stockanalysis.com

Management also described V2 as adding chemotherapy and extended-endocrine-therapy decision support. That investor-facing development plan goes beyond what the Sparano prognostic paper itself establishes. stockanalysis.com

The broader financial story appears more concentrated on molecular-profiling volume and reimbursement, margins, and Caris Detect’s early-detection opportunity. For example, the company attributed its Q2 revenue growth primarily to molecular profiling, while an Investing.com earnings preview identified Detect as a major analyst focus. sec.gov

My reading: MI Clarity is a credible commercial initiative that investors are watching, but I did not locate a clear analyst price-target increase, a quantified valuation contribution, or a stock-price move specifically attributed to it—or to the Sparano publication. Publicly accessible coverage is limited; that finding does not exclude discussion in subscription analyst reports.