Friday, July 31, 2026

Major News: FDA Posts Its 28pp Review of ArteraAI - Breast Cancer Computational Pathology

In May 2026, Arera AI got FDA clearance for its computational pathology test for breast cancer.  FDA can take "weeks to months" to post its 20- to 30-page review packet - and the breast cancer review is now public.

Artera AI Breast Cancer (NEW)

See the K254114 product home page here;

https://www.accessdata.fda.gov/scripts/cdrh/cfdocs/cfPMN/pmn.cfm?ID=K254115

See the FDA's 28-page 510(k) review here:

https://www.accessdata.fda.gov/cdrh_docs/reviews/K254115.pdf

FDA notes that the application includes a Predetermined Change Control Plan PCCP for updates.

See the four-page letter here:

https://www.accessdata.fda.gov/scripts/cdrh/cfdocs/cfPMN/pmn.cfm?ID=K254115

See the 864.3755 regulatory classification here, product classification SHW:

https://www.accessdata.fda.gov/scripts/cdrh/cfdocs/cfpcd/classification.cfm?id=5520

Pathology software algorithm device analyzing digital images for breast cancer prognosis

From 2025: Artera AI Prostate Cancer (2025)

FDA's de novo approval of Artera AI Prostate dates to last summer, July 2025.  See the product home page here for DEN240068:

https://www.accessdata.fda.gov/scripts/cdrh/cfdocs/cfpmn/denovo.cfm?id=DEN240068

See the classification order (letter) here:

https://www.accessdata.fda.gov/cdrh_docs/pdf24/DEN240068.pdf

See the 24-page Decision Summary here:

https://www.accessdata.fda.gov/cdrh_docs/reviews/DEN240068.pdf

See the 864.3755 regulatory classification here, product classification SFH

https://www.accessdata.fda.gov/scripts/cdrh/cfdocs/cfpcd/classification.cfm?id=5519

Pathology software algorithm device analyzing digital images for cancer prognosis

NOTE:

As shown above, the breast cancer home page product classification links to a product classification with "5520" in the link, and states it is under 864.3755, 510(k), last update 7/27/2026. Text re: breast cancer.

The prostate cancer product profile product classification links to a product classification with "5519" in the link, and ALSO states it is under 864.3755, but gives somewhat different text (re: "cancer prognosis").  Last update 7/27/2026.  


##  CALLBACKS ##

  • I discussed then-available Artera digital pathology documents in a May 14, 2026 blog here.
  • In mid July 2026,  I released a detailed AI generated white paper on H&E prognostics vs wet-lab prognostics (Oncotype), find it here.
  • FDA by law places new lab tests into CLIA classifications.  See FDA rules here. See CLIA discussion of FDA classification of CLIA tests here.   Search the "FDA CLIA Database" here. For example, search PAIGE and get the FDA-CLIA classification for Paigue system ("high complexity.")  See also this view.


###  DETAILED AI REVIEW OF 

NEW FDA ARTERA BREAST DOCUMENT ###


Chat GPT 5.6 writes, "I read the 2026 ArteraAI Breast 510(k) decision summary (K254115) primarily against the 2025 ArteraAI Prostate De Novo (DEN240068). The breast document is interesting not merely as a second Artera product, but because it shows what happens once FDA has created a Class II pathway for this kind of pathology-AI prognostic device: the prostate De Novo established the regulatory beachhead; a year later, Artera could bring breast cancer through a conventional 510(k), explicitly citing prostate as the predicate."

Bottom line

ArteraAI Breast is a prognostic, not predictive, AI test. It takes H&E whole-slide images from an already diagnosed, resected HR+/HER2−, N0/N1 early breast cancer, plus three conventional clinical variables—age, tumor size, and nodal status—and generates an ArteraAI score, a binary Low/High classification, and associated observed 5- and 10-year distant-metastasis risks. It is not cleared to say whether chemotherapy, endocrine therapy, or another treatment will work better.

Its pivotal validation is reasonably substantial: 1,271 patients, three US sites, with clear separation between the two risk groups. At 5 years, DM risk was 0.9% Low versus 8.7% High; at 10 years, 2.8% versus 16.6%. Sixty-five percent of patients were Low and 35% High.

The regulatory story may be almost as interesting as the clinical story.

Prostate 2025 → Breast 2026

ArteraAI Prostate, 2025ArteraAI Breast, 2026
FDA routeDe Novo DEN240068510(k) K254115
Regulatory roleCreated new Class II device typeUses prostate as predicate
Regulation21 CFR 864.3755Same, 21 CFR 864.3755
Product codeSFHSHW
Input tissueProstate core biopsy, H&E WSIBreast resection, H&E WSI
Other algorithm inputsEssentially image onlyAge + tumor size + nodal status
Primary clinical output10-y DM + PCSM5-y and 10-y DM
CategoriesLow / Intermediate / HighLow / High
Pivotal validationn=886n=1,271
PCCPAdd FDA-cleared WSI scannersAdd FDA-cleared scanners and file formats

FDA's breast decision summary expressly identifies DEN240068 ArteraAI Prostate as the predicate and then lays out the similarities and differences. The breast indication is substantially different biologically, but FDA regarded the underlying device concept—locked deep-learning software operating on FFPE H&E WSIs to produce prognostic cancer-risk information—as sufficiently similar for substantial equivalence.

That is a rather important precedent. The De Novo was the hard regulatory step. Once FDA classified “software algorithm device analyzing digital images for cancer prognosis” as Class II and established the special-control framework, Artera had a predicate from which to extend the platform into another tumor type. The prostate decision explicitly concluded that general controls alone were insufficient but that the special controls made the benefit-risk acceptable and created the Class II device type. The breast review, by contrast, ends with the much shorter 510(k) conclusion that the evidence supports substantial equivalence.

What exactly is ArteraAI Breast?

This is worth emphasizing because it is easy to describe it too loosely as an “AI pathology test.”

It is actually a multimodal prognostic algorithm. The AI consumes:

  1. H&E WSIs from breast resection tissue; and

  2. physician-supplied age, tumor size, and nodal status.

The deep-learning engine combines clinical variables with image-derived features; the model is locked, rather than continuously learning.

The report gives:

  • an ArteraAI risk score;

  • Low versus High;

  • observed 5-year DM risk from the clinical validation data; and

  • observed 10-year DM risk from that data.

Thus it is somewhat different conceptually from a pure “pixels-to-prognosis” device. Age, tumor size and nodes are actually inside the algorithm, not merely displayed alongside its result. In prostate, FDA says that additional clinical data entered into the portal were not used as algorithm inputs.

That distinction matters both scientifically and commercially: some of the breast score's performance presumably comes from three already quite prognostic clinical variables, in combination with the morphology signal.

The breast clinical evidence

The pivotal cohort contained 1,271 HR+/HER2−, pT1–T3, pN0–N1, pM0 patients, diagnosed between 2006 and 2019, retrospectively assembled from three US sites. Patients had undergone surgery and endocrine therapy; chemotherapy and other standard care were permitted.

The cohort looks clinically recognizable for the intended population:

  • median age 62;

  • 84% N0, 16% N1;

  • 65% received endocrine therapy only;

  • 23% had chemotherapy;

  • 34% grade 1, 55% grade 2, 11% grade 3;

  • 65% classified Artera Low, 35% High.

One weakness worth noting is that although there are three sites, one site supplies 1,031 of the 1,271 patients—81% of the entire validation cohort. The other two contribute only 186 and 54. That is much less balanced than the prostate pivotal study, whose three sites contributed 33%, 50% and 17%, respectively. The prostate study had 886 patients.

Clinical discrimination is quite clear

The 5-year result is:

  • Low: 827 patients; 7 DM events; 0.9% estimated risk.

  • High: 444 patients; 38 events; 8.7%.

  • Overall: 3.6%.

At 10 years:

  • Low: 15 events; 2.8%.

  • High: 57 events; 16.6%.

  • Overall: 7.6%.

Those are clinically meaningful separations, and FDA explicitly calls the differences statistically and clinically significant.

There is also evidence that the continuous score contains information beyond the single Low/High cutoff. FDA presents four score bins:

  • score 5.9–<25 → 10-y DM 2.1%

  • 25–30 → 3.5%

  • 30–<50 → 13.1%

  • 50–68 → 27.7%

The same monotonic pattern occurs at five years.

The graph on page 26 is in some respects the most informative figure in the document: it shows that the binary cutoff at about 30 is convenient for reporting, but the underlying score behaves more like a graded prognostic variable.

One small but interesting caveat is also visible there: the validation cohort had no patients with scores >70. The figure assigns the >70 region the same 15.6% 5-year and 27.7% 10-year risks as the 50–68 group, with the explicit footnote that those estimates are based on patients scoring 50–68.

So I would not describe the FDA-cleared output as a perfectly calibrated individualized probability across an unrestricted 0–100 continuum. FDA's own description is more careful: the report includes the score, classification, and observed risks from the clinical-validation dataset.

Comparison with the prostate clinical result

Prostate had a somewhat more dramatic High-risk separation:

  • Low: 3.3% 10-year DM

  • Intermediate: 6.6%

  • High: 28.1%

  • Overall: 8.1%.

It also predicted prostate-cancer-specific mortality, with 10-year PCSM of 0.6% Low, 1.1% Intermediate and 10.2% High.

Breast is therefore in one respect simpler: two risk classes and one clinical endpoint, distant metastasis, but it adds the five-year horizon and a continuous score.

I would not compare the 28.1% prostate High risk with the 16.6% breast High risk as though this shows that one algorithm is better. They are entirely different diseases, populations, treatments and cutoffs. What is comparable is that FDA accepted essentially the same paradigm: use long-term outcomes from archived randomized-trial/clinical-study material for model development, and then demonstrate clinically meaningful risk stratification in an independent retrospective US cohort.

An interesting point about model development

The breast model was developed from multiple prospective trial datasets and, notably, included both pretreatment biopsies and surgical slides during development, although the final intended-use specimen is the breast resection. FDA says the development data came from WSG ADAPT, WSG PlanB, NSABP B34 and ABCSG 6 and incorporated the image data plus age, tumor size and nodal status.

The pivotal test set, however, properly corresponds to the cleared use: pretreatment H&E slides from the surgical specimen, with the highest-grade/highest-tumor-content slide selected.

Race subgroup: reassuring ordering, but I would be cautious about calibration

There are 140 African American patients versus 1,069 White patients. At five years the overall DM rate was almost identical—3.6% versus 3.7%—and Low remained lower than High in both groups.

At ten years, however, there is a notable absolute-risk difference:

  • White High: 20.1%

  • African American High: 6.4%

  • White Low: 2.8%

  • African American Low: 3.2%.

The African American High group contains only 63 patients and four 10-year DM events, with a wide 95% CI of 2.4%–16.1%. So the directional discrimination remains, but I would be reluctant to infer that the absolute risk estimates are equally calibrated across racial groups.

Interestingly, the prostate De Novo had almost the opposite-looking subgroup observation: among its 72 African American patients, the estimated 10-year DM risks were higher than among non-African American patients, and FDA explicitly cautioned that the African American subgroup was limited.

PCCP — this is more significant than a housekeeping paragraph

The breast submission actually lists establishment of a PCCP as one of its two purposes, right alongside clearance of the new device. The authorized change is the ability to add additional FDA-cleared interoperable WSI scanners and file formats.

This is not Artera's first PCCP. The prostate De Novo already contained one. The prostate plan authorized later addition of FDA-cleared WSI scanners; its intended-use language explicitly contemplated either the originally authorized scanner or another 510(k)-cleared scanner qualified under the PCCP.

The evolution is:

Prostate 2025:
Philips Ultra Fast was the initial scanner. PCCP created a mechanism to qualify additional FDA-cleared WSI scanners.

Breast 2026:
The device starts out cleared with both Philips Ultra Fast and Leica Aperio GT450 DX, and the PCCP covers future FDA-cleared scanners plus their file formats.

That is quite practical. Scanner dependence has been one of the potential regulatory bottlenecks for pathology AI. Without some such mechanism, every new scanner or image format potentially becomes a device modification requiring FDA analysis about whether another premarket submission is needed. The PCCP pre-specifies how Artera can make this class of modification and validate it.

What FDA allows Artera to change

For a new scanner, Artera may:

  1. modify the UI to offer the new scanner in the upload workflow;

  2. modify the backend so that it recognizes/verifies the new scanner's image metadata; and

  3. where needed, modify the image converter so that the scanner's file can be converted into the format expected by the AI engine.

The core prognostic algorithm is not what is being opened for modification. The AI remains a locked model. The PCCP is basically an interoperability PCCP, not permission to retrain the breast prognostic model or alter its cutoff.

Before adding a scanner, Artera must use the specified verification/validation procedure and acceptance criteria to demonstrate comparable performance and assure that the software modification has not adversely affected already supported scanners or other device functions. Once validation succeeds, the labeling can be updated to list the additional interoperable scanner.

The analytical work in the breast submission explains why the PCCP is credible

FDA already has direct evidence that the same locked algorithm can operate with two quite different scanner/file ecosystems:

  • Philips: 40×, 0.25 μm/pixel, proprietary iSyntax

  • Leica GT450 DX: 40×, 0.26 μm/pixel, SVS.

They tested 50 breast specimens across three labs, over five nonconsecutive days, on both systems, including cases close to the Low/High cutoff.

Reproducibility CVs were up to 8.4% for Philips and 4.7% for Leica. Most samples had 100% categorical agreement; unsurprisingly, the few disagreements clustered around the cutoff. For example, a Philips borderline case with mean score 30.4 was Low in 53% and High in 47% of replicate measurements. A couple of Leica borderline cases likewise crossed the cutoff.

That is actually useful disclosure rather than alarming: scanner/day variability matters very little away from the threshold and can flip classifications close to it. It also explains why a rigorous scanner-qualification PCCP is needed.

The prostate PCCP was strikingly similar in architecture: qualify another FDA-cleared scanner, update UI/backend/converter as necessary, validate against the original Philips performance, then update labeling. Breast essentially inherits and modestly broadens that playbook.

Chat GPT writes:
What I think is most noteworthy

[The AI writes] My read would be:

1. FDA has now validated a regulatory platform, not merely two isolated algorithms.
The 2025 De Novo created the Class II category; the 2026 breast clearance demonstrates that a substantially different cancer prognosis application can use that prostate device as the predicate. That should make additional Artera tumor types—assuming comparable evidence—much more straightforward regulatory projects.

2. The breast product is less “pure AI pathology” than one might assume.
It mixes morphology with age, tumor size and nodal status. Consequently, one should not attribute all of its discrimination to AI extraction of previously invisible H&E morphology.

3. The clinical separation looks quite respectable.
A 65% Low population with only 0.9% 5-year and 2.8% 10-year DM, against 8.7% and 16.6% in High, is a potentially clinically useful prognostic distinction.

4. But this FDA clearance does not establish treatment prediction.
The chemotherapy/endocrine subgroup tables show that the prognostic stratification persists within treatment groups; they do not show an interaction proving that one Artera group derives more or less benefit from chemotherapy. FDA consistently describes the device as assisting prognostic risk-based decisions.

  • That distinction will be especially important if Artera eventually positions this against genomic breast assays whose clinical value propositions include informing chemotherapy decisions.

5. PCCP is strategically important.
I would give it more attention than the few pages FDA devotes to it. Artera has effectively separated two kinds of change: the medical algorithm stays locked, while the digital-pathology infrastructure around it can expand in a pre-authorized, validated fashion. The prostate De Novo introduced this concept; breast demonstrates its reuse and extends it to file formats as well as scanners.

6. The breast validation is larger than prostate, but less evenly multicenter.
n=1,271 versus 886 is favorable, but the 81% contribution from one breast site is worth keeping in mind. Likewise, the race subgroup analyses preserve risk ordering but are too small to give me great confidence in subgroup-specific absolute calibration.

If the purpose of your review [of the new FDA breast cancer device]  is partly reimbursement/clinical adoption, the next question I would focus on is quite different from FDA's question: 

  • What does ArteraAI Breast add over clinicopathologic variables alone  ----  and 
  • Over established breast genomic assays such as Oncotype DX? 

The FDA summary convincingly establishes prognosis and substantial equivalence; it does not in this document demonstrate the incremental prognostic value of the image AI over age + tumor size + nodes, or demonstrate predictive utility for chemotherapy. Those are likely to be much more central to payer evidence than the 510(k) itself.