Neural networks in digital pathology: how a model reads a slide
What a whole-slide image is, how a model learns from annotated tiles, what studies and clearances report, what EU rules require and what a hospital needs.
A tissue slide stained with haematoxylin and eosin is, for a pathologist, an image read in a few minutes. For a machine-learning model, the same slide is a file of billions of pixels that cannot be looked at in one pass. This article explains how a slide becomes data, how a neural network learns from that data, what results published studies and market clearances report, what the European framework says, and what a pathology department needs in order to take part in such projects.
What a whole-slide image is
A whole-slide image (WSI) is a complete scan of a glass slide at microscopic resolution. The DICOM standards committee, which published Supplement 145 on whole-slide imaging, gives a concrete example: a 20 × 15 mm sample scanned at 0.25 micrometres per pixel (what practice calls "40x") yields an image of roughly 80,000 × 60,000 pixels, or 4.8 gigapixels. At 24-bit colour, the raw data is about 15 GB. Scanning at 0.5 micrometres per pixel is called "20x".
According to DICOM, lossless compression reduces size 3–5 times, and JPEG2000 30–50 times, so the 15 GB example comes down to about 300 MB. The image is stored as a pyramid of resolutions: the base level holds full detail, upper levels hold progressively smaller versions, and the viewer loads only the tile it needs at the current zoom. In production, a large cancer centre in the United States reports an average of 2 GB per slide scanned in a single focal plane, according to Ardon et al. (Journal of Pathology Informatics, 2023).
How a model learns from annotated tiles
A neural network cannot process 4.8 gigapixels in one step. The slide is cut into small fragments, usually squares of a few hundred pixels, called tiles. The model learns on these tiles, and the results are then aggregated at slide level.
Two families of architecture are in use today. Convolutional neural networks (CNNs) apply local filters that detect edges, textures and then increasingly complex structures, from nuclei to glands. Vision transformers (ViTs) split a tile into patches and learn the relationships between all patches at once, through an attention mechanism. The UNI model, published by Chen et al. in Nature Medicine in 2024, is a vision transformer pretrained without labels, through self-supervised learning, on more than 100 million tiles from more than 100,000 haematoxylin-and-eosin slides across 20 major tissue types, then evaluated on 33 clinical tasks. Such a "foundation model" diagnoses nothing by itself: it provides a general representation of tissue, on top of which task-specific models are trained with far less data.
Labels can come at two levels. Pixel-level annotation means a pathologist outlines every tumour region; it is precise but time-consuming. The alternative is weak supervision: the model receives only the slide's diagnosis, as it appears in the pathology report, and learns by itself which tiles explain that diagnosis. Campanella et al. (Nature Medicine, 2019) showed that this works at scale: 44,732 slides from 15,187 patients, scanned at 20x (0.5 micrometres per pixel), with labels taken exclusively from diagnostic reports and no manual annotation.
The slide does not enter the model whole. It enters cut into tens of thousands of tiles, and the quality of the labels attached to those tiles decides the quality of the model.
What tasks are realistic today
The literature and the authorised products currently cover four types of task.
- Tumour detection. The model flags suspicious regions or classifies the slide as "suspicious" or "no suspicion". It is the task covered by the clearances and studies cited below.
- Grading support. The model proposes, for example, Gleason patterns in a prostate biopsy, or identifies a specific pattern such as grade 4 cribriform, for the pathologist to confirm.
- Biomarker quantification. The model counts positive and negative cells on an immunohistochemical stain, for example HER2 in breast cancer, and calculates a percentage that the pathologist verifies.
- Quality control. The model detects blurred areas, tissue folds or scanning artefacts before the slide reaches the physician.
In all four cases the diagnosis remains the pathologist's. The wording of the clearances is explicit: Paige Prostate is, according to the FDA decision summary, a "software only device intended to assist pathologists in the detection of foci that are suspicious for cancer", and the Aiforia CE-IVD models "assist pathologists" in assessing HER2 or detecting prostate patterns.
What published studies and clearances report
The first public benchmark was the CAMELYON16 competition, published by Ehteshami Bejnordi et al. in JAMA in 2017. Teams received 270 training slides and were evaluated on 129 test slides with or without lymph-node metastases of breast cancer. The metric was the area under the ROC curve (AUC), where 1.0 means perfect separation and 0.5 means guessing. The best algorithm reached an AUC of 0.994. A pathologist without time constraint reached 0.966. A panel of 11 pathologists who read the same slides in a two-hour session simulating routine pace averaged 0.810.
The authors noted that the results came from a competition, not a clinical setting. The 2019 Campanella et al. study took the next step, with high-volume routine data and no manual annotation: AUC 0.991 for prostate cancer, 0.988 for basal cell carcinoma and 0.966 for breast-cancer lymph-node metastases. The authors estimate that, used as a filter, the prostate model could remove more than 75% of slides from the pathologist's workload without loss of sensitivity at patient level.
Market clearances add a different kind of evidence, because they measure the effect on the physician, not only the algorithm's performance. Paige Prostate was authorised by the FDA in 2021 through the De Novo pathway (DEN200080), in class II, for H&E-stained prostate biopsies scanned with a Philips Ultra Fast Scanner. In the pivotal study, on 171 slides with cancer, assisted reading produced a net sensitivity gain of 7.3% per biopsy, and specificity rose from 88.45% to 89.50%. The FDA notes that the per-patient benefit would be "substantially lower", because a patient usually has several biopsy cores.
In the European Union, Aiforia Technologies (Finland) announced in February 2025 its certification under the IVDR, granted by the notified body BSI, together with CE-IVD models for HER2 in breast cancer, for the Gleason 4 cribriform pattern and for perineural invasion in prostate cancer.
The data pipeline: from archive to training set
An AI project in pathology does not start with the model. It starts with the data. The steps are the same regardless of vendor.
- Legal basis and de-identification. Health data is a special category under Regulation (EU) 2016/679 (GDPR): Article 9(1) prohibits its processing, and Article 9(2)(j) allows the exception for scientific research, on the basis of Union or national law and with suitable safeguards. Recital 26 is essential for pathology: pseudonymised data, which can be attributed to a person with additional information, remains personal data; only anonymous information falls outside the regulation. De-identifying a slide means removing the scanned label, the metadata in the file and any identifier in the file name.
- The research contract. Who is the controller, who is the processor, where the data sits, who has access, what happens to the resulting model and to the data at the end.
- Selection and scanning. Documented inclusion criteria, an agreed resolution (20x or 40x), a quality check.
- Annotation. Region-level outlining by pathologists or slide-level labels from reports, as in the Campanella approach; the protocol is written beforehand and disagreement between annotators is measured.
- Validation on external cohorts. The model is tested on slides from another centre, scanned on another device, before any conclusion is drawn.
Why external validation matters
A model that reaches an AUC of 0.99 on its own hospital's slides can do much worse at the hospital next door. Howard et al. (Nature Communications, 2021) analysed more than 3,000 patients with six cancer types from the public TCGA archive and showed that a model can recognise from the image which centre submitted the slide, with an AUROC between 0.964 and 0.998 depending on cancer type. Colour normalisation reduces the effect, but the centre remains recognisable, with an average AUROC above 0.850. The differences come from fixation, staining and the scanner.
The consequence is that the model learns the centre's "signature" instead of biology. When the authors retrained models with training and test centres kept separate, 51 of 56 predictable features (91.1%) lost AUROC, and 20 (35.7%) were no longer significantly detectable. The same phenomenon appears in clinical studies: Campanella et al. report a drop of about 6 AUC points when the prostate model trained on in-house slides was tested on consultation slides received from other institutions.
A result on a single cohort, from a single scanner, describes that laboratory, not the model. External validation is not a formality; it is the main test.
The regulatory framework in the European Union
Software that provides information for diagnostic decisions is a medical device or an in vitro diagnostic (IVD) medical device, as the case may be. The European Commission's guidance MDCG 2019-11 rev.1 (June 2025) explains the qualification; the two regulations are Regulation (EU) 2017/745 on medical devices (MDR) and Regulation (EU) 2017/746 on in vitro diagnostic medical devices (IVDR).
| Legal act | What it says for pathology software | Source |
|---|---|---|
| MDR, Annex VIII, Rule 11 | Software for decisions with diagnostic or therapeutic purposes: class IIa; class III if the decision may cause death or an irreversible deterioration; class IIb if it may cause a serious deterioration or a surgical intervention. | MDCG 2019-11 rev.1 |
| IVDR, Annex VIII, implementing rule 1.4 | Software that drives or influences an IVD device takes the class of that device; independent software is classified in its own right. Example from the guidance: software classifying a cytology smear as normal or suspicious, class C under Rule 3(h). | MDCG 2019-11 rev.1 |
| AI Act, Article 6(1) and Annex I | High-risk if the system is a product, or a safety component of a product, under the legislation in Annex I, and the product requires third-party conformity assessment. Annex I lists the MDR (point 11) and the IVDR (point 12). | Regulation (EU) 2024/1689 |
| AI Act, Article 113 | General application from 2 August 2026; the obligations under Article 6(1) were originally set for 2 August 2027. | Regulation (EU) 2024/1689 |
| Digital Omnibus on AI | Regulation (EU) 2026/1744 (adopted 8 July 2026, published 24 July 2026) defers the obligations for AI in Annex I products, including medical devices and IVDs, to 2 August 2028; Annex III: 2 December 2027. | Regulation (EU) 2026/1744 |
For a hospital the practical reading is simple: a tool used in diagnosis must carry a CE mark as a medical device or IVD, with a written intended purpose, and a research prototype cannot enter clinical routine until it has completed that path.
What this means for a hospital in Romania or Moldova
Taking part in a pathology AI project does not require a model of your own. It requires good data, scanning infrastructure and an organisation that can sustain the workflow. The figures below come from a tertiary cancer centre in New York, Memorial Sloan Kettering (MSK), which published in 2023 the costs of its digital pathology operation for 2021, according to Ardon et al.
The scale at MSK is not that of a county hospital: 25 scanners in 2021, more than 6.1 million slides scanned in total and an on-premises storage cost of about USD 1 million per petabyte, redundancy included. The authors sum up the requirements in six words: sponsorship from leadership, staff, space, scanners, software and servers. Scanners accounted for 31.5% of the total cost of the operation, service contracts cost between 7% and 20% of the scanner price per year, and the cost per slide depended above all on how intensively each scanner was used.
Concrete steps for a laboratory that wants to join a project:
- Take stock of the archive. Slides per year, cases in the tumour types of interest, stains, age of the blocks.
- Size storage from 2 GB per slide. For 20,000 slides a year, roughly 40 TB annually, before redundancy.
- Fix the resolution and the format. 20x or 40x, proprietary format or DICOM; the choice affects file size and compatibility.
- Write down the de-identification workflow. Who removes the label and the metadata, who checks, where the pseudonymisation key is kept and when it is destroyed. As long as the key exists, the data remains personal.
- Plan the staff. The MSK benchmark is one full-time employee per 3–4 scanners, covering loading, quality control and rescanning.
- Define the pathologists' role. Protected time for annotation and for checking results, under a written protocol.
Questions to ask an AI vendor:
- On how many centres and how many different scanners was the model validated, and with what results per cohort, not only on average?
- What is the intended purpose on the CE mark, and in which class (MDR or IVDR) is the product placed?
- How does the model perform on slides scanned with our own device, on a test set we choose, before the contract?
- What happens to our data after training, and who owns the resulting model?
Consdinamic builds software and artificial intelligence to order, with its deepest specialisation in healthcare, and collaborates with Aiforia (Finland) on de-identified digital pathology data sets; for partners in the European Union, work is delivered through its sister company Softmed Labs SRL in Romania.
Conclusion
A neural network does not "see" a slide the way a pathologist does. It learns from tens of thousands of tiles, with labels that come either from manual outlines or from diagnostic reports, and it produces a probability that the physician interprets. The published evidence is solid for detection, and the FDA authorisations and IVDR certifications show that these tools can enter routine use, provided the intended purpose is clear and the pathologist remains the one who decides. Two things separate a useful project from a decorative one: validation on external cohorts, with other scanners, and a data pipeline built correctly from the first step, with de-identification and a contract.
- DICOM Standards Committee, Whole Slide Imaging in Pathology (Supplement 145) — size of a whole-slide image: 80,000 × 60,000 pixels, 4.8 gigapixels, 15 GB uncompressed, ~300 MB in JPEG2000, resolution pyramid
- Ehteshami Bejnordi et al., Diagnostic Assessment of Deep Learning Algorithms for Detection of Lymph Node Metastases, JAMA, 2017 — CAMELYON16: AUC 0.994 best algorithm, 0.966 pathologist without time constraint, 0.810 mean of 11 pathologists with time constraint
- Campanella et al., Clinical-grade computational pathology using weakly supervised deep learning on whole slide images, Nature Medicine, 2019 — 44,732 slides, 15,187 patients, slide-level labels only, AUC 0.991 / 0.988 / 0.966, 20x scanning, drop on external slides
- FDA, De Novo Decision Summary DEN200080, Paige Prostate, 2021 — indications for use, class II, 171 slides with cancer, +7.3% sensitivity per biopsy, specificity 88.45% → 89.50%
- Aiforia Technologies, IVDR certification and new CE-IVD marked products, 2025 — IVDR certification from BSI, CE-IVD models for HER2, Gleason 4 cribriform and perineural invasion, intended to assist pathologists
- Chen et al., Towards a general-purpose foundation model for computational pathology (UNI), Nature Medicine, 2024 — pretraining on over 100 million tiles from over 100,000 H&E slides, 20 tissue types, 33 clinical tasks
- Howard et al., The impact of site-specific digital histology signatures on deep learning model accuracy and bias, Nature Communications, 2021 — over 3,000 patients, 6 cancer types; the submitting site is recognised with AUROC 0.964–0.998; 91.1% of features lose AUROC under site-preserved validation
- Ardon et al., Digital pathology operations at a tertiary cancer center: infrastructure requirements and operational cost, Journal of Pathology Informatics, 2023 — 25 scanners, over 6.1 million slides scanned, 2 GB per slide, 9 PB of storage, USD 1 M per PB, USD 0.55–19.53 per scan, 1 FTE per 3–4 scanners
- MDCG 2019-11 rev.1, Qualification and classification of software, Regulation (EU) 2017/745 and Regulation (EU) 2017/746, European Commission, 2025 — text of MDR Rule 11, IVDR implementing rule 1.4, example of cytology screening software in class C
- Regulation (EU) 2024/1689 on artificial intelligence (AI Act), EUR-Lex — Article 6(1) and Annex I points 11–12: AI systems in medical devices and IVDs assessed by a notified body are high-risk
- Regulation (EU) 2026/1744 (Digital Omnibus on AI), EUR-Lex — adopted 8 July 2026, published 24 July 2026; obligations for AI in Annex I products apply from 2 August 2028
- Regulation (EU) 2016/679 (GDPR), EUR-Lex — Article 9 on health data and the scientific-research exception; Recital 26 on pseudonymised and anonymous data
Tell us what you need. We come back with a prototype, not with slides.