Redefining TSR scoring with AI: insights from the data scientists – part 2

Discover how AI is revolutionizing TSR scoring, enhancing clinical decisions, and shaping the future of healthcare.

The ARABESC project is the result of a collaborative partnership between WSK Medical and IMP Diagnostics. Since 2023, the two companies have been working together to develop an AI-driven algorithm for tumor-stroma ratio (TSR) scoring.

Building on Part 1, Felix and Cyrine now tackle the “last mile” problems—embedding ARABESC within busy diagnostic labs, aligning with IVDR and FDA pathways, and measuring real-world value beyond pure accuracy.

COMBINING THE FIELDS OF TECHNOLOGY AND MEDICINE

I – How could this tool be integrated into existing digital pathology workflows without disrupting diagnostic timelines?

FD – The final product is a pipeline that performs preprocessing, prediction and postprocessing on a snapshot of a whole slide image. This pipeline is deployed on-premise in a hospital or pathology lab or using cloud services. Many viewers for digital pathology slides have their own way of evoking models within their viewer environment. By adjusting the endpoints of the deployment to be compatible with the specific input and output format of these viewers, the model can be used by pathologists as easy as a right mouse click action.

I – How would you structure collaboration between AI developers, pathologists, and oncologists during tool refinement?

CF – Collaboration can be structured as a continuous feedback loop. Pathologists could provide their feedback and whether the tool’s output aligns with therapeutic decision-making needs. Besides, regular review meetings can be held to discuss model performance on edge cases, identify clinical priorities, and guide iterative updates.

I – What training data requirements (e.g., annotated regions, clinical outcomes) would you specify for pathologist collaborators?

FD – Especially with the rise of foundation models, the focus of data requirements has shifted from “big data” to “quality data”. We have found a way to ensure that our training data is of the highest quality, by asking our pathologist collaborators to focus on a relatively small area when annotating but doing this in the highest detail. These annotations are than passed through a set of quality assessment steps and reannotated, after which they are used for training.

CF – High quality annotations from pathologists, including clearly delineated tumour and stroma regions on representative H&E slides, is a key component to train a model effectively.

I – How would you communicate the tool’s limitations to non-technical stakeholders in clinical settings?

FD – The user needs to understand that the tool is never 100% accurate even in production. These errors are normal behaviour of any AI model of the current age. The tools score is as traceable as possible and achieves this by displaying a full segmentation map. The user should realize that this map is not there to convince it to trust the score, but instead to highlight any errors that the model might have made. However, it is important that the transparency of these errors does not convince the user to disregard the model.

KEY CONSIDERATIONS FOR IMPLEMENTING AI IN CLINICAL PRACTICE

I – What regulatory considerations (e.g., FDA/CE-IVDR compliance) are critical for clinical deployment of this tool?

CF – Regulatory compliance is essential for clinical deployment, and it is crucial that our tool meets the requirements of the IVDR. For successful deployment, the tool must demonstrate its scientific validity and clinical performance, while also ensuring a sufficient level of explainability so that end users can interpret and trust the results as part of the diagnostic process.

FD – For CE-IVDR we thoroughly test the functional requirements of the model, by validating it to specific benchmarks as well as performing an extensive literature review to research the clinical relevance and useability of the TSR itself. Besides this, the safety requirements are tested with using a risk analysis, that is updated with each step in development, as each added functionality results in an added risk. To comply with CE-IVDR each of these risks will be met with an appropriate mitigation.

I – How would you ensure the algorithm remains adaptable to new colorectal cancer grading systems or emerging biomarkers?

FD – As the model in its core is a tissue segmenter it is highly adaptable to other biomarkers such using its ability to identify lymphatic tissue for lymph nodes, small clusters of tumour as budding detector, its ability to segment stroma as input for stromal configuration detection, and its ability to identify malignant tissue in general to possibly draw the attention to lymphatic or vascular invasion.

I – What KPIs would you track to demonstrate the tool’s clinical utility beyond technical accuracy (e.g., workflow efficiency gains)?

FD – Important KPIs would be bias KPIs that identify unexplainable model behaviour in specific tumour subtypes and the rate of adoptability. Both could be tracked using logging which counts the number of users of the model and the number of uses over time in general and per user.

CF – I think it’s important to track KPIs that reflect the tool’s impact on real-world clinical practice such as the time saved per case or per day when using the tool compared to manual TSR assessment or the improvement in consistency of stroma-high vs. stroma-low classification among pathologists. A survey on the tool’s perceived usefulness and user trust can be conducted.

AI is steadily reshaping how we approach TSR scoring, offering new levels of consistency, transparency, and clinical support. By addressing current challenges with care and collaboration, we move closer to tools that are not only powerful but practical. The future of TSR scoring is not just automated – it’s augmented, collaborative, and clinically meaningful.

Related posts