A leading life sciences company improves cancer variant prediction by 21% with ESM models

Industries

Life Sciences,
Cancer Research

Tech & TOOLS

Meta ESM1b, ESM2, COSMIC, ClinVar, AlphaMissense, SageMaker, MLFlow, AWS CDK, LoRA

Teams & Services

AI & ML, Data Engineering,
Cloud Engineering, MLOps

milestones

21% higher PR-AUC than zero-shot methods, 96% precision at 20% recall, and 27% more data available for inference

Industry

Life Sciences,
Cancer Research

Tech & TOOLS

Meta ESM1b, ESM2, COSMIC, ClinVar, AlphaMissense, SageMaker, MLFlow, AWS CDK, LoRA

Teams & Services

AI & ML, Data Engineering,
Cloud Engineering, MLOps

milestones

21% higher PR-AUC than zero-shot methods, 96% precision at 20% recall, and 27% more data available for inference

The company partnered with Loka to improve Variant Effect Prediction for cancer research, achieving 21% higher PR-AUC than zero-shot methods and 96% precision at 20% recall.
‍

‍

The situation

A life sciences company focused on advancing cancer treatments wanted to improve Variant Effect Prediction (VEP) for cancer research by developing a model trained specifically on cancer-related data.

The company partnered with Loka to enhance an ESM-based approach using the COSMIC cancer dataset. The goal was to improve prediction accuracy beyond zero-shot models and establish machine learning infrastructure for continued model development.

The challenge

Existing ESM models generate powerful protein-sequence representations, but zero-shot approaches are not tailored to predict cancer-related variants.

The company needed to determine whether fine-tuning ESM models on cancer-specific data could improve prediction performance. The work also required processing protein sequences beyond ESM's context window and establishing repeatable pipelines for data preparation, model training, and inference.

The solution

Loka enhanced the company's ESM-based workflow by building AWS-native pipelines for preprocessing, training, and inference. The team trained classifiers using ESM embeddings and the COSMIC cancer dataset, evaluating approaches including Support Vector Classifiers, MLPs, and XGBoost.

Loka also engineered an approach for processing sequences beyond ESM's context window and evaluated fine-tuning with a Siamese network and Parameter Efficient Fine Tuning (PEFT) methods, including LoRA. External datasets were also evaluated by Loka for pre-finetuning.

SageMaker, MLFlow, and AWS CDK provided the infrastructure for repeatable model development and evaluation.

What we delivered

Cancer-specific prediction
Trained classifiers using ESM embeddings and the COSMIC cancer dataset to improve cancer-specific variant prediction.

Long-sequence processing

Engineered a solution to process protein sequences beyond ESM's context window.

Model evaluation
Compared zero-shot and fine-tuned approaches across cancer-specific and generalized datasets.

ML infrastructure
Built repeatable AWS pipelines for data preparation, training, evaluation, and inference using SageMaker, MLFlow, and AWS CDK.

Fine-tuning foundation
Evaluated LoRA and other PEFT approaches for continued model development.

The results

Loka helped the company improve cancer-specific VEP while establishing an infrastructure foundation for continued model development.

21% higher PR-AUC: The ESM1b + SVC model achieved a PR-AUC of 0.63 on the COSMIC dataset, a 21% improvement over the zero-shot approach.

96% precision: Achieved at 20% recall, a level relevant to the company's production operating regime.

26% relative improvement in precision: Increased precision at 20% recall compared with the zero-shot approach.

27% more data: Enabled inference on additional data by addressing ESM's sequence-length limitation.

100% of COSMIC data: Supported training with the full cancer dataset.