A preprint posted on arXiv proposes a way to improve contrastive pretraining for molecular embeddings by incorporating deep-learning-based morphology profiles. These profiles are derived from cell imaging, which has recently advanced to the point of collecting high-volume morphology data.
The authors argue that such data can serve as a training signal by representing the experimental phenotypic perturbations a molecule induces. This allows molecular embedding models to learn from observed biological effects rather than only from chemical structure or textual data.
The abstract is brief and does not include implementation details or experimental results. As this is the only source, there is no independent comparison to note; the claims should be read as a proposal awaiting full validation.