Image2SMILES: Transformer‐Based Molecular Optical Recognition Engine**
The rise of deep learning in various scientific and technology areas promotes the development of AI‐based tools for information retrieval. Optical recognition of organic structures is a key part of the automated extraction of chemical information. However, this is a challenging task because there is a large variety of representation styles. In this research, we present a Transformer‐based artificial neural network to convert images of organic structures to molecular structures. To train the model, we created a comprehensive data generator that stochastically simulates various drawing styles, functional groups, functional group placeholders (R‐groups), and visual contamination. We demonstrate that the Transformer‐based architecture can gather chemical insights from our generator with almost absolute confidence. That means that, with Transformer, one can fully concentrate on data simulation to build a good recognition model. A web demo of our optical recognition engine is available online at Syntelly platform, and the code for dataset generation is available on GitHub.
Citations by journals
1
2
|
|
Journal of Cheminformatics
|
Journal of Cheminformatics
2 publications, 12.5%
|
Journal of Chemical Information and Modeling
|
Journal of Chemical Information and Modeling
2 publications, 12.5%
|
Briefings in Bioinformatics
|
Briefings in Bioinformatics
1 publication, 6.25%
|
Molecular Informatics
|
Molecular Informatics
1 publication, 6.25%
|
Lecture Notes in Computer Science
|
Lecture Notes in Computer Science
1 publication, 6.25%
|
28th International Conference on Intelligent User Interfaces
|
28th International Conference on Intelligent User Interfaces
1 publication, 6.25%
|
npj Computational Materials
|
npj Computational Materials
1 publication, 6.25%
|
Nature Communications
|
Nature Communications
1 publication, 6.25%
|
Macromolecules
|
Macromolecules
1 publication, 6.25%
|
Energy
|
Energy
1 publication, 6.25%
|
1
2
|
Citations by publishers
1
2
3
4
5
|
|
Springer Nature
|
Springer Nature
5 publications, 31.25%
|
American Chemical Society (ACS)
|
American Chemical Society (ACS)
3 publications, 18.75%
|
IEEE
|
IEEE
2 publications, 12.5%
|
Oxford University Press
|
Oxford University Press
1 publication, 6.25%
|
Wiley
|
Wiley
1 publication, 6.25%
|
Association for Computing Machinery (ACM)
|
Association for Computing Machinery (ACM)
1 publication, 6.25%
|
Elsevier
|
Elsevier
1 publication, 6.25%
|
1
2
3
4
5
|
- We do not take into account publications that without a DOI.
- Statistics recalculated only for publications connected to researchers, organizations and labs registered on the platform.
- Statistics recalculated weekly.