Open Access
Open access
volume 2 issue 1 publication number e202100069

Image2SMILES: Transformer‐Based Molecular Optical Recognition Engine**

Ivan Khokhlov 1, 2
Lev Krasnov 1, 2, 3, 4
Maxim V. Fedorov 1, 2, 5, 6, 7, 8
Sergey Sosnin 1, 2, 6, 8
Publication typeJournal Article
Publication date2022-01-11
scimago Q1
wos Q2
SJR1.658
CiteScore11.6
Impact factor3.6
ISSN26289725
Materials Science (miscellaneous)
Abstract

The rise of deep learning in various scientific and technology areas promotes the development of AI‐based tools for information retrieval. Optical recognition of organic structures is a key part of the automated extraction of chemical information. However, this is a challenging task because there is a large variety of representation styles. In this research, we present a Transformer‐based artificial neural network to convert images of organic structures to molecular structures. To train the model, we created a comprehensive data generator that stochastically simulates various drawing styles, functional groups, functional group placeholders (R‐groups), and visual contamination. We demonstrate that the Transformer‐based architecture can gather chemical insights from our generator with almost absolute confidence. That means that, with Transformer, one can fully concentrate on data simulation to build a good recognition model. A web demo of our optical recognition engine is available online at Syntelly platform, and the code for dataset generation is available on GitHub.

Found 
Found 

Top-30

Journals

1
2
3
4
Journal of Cheminformatics
4 publications, 12.5%
Journal of Chemical Information and Modeling
3 publications, 9.38%
Briefings in Bioinformatics
1 publication, 3.13%
Molecular Informatics
1 publication, 3.13%
Lecture Notes in Computer Science
1 publication, 3.13%
28th International Conference on Intelligent User Interfaces
1 publication, 3.13%
npj Computational Materials
1 publication, 3.13%
Nature Communications
1 publication, 3.13%
Macromolecules
1 publication, 3.13%
Energy
1 publication, 3.13%
RSC Advances
1 publication, 3.13%
Complex & Intelligent Systems
1 publication, 3.13%
Scientific Reports
1 publication, 3.13%
Journal of Pharmaceutical Analysis
1 publication, 3.13%
Journal of Physical Chemistry Letters
1 publication, 3.13%
Nature Machine Intelligence
1 publication, 3.13%
Journal of Supercomputing
1 publication, 3.13%
Chemical Reviews
1 publication, 3.13%
Chemical Society Reviews
1 publication, 3.13%
Environmental Science and Technology Letters
1 publication, 3.13%
Plants
1 publication, 3.13%
1
2
3
4

Publishers

2
4
6
8
10
12
Springer Nature
11 publications, 34.38%
American Chemical Society (ACS)
7 publications, 21.88%
Institute of Electrical and Electronics Engineers (IEEE)
4 publications, 12.5%
Elsevier
2 publications, 6.25%
Royal Society of Chemistry (RSC)
2 publications, 6.25%
Oxford University Press
1 publication, 3.13%
Wiley
1 publication, 3.13%
Association for Computing Machinery (ACM)
1 publication, 3.13%
MDPI
1 publication, 3.13%
2
4
6
8
10
12
  • We do not take into account publications without a DOI.
  • Statistics recalculated weekly.

Are you a researcher?

Create a profile to get free access to personal recommendations for colleagues and new articles.
Metrics
32
Share
Cite this
GOST |
Cite this
GOST Copy
Khokhlov I. et al. Image2SMILES: Transformer‐Based Molecular Optical Recognition Engine** // Chemistry - Methods. 2022. Vol. 2. No. 1. e202100069
GOST all authors (up to 50) Copy
Khokhlov I., Krasnov L., Fedorov M. V., Sosnin S. Image2SMILES: Transformer‐Based Molecular Optical Recognition Engine** // Chemistry - Methods. 2022. Vol. 2. No. 1. e202100069
RIS |
Cite this
RIS Copy
TY - JOUR
DO - 10.1002/cmtd.202100069
UR - https://chemistry-europe.onlinelibrary.wiley.com/doi/10.1002/cmtd.202100069
TI - Image2SMILES: Transformer‐Based Molecular Optical Recognition Engine**
T2 - Chemistry - Methods
AU - Khokhlov, Ivan
AU - Krasnov, Lev
AU - Fedorov, Maxim V.
AU - Sosnin, Sergey
PY - 2022
DA - 2022/01/11
PB - Wiley
IS - 1
VL - 2
SN - 2628-9725
ER -
BibTex
Cite this
BibTex (up to 50 authors) Copy
@article{2022_Khokhlov,
author = {Ivan Khokhlov and Lev Krasnov and Maxim V. Fedorov and Sergey Sosnin},
title = {Image2SMILES: Transformer‐Based Molecular Optical Recognition Engine**},
journal = {Chemistry - Methods},
year = {2022},
volume = {2},
publisher = {Wiley},
month = {jan},
url = {https://chemistry-europe.onlinelibrary.wiley.com/doi/10.1002/cmtd.202100069},
number = {1},
pages = {e202100069},
doi = {10.1002/cmtd.202100069}
}