Chemistry line notations from 1949 to today
How chemists write a molecule as one line of text, from WLN and SMILES to InChI and the notations made for machine learning, with caffeine as the running example.
A line notation writes a molecule, which is a graph of atoms and bonds, as a single line of text. Chemists have devised several since 1949, each to fix a weakness of the ones before, and most are still in use. This post takes them in order and writes caffeine in each.
One molecule, many strings#
To write the graph as text, a program walks through it and records the atoms and bonds it meets. Another starting atom or branch order gives another string, so a plain text comparison misses records that are chemically identical. The strings below do one of three jobs: a notation describes one molecule, an identifier names it so that independent programs produce the same string, and a pattern describes a set of molecules. Patterns such as SMARTS are covered in Reactions and retrosynthesis.
Before SMILES#
WLN, 1949#
William Wiswesser designed his line notation for the typewriter: letters and digits encode rings, functional groups and substituents in a fixed order. The Institute for Scientific Information kept its registry of new compounds in WLN from 1968 to 1987, and pharmaceutical companies used it for their internal files.1 An open-source parser now converts the WLN records that remain in ChEMBL, ChemSpider and PubChem.2
CAS Registry Number, 1965#
Chemical Abstracts Service assigns every registered substance a number; caffeine is 58-08-2. The number says nothing about the structure. Its last digit is a check digit: weight the other digits 1, 2, 3 and so on from the right and keep the last digit of the sum (8×1 + 0×2 + 8×3 + 5×4 = 52). Bottles, safety data sheets and regulatory filings all cite it.
IUPAC names, 1979 onwards#
The systematic name of caffeine is 1,3,7-trimethyl-3,7-dihydro-1H-purine-2,6-dione. A name can be read aloud and checked by hand, but the rules, collected in the Blue Book, run to more than 1,500 pages, so generating names by program is hard. Commercial generators exist, and STOUT is an open-source neural model that translates SMILES into names; its output needs checking.3
SMILES and its dialects#
SMILES, 1988#
David Weininger published SMILES in 1988, from the US Environmental Protection Agency in Duluth, as a notation a chemist could type on one line.4
CN1C=NC2=C1C(=O)N(C(=O)N2C)CThe string records a depth-first walk through the graph. Common organic atoms are bare element symbols, other atoms go in square brackets, and hydrogens follow from normal valences. Single and aromatic bonds are implied, = and # mark double and triple bonds, matching digits close rings, parentheses hold branches, and @, / and \ record stereochemistry. These 28 characters describe all 24 atoms of C8H10N4O2.
Canonical SMILES, 1989#
Starting the walk elsewhere gives CN1C(=O)N(C)c2ncn(C)c2C1=O, which is also caffeine. To compare strings, each must first be written in a canonical atom order. Morgan's 1965 algorithm ranks atoms by extended connectivity,5 and Weininger published an algorithm for unique SMILES in 1989.6 Every toolkit implements its own:
RDKit 2026.03 Cn1c(=O)c2c(ncn2C)n(C)c1=O
PubChem (OEChem) CN1C=NC2=C1C(=O)N(C(=O)N2C)C
Open Babel 3.1 Cn1cnc2c1c(=O)n(C)c(=O)n2C
Indigo 1.46 CN1C(=O)C2=C(N=CN2C)N(C)C1=OAll four give the same InChIKey, and no two are identical. A canonical SMILES is a reliable key for duplicates within one toolkit, not between toolkits. The OpenSMILES specification, started by Craig James in 2007, documents the grammar that parsers actually accept.7
Balsa, 2022#
Richard Apodaca's Balsa keeps the SMILES character set but defines the grammar formally.8 It has no aromatic bond symbol, so aromaticity lives only in lowercase atoms, and the atoms allowed without brackets are a closed list. The aim is a notation that independent parsers read identically.
InChI and InChIKey#
The IUPAC International Chemical Identifier (2005) has the opposite aim to SMILES: not a string that is easy to write, but one that every implementation computes identically. IUPAC and NIST developed it, and the InChI Trust maintains one reference implementation.9
InChI=1S/C8H10N4O2/c1-10-4-9-6-5(10)7(13)12(3)8(14)11(6)2/h4H,1-3H3Slashes separate the layers: formula, connectivity (c), hydrogens (h), then charge, stereochemistry and isotopes where needed. The InChIKey (2007) is a 27-character hash of the InChI for indexes and web searches; caffeine's is RYYVLZVUVIJVGH-UHFFFAOYSA-N. The first 14 characters encode the connectivity, the next block the other layers, ending in S for standard and A for version 1, and the last character the protonation.10 Tests on very large sets found hash collisions at the rate theory predicts.11
Notations for machine learning#
Generative models write strings token by token, and in SMILES one misplaced parenthesis or ring digit invalidates the whole string. A large fraction of generated SMILES are not valid molecules.12
DeepSMILES, 2018#
Noel O'Boyle and Andrew Dalke removed the paired symbols: a branch closes with parentheses whose number gives its length, and a ring closes with one digit that gives its size.13 The RDKit canonical SMILES of caffeine becomes Cnc=O)ccncn5C))))nC)c6=O.
SELFIES, 2020#
In SELFIES every string decodes to a valid molecule.12 Branch and ring tokens read the next token as a length, so nothing can be left open, and the decoder respects valences.
[C][N][C][=N][C][=C][Ring1][Branch1][C][=Branch1][C][=O][N][Branch1][=Branch2][C][=Branch1][C][=O][N][Ring1][Branch2][C][C]Group SELFIES and SAFE#
Group SELFIES (2023) uses whole fragments, such as a benzene ring, as tokens, which shortens the strings.14 SAFE (2024) writes a molecule as fragments joined by attachment-point digits. The result remains compatible with SMILES, and a model can keep a scaffold fixed while it generates the side chains.15
Side by side#
| Notation | Year | Kind | What it adds |
|---|---|---|---|
| WLN | 1949 | Notation | A structure code for the typewriter |
| CAS Registry Number | 1965 | Identifier | The number on bottles and safety data sheets |
| IUPAC name | 1979 | Identifier | Readable and checkable by hand |
| SMILES | 1988 | Notation | A short string anyone can type |
| Canonical SMILES | 1989 | Notation | One string per molecule within a toolkit |
| InChI | 2005 | Identifier | The same string in every program |
| InChIKey | 2007 | Identifier | A fixed-length hash |
| DeepSMILES | 2018 | Notation | No paired symbols |
| SELFIES | 2020 | Notation | Every string is a valid molecule |
| Balsa | 2022 | Notation | A formal SMILES grammar |
| Group SELFIES | 2023 | Notation | Fragment tokens |
| SAFE | 2024 | Notation | Fragments, still compatible with SMILES |
On this site#
The SMILES to structure converter draws any SMILES as a skeletal formula, Covalent pastes and copies SMILES as you draw, and every compound in the molecule catalogue lists its CAS number, SMILES, InChI and InChIKey. File formats, spectra and reactions are covered in Cheminformatics file formats, Spectroscopy file formats and Reactions and retrosynthesis.
References#
- Garfield E. From laboratory to information explosions: the evolution of chemical information services at ISI. J. Inf. Sci. 27, 119–125 (2001). doi:10.1177/016555150102700208
- Blakey M et al. Zombie cheminformatics: extraction and conversion of Wiswesser Line Notation (WLN) from chemical documents. J. Cheminform. 16, 42 (2024). doi:10.1186/s13321-024-00831-2
- Rajan K et al. STOUT: SMILES to IUPAC names using neural machine translation. J. Cheminform. 13, 34 (2021). doi:10.1186/s13321-021-00512-4
- Weininger D. SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. J. Chem. Inf. Comput. Sci. 28, 31–36 (1988). doi:10.1021/ci00057a005
- Morgan HL. The generation of a unique machine description for chemical structures: a technique developed at Chemical Abstracts Service. J. Chem. Doc. 5, 107–113 (1965). doi:10.1021/c160017a018
- Weininger D et al. SMILES. 2. Algorithm for generation of unique SMILES notation. J. Chem. Inf. Comput. Sci. 29, 97–101 (1989). doi:10.1021/ci00062a008
- James CA et al. OpenSMILES specification, version 1.0 (2016). opensmiles.org
- Apodaca RL. Balsa: a compact line notation based on SMILES. ChemRxiv (2022). doi:10.26434/chemrxiv-2022-01ltp
- Heller SR et al. InChI, the IUPAC International Chemical Identifier. J. Cheminform. 7, 23 (2015). doi:10.1186/s13321-015-0068-4
- InChI Trust. InChI Technical Manual. inchi-trust.org
- Pletnev I et al. InChIKey collision resistance: an experimental testing. J. Cheminform. 4, 39 (2012). doi:10.1186/1758-2946-4-39
- Krenn M et al. Self-referencing embedded strings (SELFIES): a 100% robust molecular string representation. Mach. Learn.: Sci. Technol. 1, 045024 (2020). doi:10.1088/2632-2153/aba947
- O'Boyle N, Dalke A. DeepSMILES: an adaptation of SMILES for use in machine-learning of chemical structures. ChemRxiv (2018). doi:10.26434/chemrxiv.7097960.v1
- Cheng AH et al. Group SELFIES: a robust fragment-based molecular string representation. Digital Discovery 2, 748–758 (2023). doi:10.1039/D3DD00012E
- Noutahi E et al. Gotta be SAFE: a new framework for molecular design. Digital Discovery 3, 796–804 (2024). doi:10.1039/D4DD00019F