All posts

Chemistry line notations from 1949 to today

How chemists write a molecule as one line of text, from WLN and SMILES to InChI and the notations made for machine learning, with caffeine as the running example.

A line notation writes a molecule, which is a graph of atoms and bonds, as a single line of text. Chemists have devised several since 1949, each to fix a weakness of the ones before, and most are still in use. This post takes them in order and writes caffeine in each.

One molecule, many strings#

To write the graph as text, a program walks through it and records the atoms and bonds it meets. Another starting atom or branch order gives another string, so a plain text comparison misses records that are chemically identical. The strings below do one of three jobs: a notation describes one molecule, an identifier names it so that independent programs produce the same string, and a pattern describes a set of molecules. Patterns such as SMARTS are covered in Reactions and retrosynthesis.

Before SMILES#

WLN, 1949#

William Wiswesser designed his line notation for the typewriter: letters and digits encode rings, functional groups and substituents in a fixed order. The Institute for Scientific Information kept its registry of new compounds in WLN from 1968 to 1987, and pharmaceutical companies used it for their internal files.1 An open-source parser now converts the WLN records that remain in ChEMBL, ChemSpider and PubChem.2

CAS Registry Number, 1965#

Chemical Abstracts Service assigns every registered substance a number; caffeine is 58-08-2. The number says nothing about the structure. Its last digit is a check digit: weight the other digits 1, 2, 3 and so on from the right and keep the last digit of the sum (8×1 + 0×2 + 8×3 + 5×4 = 52). Bottles, safety data sheets and regulatory filings all cite it.

IUPAC names, 1979 onwards#

The systematic name of caffeine is 1,3,7-trimethyl-3,7-dihydro-1H-purine-2,6-dione. A name can be read aloud and checked by hand, but the rules, collected in the Blue Book, run to more than 1,500 pages, so generating names by program is hard. Commercial generators exist, and STOUT is an open-source neural model that translates SMILES into names; its output needs checking.3

SMILES and its dialects#

SMILES, 1988#

David Weininger published SMILES in 1988, from the US Environmental Protection Agency in Duluth, as a notation a chemist could type on one line.4

Caffeine in SMILES (PubChem)
CN1C=NC2=C1C(=O)N(C(=O)N2C)C

The string records a depth-first walk through the graph. Common organic atoms are bare element symbols, other atoms go in square brackets, and hydrogens follow from normal valences. Single and aromatic bonds are implied, = and # mark double and triple bonds, matching digits close rings, parentheses hold branches, and @, / and \ record stereochemistry. These 28 characters describe all 24 atoms of C8H10N4O2.

Canonical SMILES, 1989#

Starting the walk elsewhere gives CN1C(=O)N(C)c2ncn(C)c2C1=O, which is also caffeine. To compare strings, each must first be written in a canonical atom order. Morgan's 1965 algorithm ranks atoms by extended connectivity,5 and Weininger published an algorithm for unique SMILES in 1989.6 Every toolkit implements its own:

Caffeine, four canonical SMILES
RDKit 2026.03      Cn1c(=O)c2c(ncn2C)n(C)c1=O
PubChem (OEChem)   CN1C=NC2=C1C(=O)N(C(=O)N2C)C
Open Babel 3.1     Cn1cnc2c1c(=O)n(C)c(=O)n2C
Indigo 1.46        CN1C(=O)C2=C(N=CN2C)N(C)C1=O

All four give the same InChIKey, and no two are identical. A canonical SMILES is a reliable key for duplicates within one toolkit, not between toolkits. The OpenSMILES specification, started by Craig James in 2007, documents the grammar that parsers actually accept.7

Balsa, 2022#

Richard Apodaca's Balsa keeps the SMILES character set but defines the grammar formally.8 It has no aromatic bond symbol, so aromaticity lives only in lowercase atoms, and the atoms allowed without brackets are a closed list. The aim is a notation that independent parsers read identically.

InChI and InChIKey#

The IUPAC International Chemical Identifier (2005) has the opposite aim to SMILES: not a string that is easy to write, but one that every implementation computes identically. IUPAC and NIST developed it, and the InChI Trust maintains one reference implementation.9

Caffeine, standard InChI
InChI=1S/C8H10N4O2/c1-10-4-9-6-5(10)7(13)12(3)8(14)11(6)2/h4H,1-3H3

Slashes separate the layers: formula, connectivity (c), hydrogens (h), then charge, stereochemistry and isotopes where needed. The InChIKey (2007) is a 27-character hash of the InChI for indexes and web searches; caffeine's is RYYVLZVUVIJVGH-UHFFFAOYSA-N. The first 14 characters encode the connectivity, the next block the other layers, ending in S for standard and A for version 1, and the last character the protonation.10 Tests on very large sets found hash collisions at the rate theory predicts.11

Notations for machine learning#

Generative models write strings token by token, and in SMILES one misplaced parenthesis or ring digit invalidates the whole string. A large fraction of generated SMILES are not valid molecules.12

DeepSMILES, 2018#

Noel O'Boyle and Andrew Dalke removed the paired symbols: a branch closes with parentheses whose number gives its length, and a ring closes with one digit that gives its size.13 The RDKit canonical SMILES of caffeine becomes Cnc=O)ccncn5C))))nC)c6=O.

SELFIES, 2020#

In SELFIES every string decodes to a valid molecule.12 Branch and ring tokens read the next token as a length, so nothing can be left open, and the decoder respects valences.

SELFIES of the PubChem SMILES
[C][N][C][=N][C][=C][Ring1][Branch1][C][=Branch1][C][=O][N][Branch1][=Branch2][C][=Branch1][C][=O][N][Ring1][Branch2][C][C]

Group SELFIES and SAFE#

Group SELFIES (2023) uses whole fragments, such as a benzene ring, as tokens, which shortens the strings.14 SAFE (2024) writes a molecule as fragments joined by attachment-point digits. The result remains compatible with SMILES, and a model can keep a scaffold fixed while it generates the side chains.15

Side by side#

NotationYearKindWhat it adds
WLN1949NotationA structure code for the typewriter
CAS Registry Number1965IdentifierThe number on bottles and safety data sheets
IUPAC name1979IdentifierReadable and checkable by hand
SMILES1988NotationA short string anyone can type
Canonical SMILES1989NotationOne string per molecule within a toolkit
InChI2005IdentifierThe same string in every program
InChIKey2007IdentifierA fixed-length hash
DeepSMILES2018NotationNo paired symbols
SELFIES2020NotationEvery string is a valid molecule
Balsa2022NotationA formal SMILES grammar
Group SELFIES2023NotationFragment tokens
SAFE2024NotationFragments, still compatible with SMILES

On this site#

The SMILES to structure converter draws any SMILES as a skeletal formula, Covalent pastes and copies SMILES as you draw, and every compound in the molecule catalogue lists its CAS number, SMILES, InChI and InChIKey. File formats, spectra and reactions are covered in Cheminformatics file formats, Spectroscopy file formats and Reactions and retrosynthesis.

References#

  1. Garfield E. From laboratory to information explosions: the evolution of chemical information services at ISI. J. Inf. Sci. 27, 119–125 (2001). doi:10.1177/016555150102700208
  2. Blakey M et al. Zombie cheminformatics: extraction and conversion of Wiswesser Line Notation (WLN) from chemical documents. J. Cheminform. 16, 42 (2024). doi:10.1186/s13321-024-00831-2
  3. Rajan K et al. STOUT: SMILES to IUPAC names using neural machine translation. J. Cheminform. 13, 34 (2021). doi:10.1186/s13321-021-00512-4
  4. Weininger D. SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. J. Chem. Inf. Comput. Sci. 28, 31–36 (1988). doi:10.1021/ci00057a005
  5. Morgan HL. The generation of a unique machine description for chemical structures: a technique developed at Chemical Abstracts Service. J. Chem. Doc. 5, 107–113 (1965). doi:10.1021/c160017a018
  6. Weininger D et al. SMILES. 2. Algorithm for generation of unique SMILES notation. J. Chem. Inf. Comput. Sci. 29, 97–101 (1989). doi:10.1021/ci00062a008
  7. James CA et al. OpenSMILES specification, version 1.0 (2016). opensmiles.org
  8. Apodaca RL. Balsa: a compact line notation based on SMILES. ChemRxiv (2022). doi:10.26434/chemrxiv-2022-01ltp
  9. Heller SR et al. InChI, the IUPAC International Chemical Identifier. J. Cheminform. 7, 23 (2015). doi:10.1186/s13321-015-0068-4
  10. InChI Trust. InChI Technical Manual. inchi-trust.org
  11. Pletnev I et al. InChIKey collision resistance: an experimental testing. J. Cheminform. 4, 39 (2012). doi:10.1186/1758-2946-4-39
  12. Krenn M et al. Self-referencing embedded strings (SELFIES): a 100% robust molecular string representation. Mach. Learn.: Sci. Technol. 1, 045024 (2020). doi:10.1088/2632-2153/aba947
  13. O'Boyle N, Dalke A. DeepSMILES: an adaptation of SMILES for use in machine-learning of chemical structures. ChemRxiv (2018). doi:10.26434/chemrxiv.7097960.v1
  14. Cheng AH et al. Group SELFIES: a robust fragment-based molecular string representation. Digital Discovery 2, 748–758 (2023). doi:10.1039/D3DD00012E
  15. Noutahi E et al. Gotta be SAFE: a new framework for molecular design. Digital Discovery 3, 796–804 (2024). doi:10.1039/D4DD00019F