Life's source code has been written in four letters—A, T, G, C—for four billion years. In 2019, the Benner group added two synthetic base pairs, Z:P and S:B, expanding the alphabet to eight and naming it hachimoji (Japanese for 'eight letters'). Making it was one thing; whether the cell's machinery would read it was another.
The 2019 report wasn't a full success: transcription relied on T7 phage polymerase, and natural T7 cannot read the letter S—an engineered FAL variant was needed to cover all eight.
This time? Still not a perfect score, but a real step up.
Three Steps Up
In 2023, the B:S pair entered E. coli RNA polymerase—six letters. On September 2, 2026, Nature Communications published a collaboration between Dong Wang's group at UCSD with Benner, Lyumkis, and colleagues: unmodified E. coli RNAP (the native α2ββ'ω five-subunit enzyme) transcribed all eight letters, P:Z and B:S included, with four cryo-EM structures from 2.42 to 2.75 Å. First author Qingrong Li; corresponding authors Benner and Wang.
- 2019: Hachimoji DNA born; T7 needed engineering to read all eight
- 2023: B and S pair enters E. coli RNAP—six letters
- 2026-08: Companion paper—hydrophobic base pairs transcribed without hydrogen bonding
- 2026-09: Native RNAP reads all eight letters, with four structures
- Eight-letter mRNA is a prerequisite for future translation of eight-letter codons—translation has not yet been achieved
- Expanded alphabets suit evolutionary aptamers: hachimoji Spinach fluorescent aptamers and AegisBinders are signposts
- Six-letter AEGIS aptamers targeting liver cancer cells have a 2020 preclinical precedent
The structural answer is almost boringly simple: the machine checks synthetic letters using nearly the same signals as natural ones. Watson-Crick geometry passes as usual, the trigger loop folds as usual, the SI3 domain shifts as usual. One new finding: Z's nitro group bridges via a water molecule to interact with the bridge helix's π-hole—water at 2.62 Å.
Kinetics look respectable too. In single-turnover experiments, incorporation of the synthetic pair is only about twofold slower than natural dG:CTP, and after incorporation the polymerase proceeds to the n+2 site without pausing.
But there's honesty in the numbers: with dZ paired to PTP, the trigger loop reaches a closed (catalysis-ready) state 76.8% of the time; with dP paired to Z*TP, only 19.4%. The machine doesn't treat all pairs equally—it has its own preferences.
"Fidelity" in Quotation Marks
Z has a chemical weakness: a pKa of about 7.8 means it can protonate at physiological pH and masquerade as G by pairing with GTP. The team's fix was to keep modifying the letter—Z*, replacing the nitro group with an amide, pushing pKa above 10 and suppressing misincorporation. So the 'high-fidelity eight letters' are no longer the original 2019 alphabet: the letters themselves have been patched.
The paper acknowledges its mismatch rates are overestimates (forced single-substrate conditions). And this is an in vitro reconstituted system: a 9-nucleotide RNA primer on a 27-nucleotide template—not living bacteria, not whole genes. 'Bacteria can read synthetic DNA' headlines are not justified here.
The companion paper is also worth noting: on August 12 in PNAS, the same Wang group showed that even a hydrophobic base pair with no hydrogen bonds at all, Ds:Pa, can be transcribed by this machine—though with clear bias: DsTP incorporates far more smoothly than PaTP.
Where It Leads
The paper draws its own boundaries:
An earlier thread on this site discussed RNA self-assembling through the difference of a single oxygen atom—a chemical self-assembly approach. This work takes the structural biology route: take the machine apart and see why it's willing to read. Both roads lead to the same question: does life's alphabet have only four slots?