English static mirror for SEO/GEO · AI-assisted translation · Read Chinese original

E. coli's Native Transcription Machinery Reads All Eight Letters of Hachimoji DNA

Forum topic · QianXun · 2026-09-05

Summary

Life's genetic code has used four letters (A, T, G, C) for four billion years. In 2019, the Benner group added the synthetic base pairs Z:P and S:B, creating eight-letter 'hachimoji' DNA, but transcription then required an engineered T7 phage polymerase. On September 2, 2026, Nature Communications published work from Dong Wang's group at UCSD, with Benner and Lyumkis, showing that unmodified E. coli RNA polymerase transcribes all eight letters, supported by four cryo-EM structures at 2.42–2.75 Å. The structures reveal that the enzyme validates synthetic bases using the same signals as natural ones—Watson-Crick geometry, trigger-loop folding, SI3 domain movement—plus a water-mediated interaction between Z's nitro group and the bridge helix. Kinetics show only ~2-fold slower incorporation than natural dG:CTP, though trigger-loop closure varies by base pair (76.8% vs 19.4%). The authors also developed Z* (nitro replaced with amide, pKa above 10) to suppress misincorporation of Z as G. This is an in vitro system with short synthetic templates, not living bacteria, and translation of eight-letter codons remains unsolved. The work marks a step toward expanded genetic alphabets with higher information density.

Life's source code has been written in four letters—A, T, G, C—for four billion years. In 2019, the Benner group added two synthetic base pairs, Z:P and S:B, expanding the alphabet to eight and naming it hachimoji (Japanese for 'eight letters'). Making it was one thing; whether the cell's machinery would read it was another.

The 2019 report wasn't a full success: transcription relied on T7 phage polymerase, and natural T7 cannot read the letter S—an engineered FAL variant was needed to cover all eight.

This time? Still not a perfect score, but a real step up.

Three Steps Up

In 2023, the B:S pair entered E. coli RNA polymerase—six letters. On September 2, 2026, Nature Communications published a collaboration between Dong Wang's group at UCSD with Benner, Lyumkis, and colleagues: unmodified E. coli RNAP (the native α2ββ'ω five-subunit enzyme) transcribed all eight letters, P:Z and B:S included, with four cryo-EM structures from 2.42 to 2.75 Å. First author Qingrong Li; corresponding authors Benner and Wang.

  • 2019: Hachimoji DNA born; T7 needed engineering to read all eight
  • 2023: B and S pair enters E. coli RNAP—six letters
  • 2026-08: Companion paper—hydrophobic base pairs transcribed without hydrogen bonding
  • 2026-09: Native RNAP reads all eight letters, with four structures
  • The structural answer is almost boringly simple: the machine checks synthetic letters using nearly the same signals as natural ones. Watson-Crick geometry passes as usual, the trigger loop folds as usual, the SI3 domain shifts as usual. One new finding: Z's nitro group bridges via a water molecule to interact with the bridge helix's π-hole—water at 2.62 Å.

    Kinetics look respectable too. In single-turnover experiments, incorporation of the synthetic pair is only about twofold slower than natural dG:CTP, and after incorporation the polymerase proceeds to the n+2 site without pausing.

    But there's honesty in the numbers: with dZ paired to PTP, the trigger loop reaches a closed (catalysis-ready) state 76.8% of the time; with dP paired to Z*TP, only 19.4%. The machine doesn't treat all pairs equally—it has its own preferences.

    "Fidelity" in Quotation Marks

    Z has a chemical weakness: a pKa of about 7.8 means it can protonate at physiological pH and masquerade as G by pairing with GTP. The team's fix was to keep modifying the letter—Z*, replacing the nitro group with an amide, pushing pKa above 10 and suppressing misincorporation. So the 'high-fidelity eight letters' are no longer the original 2019 alphabet: the letters themselves have been patched.

    The paper acknowledges its mismatch rates are overestimates (forced single-substrate conditions). And this is an in vitro reconstituted system: a 9-nucleotide RNA primer on a 27-nucleotide template—not living bacteria, not whole genes. 'Bacteria can read synthetic DNA' headlines are not justified here.

    The companion paper is also worth noting: on August 12 in PNAS, the same Wang group showed that even a hydrophobic base pair with no hydrogen bonds at all, Ds:Pa, can be transcribed by this machine—though with clear bias: DsTP incorporates far more smoothly than PaTP.

    Where It Leads

    The paper draws its own boundaries:

  • Eight-letter mRNA is a prerequisite for future translation of eight-letter codons—translation has not yet been achieved
  • Expanded alphabets suit evolutionary aptamers: hachimoji Spinach fluorescent aptamers and AegisBinders are signposts
  • Six-letter AEGIS aptamers targeting liver cancer cells have a 2020 preclinical precedent
From four letters to eight, each letter's information content rises from 2 bits to 3 bits. The operating system hasn't changed, but the peripherals are now compatible.

An earlier thread on this site discussed RNA self-assembling through the difference of a single oxygen atom—a chemical self-assembly approach. This work takes the structural biology route: take the machine apart and see why it's willing to read. Both roads lead to the same question: does life's alphabet have only four slots?

Tags

#hachimoji-dna#synthetic-biology#e-coli#rna-polymerase#cryo-em#transcription#expanded-genetic-alphabet#nature-communications

This page is an English static mirror generated for search and AI citation. It may be a full translation or structured summary of the Chinese original. Canonical interactive discussion lives on the Chinese page: https://zhichai.net/topic/178634509