Let me put it bluntly: in 2026, if your AI drug-discovery model still relies purely on 'reading text literature' to guess molecular structures, the tens of millions you burned on GPUs are essentially wasted.
The Bolek model from Poland's Ingenix.ai team acts like a scalpel, cutting into a festering sore in today's AI for Science circle: we thought we had trained a group of master physicians, only to find they were bookworms who memorized the medical canon but can't even read a thermometer. 🤒💊
Why does your LLM spout nonsense on critical data?
Today's tech giants suffer from a fatal 'text fundamentalism.' They feed tons of journal papers into LLMs. Ask one about a new molecule's toxicity, and it will write two pages of pharmacological analysis with a perfectly rigorous chain of thought. But pull back the curtain on its logic and you'll find: the polar surface area and lipophilicity coefficients it cites were all guessed based on contextual probability!
The uncomfortable truth: when it comes to rigorous molecular reasoning, a pure-text LLM is a fortune teller prescribing drugs based on your star sign. 🔮📉
> Notes: > * \(\text{Morgan\_Fingerprint}\): the standard chemistry encoding that 'flattens' a 3D molecular structure into a numeric vector. > * Meaning of the formula: rather than hoping the AI spontaneously develops chemical common sense, standard 'lab reports' are forcibly translated into the model's native language and injected as a physical external module.
Bolek's brutally pragmatic approach
Bolek has no patience for the romantic notion that everything is text. It has only 4B parameters (built on Qwen3), but the researchers set an iron rule during training: 'Before giving any pharmacological conclusion, you must first report the underlying values from the structural fingerprint verbatim!' 🏗️
The results are a massacre. Across 15 real drug property prediction tasks, this tiny model with its 'numerical IV drip' crushed the chemistry-tuned 9B giant (TxGemma-9B). Bolek cites underlying ground-truth values 10–100x more frequently than the scripture-reciting large models, achieving a correlation coefficient of 0.91 with values computed by professional chemistry software.
My bet
Architects still hoping that expanding context windows and stuffing in more natural-language corpora will fix scientific computing hallucinations are walking into a dead end. The foundation of science is physical measurement, not word games.
If you disagree, keep grinding on your pure-text LLM. But next year, when competitors use a micro model mounted with physics-grounded inputs at 1/10 the cost to lock down a real antibody, while you're wasting lab reagents on AI-invented 'dream targets,' don't say nobody rang the bell for you in 2026. 🤝
Let the numbers count where numbers should count—stop letting AI do science with rhetoric. 🎙️🔥
---
Paper Information
- Title: Bolek: A Multimodal Language Model for Molecular Reasoning
- Authors: Frederic Grabowski, Tomasz Jetka, et al.
- Institutions: Ingenix.ai, Warsaw University of Technology
- Published: 2026-05-04
- Categories: cs.LG, cs.AI