Skip to main content

Neural nets screen bio-based monomers for latex — before the first lab batch

Ilustrační obrázek
Screening a new bio-based monomer normally means weeks of synthesis before you find out whether it copolymerises the way your formula needs. A newly reported machine-learning framework reverses that order: it predicts the four properties that decide whether a monomer is usable — propagation rate constant (kp), reactivity ratios (rr), glass transition temperature (Tg) and water solubility (Ws) — simultaneously, and it does so for pairs of monomers rather than one molecule at a time.

That last detail is the one worth pausing on. Emulsion polymerisation for latexes is not a single-molecule problem. What you get out of the reactor depends on who reacts with whom, and that is governed by the two reactivity ratios of a monomer pair — the Mayo–Lewis relationship that every coatings chemist has drawn on a whiteboard at least once. A tool that only scores individual monomers can tell you a molecule is "green" and "available". It cannot tell you whether it will actually build the copolymer you designed.

Four properties, four different failure modes

It is worth spelling out why these four outputs were chosen, because each one kills a different project:

kp — the propagation rate constant — decides how fast radicals add the next monomer unit. Too slow and your conversion per reactor hour collapses; you either accept a longer batch cycle or you accept residual monomer.

Reactivity ratios control sequence distribution. A pair with badly mismatched ratios gives you compositional drift: a gradient copolymer that behaves nothing like the statistical copolymer you drew in the lab notebook, with knock-on effects on adhesion and weatherability.

Tg sets the minimum film formation temperature and blocking resistance — the two properties your customer actually notices.

Ws, water solubility, is the emulsion chemist's quiet killer. A monomer that partitions too readily into the aqueous phase feeds secondary nucleation, produces coagulum, and ends up as something other than latex particles.

According to the research summarised by European Coatings, the framework combines neural networks with gradient boosting, uses an enhanced network architecture specifically for reactivity ratios, and covers a broader range of monomer families than earlier predictive tools. On water solubility it is reported to outperform the previous state of the art, SolTranNet. Validation went beyond the usual benchmark tables: the selected bio-based monomer systems were actually polymerised in coating and adhesive formulations, reaching high conversion with Tg values equal to their petroleum-based counterparts.

The number that explains the interest

Bio-based polymers currently account for roughly 4.2 million tonnes per year — about 1% of global polymer production. The projection cited is 4% to 5% of total polymer production by 2035.

Run the arithmetic and the picture gets less triumphant. If 4.2 million tonnes is 1%, total polymer output sits somewhere near 420 million tonnes a year. Hitting 5% means bio-based volume of roughly 21 million tonnes — call it a five-fold increase in under a decade. That is a real industrial ramp, not a niche. It is also far too slow to absorb the cost of screening every candidate monomer pair the way we do today, one 500 ml reactor at a time. This is precisely the gap in-silico screening is meant to fill.

Why this lands harder in Europe than anywhere else

Europe's coatings and adhesives sector — AkzoNobel, BASF, Allnex, Beckers, Tikkurila and a long tail of mid-sized formulators — is being pushed toward bio-based raw materials by two things at once: the EU's Chemicals Strategy for Sustainability and its safe-and-sustainable-by-design framing, and plain REACH pressure on the fossil-derived monomers that are easiest to reach for.

The European constraint is not ambition. It is documentation. A formulator here cannot simply swap a monomer and ship; the substitution has to be justified in a technical dossier, and a predicted Tg is not a measured Tg. What a tool like this genuinely buys a European team is ordering: it lets you throw out ninety percent of a candidate list before anyone books reactor time, then spend the lab budget proving the two or three pairs that the model says are worth proving.

The AI Act angle — briefly, because it is mostly a non-story here

Since 2 August 2026, the European AI Office and national authorities have binding enforcement powers, and general-purpose AI providers face real obligations including Article 50 transparency requirements. A narrow property-prediction model for monomers is not a general-purpose AI system and does not sit in a high-risk category. Treat it as industrial software with an uncertainty budget. The legal liability for a formulation remains with the company that signs the technical data sheet, not with the model that ranked the shortlist.

Worth noting how different this class of model is from what we stress-test in AI Arena. There we measure tokens per second, time to first token and VRAM ceilings for local LLMs. A descriptor-based property predictor is a different species entirely: small weights, tabular input, milliseconds of inference, and the entire value concentrated in the training dataset — not in the architecture.

Where I would stay skeptical

Three things the announcement does not settle.

Applicability domain. A model trained largely on conventional monomers is extrapolating when you hand it something with a furan ring and two ester groups. Neural networks do not politely decline to answer outside their training distribution; they return a confident number. Any serious deployment needs an uncertainty flag attached to every prediction.

Data sparsity on the bio side. The whole point of the tool is to guess at molecules that are underrepresented in the literature. That is exactly where the training data is thinnest.

Scale of validation. "Demonstrated in coating and adhesive applications" is a broad phrase. Pilot-scale latex production is where the secondary-nucleation problems that a Ws prediction is supposed to catch actually show up.

What a formulator can do with this on Monday

Use it as a filter, not an oracle. Build a candidate list of bio-based monomers with real supply availability in the EU, run the pairs against your incumbent copolymer's target properties, and take the top handful into a design-of-experiments campaign. The framework's value is proportional to how expensive your lab iteration is — and if you are running a 30-litre pilot reactor, one avoided trial pays for a lot of compute.

The tool itself is research output rather than a commercial product, so availability currently depends on what the authors release. That is the first thing worth checking before it enters anyone's workflow.

Do I need a GPU cluster to run a model like this?

No. Property-prediction models built on molecular descriptors and tabular kinetics data are small by modern standards. Inference on thousands of monomer pairs runs comfortably on a laptop CPU; training is a workstation job, not a datacentre job. The expensive part is assembling a trustworthy dataset, not the compute.

Can these predictions replace experimental validation?

No, and any supplier claiming otherwise should be treated with suspicion. Predicted glass transition temperatures and reactivity ratios are useful for ranking and for eliminating bad candidates early. Measured values — by DSC for Tg, by pulsed-laser polymerisation for kp — remain the reference for any formulation that ships.

Is the underlying data open?

That is the question to ask the authors. Predictive accuracy on paper is one thing; whether the training set, descriptors and uncertainty estimates are published is what determines if a European lab can actually reproduce the results and audit where the model is reliable.

Discussion

No comments yet — be the first to share your thoughts.
X

Don't miss out!

Subscribe for the latest news and updates.