Insilico Medicine releases frontier AI models for drug discovery

· News-Medical

Insilico Medicine, a clinical-stage generative artificial intelligence (AI)-driven drug discovery company, today announced the release of a series of frontier specialist models for chemistry and biology, trained through its MMAI Gym for Science framework. Across chemistry and biology, the models demonstrate state-of-the-art (SOTA) performance on more than 50 benchmark tasks.

Alex Zhavoronkov, Ph.D., founder and CEO of Insilico MedicineFor years, language models have been seen as versatile conversational tools, but when it came to complex, data-scarce problems in drug discovery, specialized computational methods held the upper hand. With MMAI Gym, we are proving that language-model architectures, when trained with domain-specific multi-task precision, can move far beyond high-level general knowledge. They are now competing with and outperforming dedicated scientific methods on real-world chemistry and geroscience benchmarks, marking a fundamental shift toward truly predictive AI for human health."

Historically, low-data drug discovery tasks such as ADMET and target potency prediction have been dominated by specialized models developed specifically for molecular data. Language models have largely remained outside this competitive landscape. The latest results from MMAI Gym suggest that this is beginning to change, with language-model-based specialists now matching or outperforming established approaches on selected tasks.

The models released through MMAI Gym are small language models trained into scientific specialists. Each model is fine-tuned for a defined category of scientific tasks and trained across multiple related problems within that category. These specialists are benchmarked directly against established computational methods using the same datasets and train-test splits. All baseline models were re-trained and run under the same evaluation setup and ground truth data splits, enabling head-to-head comparison rather than relying on scores reported on public leaderboards.

The current set of chemistry models includes language-model-based specialists for chemical synthesis, ADMET prediction, and potency prediction across GPCR and kinase panels.

For ADMET prediction, the specialist was evaluated across 28 tasks and achieved SOTA-level performance relative to established computational methods, with particularly strong results across drug-drug interaction risk, cytotoxicity, and pharmacokinetic properties. Accurate prediction of these properties can help identify liabilities earlier and support compound prioritization before more resource-intensive experimental studies. As a category-focused model, it can evaluate multiple ADMET-related endpoints using a single model rather than relying on separate models for each individual property. Furthermore, the model delivered the strongest overall performance on a held-out experimental set spanning 14 properties from the Drug Candidate Essentials benchmark, outperforming established SOTA baselines.

For target activity prediction, two protein family-focused models were trained: one for GPCRs, covering 44 receptors, and one for kinases, covering 67 enzymes. In both categories, the models achieved SOTA-level IC50 prediction performance on selected targets. These models can support a range of drug discovery applications, from virtual screening and lead prioritization to the design of multi-target compounds and selectivity profiling to identify and minimize potential off-target activity.

For chemical synthesis, the models are specialized for single-step retrosynthesis and are built on Liquid AI's compact 2.6B-parameter architecture. The retrosynthesis specialists outperform leading dedicated methods on both standard in-distribution benchmarks and more challenging out-of-distribution evaluation sets. A previous version is already available through Microsoft Marketplace (https://marketplace.microsoft.com/en-us/product/insilicomedicineusinc1751604844440.lfm2-mmai-chem-ssrs), while an updated, higher-performing version is planned for release there soon. Beyond standalone retrosynthesis prediction, these models can also serve as components in broader computational drug discovery and synthesis-planning pipelines that require a strong single-step retrosynthesis engine.

Beyond chemistry, MMAI Gym is also being applied to biology-focused specialist models trained across diverse biological data types and related prediction tasks. These models extend the same category-based approach to clinical, omics, and molecular biology applications and, on selected benchmarks, have demonstrated SOTA performance compared with substantially larger general-purpose frontier models.

Collectively, the models demonstrate SOTA-level or superior performance on more than 70 benchmark tasks across chemistry and biology. The evaluation results are documented using DDD Bench (Drug Discovery and Development Benchmark, https://dddbench.insilico.com/), a standardized benchmarking framework established by Insilico Medicine to systematically assess AI models across key drug discovery and development tasks.

Importantly, these results are significant not because language models can discuss chemistry or biology, but because language-model-based specialists can now compete with established scientific methods and, on selected tasks, outperform methods developed specifically for these prediction problems.

MMAI Gym is designed around this approach: transforming language models into domain-specific scientific specialists through targeted training across related tasks and rigorous head-to-head benchmarking against established methods. These results demonstrate that language models can move beyond general scientific knowledge and become competitive predictive and generative models for practical applications in drug discovery and aging research.

MMAI Gym for Science was first unveiled at NeurIPS 2025, where Insilico Medicine demonstrated that targeted scientific training could substantially improve the performance of a general-purpose language model across drug discovery tasks. Since then, MMAI Gym has expanded through a series of research efforts advancing multi-task chemistry modeling, benchmarking, and domain adaptation. At ICLR 2026 (https://openreview.net/forum?id=4m2TcQpwkf), Insilico Medicine and Liquid AI presented LFM2-2.6B-MMAI, a multi-task chemistry model trained through MMAI Gym to perform across a broad range of drug discovery tasks. At ICML 2026 (https://openreview.net/forum?id=ABhGgV7pov), Insilico Medicine introduced new benchmarks and evaluation metrics for single-step retrosynthesis designed to provide a more rigorous assessment beyond conventional Top-K accuracy. Using this framework, the team trained a language-model-based specialist that achieved performance competitive with frontier foundation models and established chemical specialist methods, outperforming them on several evaluation metrics. Most recently, a new study from Insilico Medicine and Liquid AI featuring 2.6B-MMAI and 24B-A2B-MMAI, two multi-task chemistry models spanning compact and larger-capacity architectures, was accepted to EMNLP 2026. Together, these studies establish a growing body of evidence that language-model architectures can be systematically adapted into high-performing scientific specialists across diverse chemistry tasks.

Recently, Insilico reported total revenue of approximately $106 million in the first half of 2026, a 287% year-over-year increase, and achieved its first profitable half-year since listing, with an adjusted net profit exceeding $51 million. This milestone was driven by a series of out-licensing, co-development, and R&D collaborations with global partners, including Eli Lilly, Servier, Takeda, SK Biopharmaceuticals, Qilu Pharmaceutical, Hygtia Therapeutics, CMS, and Tenacia. As of the latest practicable date, the total contract value of transactions announced by Insilico in 2026 reached approximately $7.3 billion, pushing the cumulative contract value of its major collaborations since 2021 approximately $11 billion.

On the AI-driven R&D front, Insilico nominated nine development candidates within nine months of 2026 as of late August, setting a new company record for annual pipeline productivity and achieving eight clinical milestones across its proprietary and co-developed programs. Leading this progress is Rentosertib (ISM001-055), the world's first drug candidate discovered and developed using generative AI, which has advanced to a Phase III trial evaluating for idiopathic pulmonary fibrosis (IPF).

Source:

InSilico Medicine