AI Models & Platforms
Insilico Medicine Releases Open Longevity AI Toolkit in Cell Study

Insilico Medicine on September 17, 2026 announced the publication of a study in Cell introducing an openly released AI toolkit for aging biology: the LongevityBench benchmark, a family of five compact Longevity-LLMs, and the Longevity Claw agentic research platform.
Publication in Cell
The paper, “An open benchmark and language models for AI in aging biology,” appears in Cell volume 189, issue 19, at pages 5980–5994.e8, dated September 17, 2026. It was published open access under a Creative Commons Attribution 4.0 license with the DOI 10.1016/j.cell.2026.08.026, and it lists 13 authors, including Alex Zhavoronkov, Vladimir Naumov, Denis Sidorenko, Alex Aliper, Ramin Hasani, Alexander Amini, Vadim N. Gladyshev and Fedor Galkin.
Insilico said the study was selected as the cover feature of the journal’s September 17, 2026 issue and was conducted with researchers from Liquid AI, the Buck Institute for Research on Aging, and Harvard Medical School and Brigham and Women’s Hospital.
According to the announcement, the Cell publication follows Insilico’s September 7, 2026 study in Nature Biotechnology, which reported that rentosertib, the company’s AI-discovered and AI-designed drug candidate for idiopathic pulmonary fibrosis, reduced biological age across six independent proteomic aging clocks in a Phase IIa clinical trial.
An Open Benchmark for Aging Biology
The study introduces LongevityBench as an open suite of 17 tasks spanning five biodata domains: clinical data, genetics, epigenetics, transcriptomics and proteomics. In the paper’s summary, the authors write that no existing benchmark evaluated whether AI systems can interpret these heterogeneous data types in the context of aging biology. According to the announcement, the benchmark was designed to reduce the likelihood that models could succeed through recall of information encountered during training, testing instead the ability to analyze biological data, recognize meaningful patterns and solve problems relevant to aging research.
The authors used LongevityBench to assess 18 frontier AI systems from six developer teams, which the announcement identifies as OpenAI, Google, Anthropic, xAI, DeepSeek and Moonshot AI. The paper reports that no single model dominated all tasks, that performance shifted with how questions were phrased, and that omics-based age prediction was the hardest task regardless of model scale.
The project’s public leaderboard, which tracks 26 models across the 17 tasks, lists datasets drawn from NHANES clinical measurements, GEO DNA methylation, GTEx bulk RNA-seq, Olink plasma proteomics, and the OpenGenes and SynergyAge genetics resources. On the leaderboard’s aggregate rank score, where lower is better, Gemini 3.1 Pro at 8.2 is listed as the best frontier model, with Claude Opus-4.6 at 9.2.
Compact Longevity-LLMs
To test whether those gaps could be closed without frontier-scale resources, the researchers fine-tuned a family of five multitask Longevity-LLMs ranging from 0.6 billion to 9 billion parameters on domain-specific aging data, and the paper reports that the compact models matched or exceeded far larger frontier systems on the benchmark. According to the announcement, the models were trained using Insilico’s MMAI Gym for Science, which the company describes as a training ground for language models that applies its proprietary data, reasoning datasets and validated models, and were built on Liquid AI’s LFM2 architecture and Alibaba’s Qwen3 and Qwen3.5 model families.
The leaderboard lists L-Qwen3.5-9B as the best overall system with a 4.4 aggregate rank score, followed by L-LFM2-2.6B at 7.6 and L-Qwen3-1.7B at 7.8. Among the leaderboard’s selected results, L-Qwen3.5-9B reached 0.868 concordance on GEO DNA-methylation age prediction versus 0.685 for the best frontier model, and L-Qwen3-0.6B recorded a 5.7-year mean absolute error on Olink proteomic age prediction versus 10.1 years for the best frontier model, along with 0.890 balanced accuracy on NHANES 10-year mortality prediction.
The models and benchmark data are publicly available in a Hugging Face collection that includes the longebench dataset and the longevity-llm 9B, Qwen3-0.6B-Longevity, Qwen3-1.7B-Longevity, LFM2-2.6B-Longevity and LFM2-1.2B-Longevity models.
Longevity Claw and Autonomous Target Discovery
The team embedded L-Qwen3.5-9B into Longevity Claw, an open-source agentic platform that combines the specialized model with tools for gene-set enrichment analysis, biological aging-clock calculation, population-level profiling, evidence retrieval and synthesis, and candidate target evaluation and prioritization. Insilico said the platform was designed to formulate and execute multi-step research workflows rather than only respond to individual questions.
Deployed across 14 recognized hallmarks of aging, the platform nominated 328 genes as potential targets for aging intervention, and the candidates showed statistically significant enrichment of up to 5.6-fold against an independently published reference set of experimentally supported aging-related targets, according to the announcement. One nominated gene, KDM1A, was independently validated in a separate published study as a dual-purpose aging and cancer target whose modulation extended lifespan in C. elegans, the announcement said.
The project’s GitHub repository, published under an MIT license, documents LongevityClaw as an agent that predicts biological age across 233 clocks spanning six modalities with 429,165 coefficients, alongside population reference datasets and a novel-target discovery module that scores candidates on six dimensions, including novelty, druggability, confidence and safety, across the 14 hallmarks.
Open Release and Stated Goals
Insilico said it is releasing the benchmark, the specialized models, training resources, evaluation code and the Longevity Claw platform to enable independent testing, validation and further development by researchers worldwide.
“We are developing benchmarked, agentic systems that can evolve into personalized longevity assistants and longevity companions, ultimately helping people monitor and improve their healthspan,” said Alex Zhavoronkov, founder and co-CEO of Insilico Medicine.
Insilico said the open framework is intended to give scientists a common foundation for measuring progress in AI-enabled aging research and to help distinguish systems that demonstrate genuine biological reasoning from those that primarily reproduce information contained in their training data.












