Bioinformatics for Drug Discovery

Explore top LinkedIn content from expert professionals.

  • View profile for Pallavi N.

    AI Product Leader | Digital Growth Strategist | Ex-AI Research Engineer | Helping products get discovered

    7,344 followers

    A Nobel laureate just did the “IMPOSSIBLE.” And then dropped the code on GitHub… for free. Five days ago, David Baker’s lab published something wild in Nature: AI-designed antibodies from scratch. Yes! You read it RIGHT! No animals, No immunization, No years of trial-and-error. - Just a target - Just code - Just weeks The old world: - Immunize animals - Wait months - Screen thousands - Pray one works - Timeline: years + millions The new world: - Run RFdiffusion - Specify your target - Design antibody loops computationally - Timeline: weeks They tested the designs on flu, cancer targets, and bacterial toxins. And Guess What? Most worked exactly as predicted. Electron microscopy? 4 out of 5 antibodies matched the computer’s design atom by atom. This isn’t incremental science. This is a new coordinate system. What changed: Baker’s team rebuilt all six protein loops that control binding, something no previous method could do. Human antibody framework stays the same (to avoid immune rejection). But the binding mechanism is rewritten like software. What this unlocks: → Drug discovery cycles collapse from 5–10 years to months → New antibodies for targets previously considered “undruggable” → A potential shake-up of the $200B antibody market → And yes : RFantibody is free on GitHub for academic and commercial use Reality check: These molecules still need optimization for stability and safety. But the bottleneck has shifted forever. We’re no longer waiting years to find a candidate. We’re designing it and spending those years making it perfect. - AlphaFold predicted nature. - RFdiffusion writes new nature. - Xaira (Baker’s $1B spinout) is already turning these designs into clinical candidates. The question isn’t “can this be done?” It’s “who builds the fastest pipeline around it?” (Check out the comment section) Is this the future of drug discovery? Or are we overhyping the earliest signs of a revolution? #AI #Biotech #DrugDiscovery #ProteinDesign #MachineLearning #Healthcare #Innovation #OpenSource #ArtificialIntelligence

  • View profile for Najat Khan, PhD
    Najat Khan, PhD Najat Khan, PhD is an Influencer

    CEO and President | Member, Board of Directors, Recursion; Former Chief Data Science Officer & SVP/Global Head, Strategy & Portfolio, Pharma, J&J

    64,607 followers

    The next generation of drug discovery will be built on integrated, AI-native systems that connect automation, experimentation, computation, and machine learning into a continuous cycle of learning where every experiment makes the next one smarter. A recent article in Scientific American by Patrick Sisson explores how this new research infrastructure is beginning to take shape and highlights Recursion's pioneering work in this space. At Recursion, we run up to 2.2 million experiments each week and leverage more than 50 petabytes of proprietary biological, chemical, and patient data as part of an end-to-end learning engine for drug discovery and development. But scale alone isn't the differentiator. The real opportunity lies in transforming multimodal data into biological understanding and ultimately into new medicines. Take our neuroscience collaboration with Roche and Genentech. For decades, neuroscience drug discovery has been constrained by repeatedly investigating the same well-studied targets. To move beyond those limitations and explore entirely new biology, our teams developed advanced cell manufacturing capabilities to produce more than 100 billion human iPSC-derived microglia, the brain's resident immune cells, which are notoriously difficult to generate and study at scale. The result is a first-of-its-kind whole-genome Microglia Map comprising 46 million cellular images across 17,000 genes. This systems-level view of biology allows our AI models to move beyond traditional approaches, uncover novel biological insights, and identify therapeutic opportunities that may have otherwise remained hidden. What excites me most is what comes next. These maps – and the AI models trained on them – are the foundation. The real opportunity is translating them into novel, first-in-class therapeutic programs. That's the frontier we're pioneering: turning systems-level biological understanding into medicines for patients. There's still important work ahead, but we're making meaningful progress, and I'm excited about what's possible as we continue to push the boundaries of AI-native drug discovery. Stay tuned. #AI #DrugDiscovery #TechBio #Biotechnology #MachineLearning

  • View profile for Ross Dawson
    Ross Dawson Ross Dawson is an Influencer

    Futurist | Board advisor | Global keynote speaker | Founder: AHT Group - Informivity - Bondi Innovation | Humans + AI Leader | Bestselling author | Podcaster | LinkedIn Top Voice

    37,192 followers

    Synthetic biology is - quite literally - our future. A goundbreaking new biological foundation model Evo2 achieves state-of-the-art prediction of genetic variation impacts and generates coherent genome sequences, spanning all domains of life. A diverse team from leading research institutions including Arc Institute Stanford University NVIDIA University of California, Berkeley trained the model on 9.3 trillion DNA base pairs and has fully shared all code, parameters, and data. A few highlights from the paper (link in comments) 🔬 Zero-shot prediction achieves state-of-the-art accuracy in genetic variant interpretation. Evo 2 can predict the functional consequences of genetic mutations across all domains of life without specialized training. It surpasses existing models in assessing the pathogenicity of both coding and noncoding variants, including BRCA1 cancer-linked mutations. This generalist capability suggests Evo 2 could revolutionize genetic disease research, reducing reliance on expensive, manually curated datasets. 🛠 Genome-scale generation paves the way for synthetic life design. Evo 2 can generate full-length genome sequences with realistic structure and function, including mitochondrial genomes, bacterial chromosomes, and yeast DNA. Unlike prior models, Evo 2 ensures natural sequence coherence, improving synthetic biology applications like engineered microbes or artificial organelles. This sets the stage for programmable biology at an unprecedented scale. 🧬 Unprecedented long-context understanding revolutionizes genomic analysis. Evo 2 operates with a context window of up to 1 million nucleotides—far beyond the capabilities of previous models—allowing it to analyze genomic features across vast distances. This ability enables it to accurately identify regulatory elements, exon-intron boundaries, and structural components critical for understanding genome function. Its long-context recall is a major breakthrough for interpreting complex biological sequences. 🎛 Inference-time search enables controllable epigenomic design. Evo 2’s generative abilities extend beyond raw DNA sequence to epigenomic features, allowing researchers to design sequences with specific chromatin accessibility patterns. This approach successfully encoded Morse code messages into synthetic epigenomes, demonstrating a new method for controlling gene regulation via AI. This could lead to breakthroughs in gene therapy and epigenetic engineering. 🔮 Future potential: Toward AI-driven biological design and virtual cell modeling. Evo 2 represents a major leap toward AI-powered genomic engineering. Future iterations could integrate additional biological layers—such as transcriptomics and proteomics—to create virtual cell models that simulate complex cellular behaviors. This could revolutionize drug discovery, genetic therapy, and even synthetic life creation.

  • View profile for Gary Monk
    Gary Monk Gary Monk is an Influencer

    LinkedIn ‘Top Voice’ >> Follow for the Latest Trends, Insights, and Expert Analysis in Digital Health & AI

    48,709 followers

    Amazon launches AI drug discovery platform to accelerate antibody design and testing 🔘Amazon has launched Amazon Bio Discovery, a platform designed to help scientists generate and test antibody drug candidates faster by combining biological foundation models with AI agents in a single environment 🔘The key shift is accessibility, AI agents guide researchers through model selection, optimisation, and experiment design, meaning scientists without deep computational expertise can run advanced drug discovery workflows 🔘The platform acts as a marketplace of models, bringing together Amazon, open source, and partner algorithms such as Apheris and Profluent, while also allowing organisations to train and deploy their own proprietary models 🔘A major innovation is the “lab in the loop” system, where AI generated candidates are physically synthesised and tested by partners like Twist Bioscience and Ginkgo Bioworks, with results fed back to continuously improve the models 🔘Early results suggest significant acceleration, work with Memorial Sloan Kettering Cancer Center generated hundreds of thousands of antibody designs and moved from design to wet lab testing in weeks rather than up to a year 💬Drug discovery is shifting from isolated AI tools to integrated systems that connect models, data, and lab testing into a continuous learning loop, making research faster and more accessible #digitalhealth #ai #pharma

  • View profile for Alexey Navolokin

    FOLLOW ME for breaking tech news & content • helping usher in tech 2.0 • GM @ AMD • Turning AI, Cloud & Emerging Tech into Revenue

    799,266 followers

    AMD, UNSW Sydney & Pawsey: Redefining Real-Time Genomics with Slorado A major milestone for open science and high-performance genomics. AMD, UNSW Sydney, and the Pawsey Supercomputing Research Centre have introduced Slorado — the world’s first fully open-source, real-time nanopore DNA basecaller designed for AMD GPUs and powered by the ROCm open software platform. This breakthrough removes long-standing vendor lock-in and dramatically accelerates genomic workflows, empowering researchers with speed, scale, and flexibility. 🔬 What Slorado Enables + Fully open-source basecalling pipeline for nanopore sequencing + Runs on AMD GPUs via ROCm and supports hybrid GPU environments + Scales across multi-GPU and HPC infrastructures + Delivers performance parity with proprietary alternatives while improving accessibility ⚡ Performance Highlights on Pawsey’s Setonix Supercomputer Powered by AMD Instinct GPUs: + Full human genome decoded in: + 2.3 hours on MI250X GPUs + Just 0.8 hours on next-gen MI300X GPUs + High-accuracy models (HAC & SUP) also show significant acceleration without compromising data quality This level of performance transforms what once took days into hours — or even minutes — enabling faster research cycles, real-time pathogen surveillance, and scalable population genomics. 🌍 Why This Matters ✅ Democratizes access to high-performance genomics ✅ Accelerates discovery and clinical research ✅ Strengthens reproducibility through open-source transparency ✅ Expands AMD’s role as a trusted platform for scientific computing and AI ✅ Bridges HPC, AI, and bioinformatics into a unified ecosystem Slorado is more than a tool — it’s a signal of where the future of genomics is heading: open, accelerated, and accessible at global scale. AMD continues to push the boundaries of what’s possible in scientific computing – from AI to genomics and beyond. 🔗 Explore more: https://lnkd.in/ghSRHX7S #AMD #Genomics #OpenScience #HPC #AIinHealthcare #ROCm #InstinctGPUs #Supercomputing #Innovation #Bioinformatics #FutureOfScience #AMDBrandAmbassador

  • View profile for Shilpa Rao

    Driving Access to Health with AI |Ex Head-AI platforms |Serial Innovator| Independent Director|Purpose Alchemist

    29,468 followers

    𝟗𝟖% 𝐨𝐟 𝐘𝐨𝐮𝐫 𝐆𝐞𝐧𝐨𝐦𝐞 𝐖𝐚𝐬 𝐈𝐠𝐧𝐨𝐫𝐞𝐝. 𝐔𝐧𝐭𝐢𝐥 𝐍𝐨𝐰 What if we could read the genome like a story — not just in fragments, but as a whole, with rhythm, meaning, and twist endings? Google just brought us one step closer. Introducing AlphaGenome — DeepMind and Google’s newest AI model that predicts how your DNA is read, regulated, and sometimes... misinterpreted. But first — a quick decode: DNA is made of 4 letters — A, T, C, G — the alphabet of life. These 3 billion letters tell your cells what to do. But they’re not read linearly like a book — they fold and loop in 3D, bringing distant parts together to turn genes on or off. This folding is everything. A mutation buried 100,000 letters away could still influence gene activity — just because folding brought it into the “wrong neighborhood.” That’s where AlphaGenome shines. It reads up to 1 million DNA letters at a time — enough to catch complex folding patterns and regulatory cues in one shot. “But wait — we have 3 billion letters!” Yes — and like reading a novel one chapter at a time, AlphaGenome moves across the genome in overlapping tiles, decoding each “functional neighborhood” in high detail. And here’s what it does inside each tile: - Predicts where genes start, stop, splice, and fold - Estimates RNA expression and protein-binding activity - Identifies how a tiny mutation might ripple across the system - Models splicing errors that cause rare diseases - Spots cancer-driving mutations in “non-coding” DNA — the 98% we used to overlook Previous models (like Enformer) had to trade off resolution vs. context. AlphaGenome offers both. It’s like watching your genome in high-def, panoramic, slow-mo… at the same time. Now available via API for non-commercial research. Not for clinical use — yet. But a massive step toward understanding how life truly runs under the hood. Sometimes, it’s not about what the DNA says — but where it folds… where it pauses… and what happens in the quiet. AlphaGenome teaches us something deeper: That meaning doesn’t always lie in the loudest signals. Often, it’s in the 98% we ignore. The background. The regulation. The timing. Life, too, is like that. It’s not just what you do. It’s when, how, and in what context you show up. The pauses between the notes matter. The unseen structure holds the story. And maybe… that’s where the real change begins Amit Saxena Ajay Nandgaonkar Suchitaa Paatil Sanju S Anju Goel Taruna Anand #AlphaGenome #GoogleDeepMind #GenomicsExplained #FutureOfHealth #VariantEffectPrediction #SyntheticBiology #PrecisionMedicine #DNAMagic #AIinHealthcare #AI #AccessForAll

  • View profile for Ganna Posternak, PhD

    Drug Discovery Scientist | Biotech | Scientific Strategy | 15+ Years in Research

    7,107 followers

    Machine Learning in Preclinical Drug Discovery 🧬💊 Machine learning (ML) is increasingly integrated into preclinical drug discovery, offering promising advancements across hit identification, mechanism-of-action elucidation, and translational investigations. A recent paper in Nature Chemical Biology, "Machine Learning in Preclinical Drug Discovery", provides a thorough analysis of how ML is being utilized to enhance efficiency in early-stage drug development. 🔬 Key Insights from the Paper 1️⃣ Hit Identification & Virtual Screening Traditionally, high-throughput screening (HTS) has been the gold standard for identifying potential drug candidates. However, it is resource-intensive and slow. ML-based virtual screening, powered by deep learning models and molecular featurization techniques, is enabling rapid exploration of chemical libraries far beyond what traditional HTS can achieve. The paper highlights the impact of message-passing neural networks (MPNNs) and Deep Docking as effective methods for prioritizing hit compounds. 2️⃣ Mechanism-of-Action (MOA) Elucidation Understanding how a compound interacts with biological targets is critical for drug development. ML is now playing a pivotal role in MOA elucidation through: AlphaFold and RoseTTAFold: AI-driven protein structure prediction is accelerating target identification and binding site analysis. Generative models: Variational autoencoders (VAEs) and diffusion models are not only aiding in de novo drug design but also helping predict chemical interactions with biological systems. 3️⃣ Translational Investigations & ADMET Predictions Many promising compounds fail in later stages due to poor pharmacokinetics and toxicity profiles. ML is being leveraged to enhance ADMET predictions, improving the likelihood of clinical success. The paper discusses advancements in: Solubility and Lipophilicity Predictions: ML-driven models now outperform traditional log(P) estimations, increasing the reliability of early-stage compound selection. Toxicity Screening: AI-powered tools are improving predictions of hERG binding and organ toxicity, reducing late-stage failures. 🚀 The Future of AI in Drug Discovery While ML is proving to be a game-changer, challenges remain, including data quality, interpretability of AI models, and integration with experimental validation. The paper underscores the importance of open-source datasets, AI transparency, and active learning strategies to enhance model accuracy. 🔗 Read the full paper here: https://lnkd.in/gMtXHrHi AI is reshaping the landscape of drug discovery. As these technologies evolve, collaboration between computational scientists, biologists, and chemists will be critical to unlocking their full potential. #AI #MachineLearning #DrugDiscovery #Pharma #Biotech #ArtificialIntelligence #ComputationalBiology #NatureChemicalBiology

  • 🚨 I'm excited to share our latest review, “New approach methodologies for drug discovery,” published in Cell by Cell Press and selected as a Featured Article. For decades, drug discovery has relied heavily on animal models. Yet, with persistently high clinical failure rates, a fundamental question remains: How predictive are animal models of human biology, and are there better alternatives? In this review, we highlight a paradigm shift toward human-centric New Approach Methodologies (NAMs), driven by rapid advances in both policy and technology. On the regulatory front, we discuss major transitions led by agencies such as the FDA (FDA Modernization Acts 1.0 → 2.0 → 3.0) and The National Institutes of Health (stem cell guidelines and the establishment of national organoid initiatives). 🔬 On the technology side, we frame NAMs evolution across three domains: ·        “New” - foundational 2D stem cell–based systems ·        “Newer” - advanced 3D organoid-based models ·        “Newest” - future-facing in silico and AI-driven platforms Across these domains, we highlight emerging therapeutic candidates, cutting-edge models, and translational and clinical applications. We also examine key biological, technical, and regulatory bottlenecks that need to be addressed to enable robust translational adoption, and discuss ongoing clinical efforts and societal considerations for responsible implementation. 🌍 Looking forward. If the past 30 years of drug discovery were shaped by animal models, the next 30 years, animal models will likely transition from a central to a supporting role, following the 3Rs principle: refinement, reduction, and ultimately replacement. Instead, the field will likely be defined by human-centric NAMs, powered by multiscale platforms, multi-omics data, and AI-enabled pipelines. This transformation is not only scientific, but also societal, aligning drug development more closely with human biology while reducing cost, inefficiency, and ethical burden. 👏 Congratulations to an outstanding team: Wenqiang (Eric) Liu, Paul Pang, Catherine Wu and Danilo Tagle from Stanford University School of Medicine, Stanford Cardiovascular Institute, Stanford Department of Medicine, Greenstone Biosciences, National Center for Advancing Translational Sciences (NCATS), The National Institutes of Health. 📄 Please check the full paper here: https://lnkd.in/estaR2cq #NAMs #DrugDiscovery #StemCell #Organoids #AI #PrecisionMedicine #TranslationalScience

  • View profile for 🎯  Ming "Tommy" Tang

    Director of Bioinformatics | Cure Diseases with Data | Author of From Cell Line to Command Line | AI x bioinformatics | >130K followers, >30M impressions annually across social platforms| Educator YouTube @chatomics

    69,739 followers

    Machine learning (ML) is revolutionizing genomics, but common pitfalls can lead to misleading results. Here's a thread on how to avoid them 🧵 1/ Pitfall 1: Distributional Differences Genomic data often exhibits inherent biological structure, leading to distributional differences. This can impact model performance when training & test sets have different distributions than the prediction set. Example: Models trained on in vitro data often perform poorly on in vivo data, as seen in transcription factor binding site prediction. This highlights the need to carefully consider the context in which a model will be applied. 2/ Pitfall 2: Dependent Examples Genomic data is often interconnected, violating the independence assumption of many ML models. This can inflate performance estimates during cross-validation. Example: When predicting protein-protein interactions, pairs sharing a protein are correlated. This can be mitigated by employing group k-fold cross-validation, ensuring dependent examples don't cross the train-test divide. 3/ Pitfall 3: Confounding Unmeasured variables can create or mask associations, leading to misinterpretations. Example: In GWAS, population structure can confound genotype-phenotype relationships. The infamous autism spectrum disorder prediction model initially seemed successful but failed to replicate after accounting for population structure. 4/ Pitfall 4: Leaky Preprocessing Data processing can inadvertently leak information from the test set into the training set, resulting in over-optimistic performance estimates. Example: Feature selection based on the entire dataset before cross-validation, common in DNA methylation analysis, introduces leakage. (this is probably one of the most common mistakes I see...) hold a test dataset that you never touch until the final step. Pitfall 5: Unbalanced Classes Datasets with uneven class distribution can lead to models overfitting the majority class. Example: Predicting enhancers is challenging due to the small proportion of positive examples. Resampling techniques and choosing appropriate performance metrics like auPR can help address this. Also use PR not ROC to evaluate your model https://lnkd.in/eh9JHmxc Key Takeaways: • Genomic data has unique characteristics that require careful consideration when applying ML. • Thoroughly inspect your data, considering potential dependencies, confounders, & class imbalance. • Employ appropriate techniques like group k-fold cross-validation, balancing methods, & robust performance metrics. By understanding these pitfalls and taking steps to mitigate them, we can ensure that ML applications in genomics yield reliable and insightful results. 💪 dive deep into the paper https://lnkd.in/eDFWnW9U I hope you've found this post helpful. Follow me for more. Subscribe to my FREE newsletter https://lnkd.in/erw83Svn

  • View profile for Jorge Bravo Abad

    Physicist at UAM · Director, AI for Materials Lab · Building AI-driven loops turning scientific discovery into infrastructure · Two books on AI and science

    31,835 followers

    AI-powered virtual screening that scores 10 trillion protein-ligand pairs in a single day Of ~20,000 human protein-coding genes, only about 10% have been successfully targeted by FDA-approved drugs or have documented small-molecule binders. The bottleneck isn't biology—it's computational scale. Traditional molecular docking takes seconds to minutes per protein-ligand pair, making genome-wide screening essentially impossible with current resources. Yinjun Jia and coauthors tackle this head-on with DrugCLIP, a contrastive learning framework that reframes virtual screening as a dense retrieval problem—similar to how modern search engines work. The key innovation: encode protein pockets and small molecules into a shared latent space using separate neural networks, then use cosine similarity for ultrafast ranking. The model is pretrained on 5.5 million synthetic pocket-ligand pairs extracted from protein structures, then fine-tuned on 40,000 experimentally determined complexes. The speed gains are staggering—up to 10 million times faster than docking. Combined with GenPack, a generative module that refines pocket detection on AlphaFold2-predicted structures, DrugCLIP enables screening at a scale previously unthinkable: 500 million compounds against ~10,000 human proteins, scoring more than 10 trillion pairs in under 24 hours on just 8 GPUs. The wet-lab validations are equally compelling. For norepinephrine transporter (NET), a 15% hit rate with two inhibitors structurally confirmed by cryo-EM. For TRIP12—a challenging E3 ubiquitin ligase with no known inhibitors or holo structures—a 17.5% hit rate using only AlphaFold2 predictions, with functional enzymatic inhibition confirmed. The resulting database, GenomeScreenDB, covers ~20,000 pockets from 10,000 proteins—nearly half the human genome—and is freely available at drugclip.com. The message is clear: by combining contrastive representation learning with generative pocket refinement and AlphaFold structures, we've entered an era where genome-wide drug discovery becomes computationally tractable, opening systematic exploration of the vast undrugged proteome. Paper: https://lnkd.in/e7aGUvAX #DrugDiscovery #ArtificialIntelligence #MachineLearning #DeepLearning #VirtualScreening #ComputationalBiology #AlphaFold #ProteinScience #Biotech #AIforScience #StructuralBiology #Bioinformatics #Pharmaceuticals #ComputationalChemistry #PrecisionMedicine

Explore categories