🦾 Great milestone for open-source robotics: pi0 & pi0.5 by Physical Intelligence are now on Hugging Face, fully ported to PyTorch in LeRobot and validated side-by-side with OpenPI for everyone to experiment with, fine-tune & deploy in their robots! π₀.₅ is a Vision-Language-Action model which represents a significant evolution from π₀ to address a big challenge in robotics: open-world generalization. While robots can perform impressive tasks in controlled environments, π₀.₅ is designed to generalize to entirely new environments and situations that were never seen during training. Generalization must occur at multiple levels: - Physical Level: Understanding how to pick up a spoon (by the handle) or plate (by the edge), even with unseen objects in cluttered environments - Semantic Level: Understanding task semantics, where to put clothes and shoes (laundry hamper, not on the bed), and what tools are appropriate for cleaning spills - Environmental Level: Adapting to "messy" real-world environments like homes, grocery stores, offices, and hospitals The breakthrough innovation in π₀.₅ is co-training on heterogeneous data sources. The model learns from: - Multimodal Web Data: Image captioning, visual question answering, object detection - Verbal Instructions: Humans coaching robots through complex tasks step-by-step - Subtask Commands: High-level semantic behavior labels (e.g., "pick up the pillow" for an unmade bed) - Cross-Embodiment Robot Data: Data from various robot platforms with different capabilities - Multi-Environment Data: Static robots deployed across many different homes - Mobile Manipulation Data: ~400 hours of mobile robot demonstrations This diverse training mixture creates a "curriculum" that enables generalization across physical, visual, and semantic levels simultaneously. Huge thanks to the Physical Intelligence team & contributors Model: https://lnkd.in/eAEr7Yk6 LeRobot: https://lnkd.in/ehzQ3Mqy
Open Source Tools for Developers
Explore top LinkedIn content from expert professionals.
-
-
Announcing the release of RACECAR, the world’s first full-scale high-speed autonomous racing open dataset! 🏎 The dataset contains 11 interesting racing scenarios across two race tracks which include solo laps, multi-agent laps, overtaking situations, high-accelerations, banked tracks, obstacle avoidance, pit entry and exit at different speeds. Multi-Sensor (LIDAR, GNSS, RADAR, and Camera) data is available in both Open Robotics #ros2 and nuTonomy #nuScenes format, providing flexibility for researchers interested in robotics, computer vision, and autonomous driving. Six university teams who raced in the Indy Autonomous Challenge during 2021-22 season have contributed to this dataset. I would like to express my sincere appreciation to these teams for their valuable contributions: Cavalier Autonomous Racing FTM Institute of Automotive Technology TUM KAIST MIT-PITT-RW PoliMOVE Autonomous Racing Team TII Unimore Racing A paper authored by Amar Kulkarni John Chrosniak Emory Ducote Florian Sauerbeck Andrew Saba Utkarsh Chirimar John L. Marcello Cellina Madhur Behl will be presented at IROS '23, providing further insights into the dataset and its applications including benchmarking problems in #mapping, #localization, and #objectdetection To access the RACECAR dataset and accompanying tutorials, please visit: Data and Code: https://lnkd.in/e723QiH3 Paper: https://lnkd.in/eDMrBeHp Demo Reel: https://lnkd.in/esxMmgcH We are also grateful to the Amazon Web Services (AWS) Open Data program for their support in facilitating the sharing of this data. This is a step towards making full-scale autonomous racing accessible to the wider research community. I invite you to explore this groundbreaking dataset and pushing the boundaries of autonomous racing research! #autonomousvehicles #autonomousracing #ai #computervision #robotics #data #iros23 #research #opendata #technology
-
📷 𝗢𝗰𝘁𝗼𝗦𝗲𝗻𝘀𝗲: Teaching Robots to See When Cameras Can't Every robot perception model today has the same blind spot: it's built 𝗮𝗹𝗺𝗼𝘀𝘁 𝗲𝗻𝘁𝗶𝗿𝗲𝗹𝘆 𝗼𝗻 𝘃𝗶𝘀𝗶𝗼𝗻. That's a problem, because 𝗰𝗮𝗺𝗲𝗿𝗮𝘀 𝗳𝗮𝗶𝗹 𝗰𝗼𝗻𝘀𝘁𝗮𝗻𝘁𝗹𝘆: in fog, at night, in tunnels, in blinding sun glare. The most capable field robots keep hitting a wall on exactly this, fusing multiple sensors so they stay robust when one of them degrades. Came across 𝗢𝗰𝘁𝗼𝗦𝗲𝗻𝘀𝗲, a new project by Anthony Bisulco and collaborators at 𝗨𝗣𝗲𝗻𝗻'𝘀 𝗚𝗥𝗔𝗦𝗣 𝗟𝗮𝗯 (with 𝗕𝗿𝗼𝘄𝗻 𝗨𝗻𝗶𝘃𝗲𝗿𝘀𝗶𝘁𝘆), tackling this head on with 𝘁𝗵𝗿𝗲𝗲 𝗽𝗶𝗲𝗰𝗲𝘀: 🔧 𝗔𝗻 𝗼𝗽𝗲𝗻 𝗵𝗮𝗿𝗱𝘄𝗮𝗿𝗲 𝗽𝗹𝗮𝘁𝗳𝗼𝗿𝗺: Eight sensors (stereo RGB, stereo event cameras, thermal, LiDAR, IMU, GNSS, CAN), all hardware synced to one clock with a custom SyncBoard, fully open sourced. 📊 𝗔 𝟱𝟵 𝗵𝗼𝘂𝗿, 𝟮,𝟰𝟳𝟰 𝗸𝗺 𝗱𝗮𝘁𝗮𝘀𝗲𝘁: One of the largest event inclusive robotics datasets out there, driven across Long Island and Philadelphia through: • day, night, sunrise, and sunset conditions • tunnels, lens flare, fog, and packet loss • urban, suburban, and rural roads • boat and quadruped platforms too, not just cars 🧠 𝗔 𝗹𝗮𝘁𝗲-𝗳𝘂𝘀𝗶𝗼𝗻 𝗺𝗮𝘀𝗸𝗲𝗱 𝗮𝘂𝘁𝗼𝗲𝗻𝗰𝗼𝗱𝗲𝗿: A method that learns one shared representation across all sensors by masking and reconstructing tokens, then runs in real time (6.68 ms on an RTX 5090, 112 ms on an embedded Jetson Orin NX). --- 👉 𝗧𝗵𝗲 𝗵𝗲𝗮𝗱𝗹𝗶𝗻𝗲 𝗿𝗲𝘀𝘂𝗹𝘁: it beats every image only foundation model (DINOv2/v3, SigLIP 2, V-JEPA) on depth, flow, segmentation, and ego motion. And the gap gets wider, not narrower, exactly when vision struggles most: 𝘢𝘵 𝘯𝘪𝘨𝘩𝘵 𝘢𝘯𝘥 𝘶𝘯𝘥𝘦𝘳 𝘴𝘦𝘯𝘴𝘰𝘳 𝘥𝘦𝘨𝘳𝘢𝘥𝘢𝘵𝘪𝘰𝘯. 👉 𝗘𝘃𝗲𝗿𝘆𝘁𝗵𝗶𝗻𝗴 𝗶𝘀 𝗼𝗽𝗲𝗻: the dataset, the CAD, the code, and a live natural language search over the driving footage. --- ❓ Curious to hear from folks working on robotics or autonomy: where have you seen vision only perception break down hardest in your own systems? 🔗 𝗟𝗶𝗻𝗸 𝗶𝗻 𝘁𝗵𝗲 𝗰𝗼𝗺𝗺𝗲𝗻𝘁𝘀! #Robotics #SLAM #Mapping #Navigation #AutonomousVehicles #ComputerVision
-
As part of the Cosmos 3 release, we are not only open-sourcing models and training frameworks, but also releasing portions of the synthetic data used to train our models. One example is SDG-RobotSim: a dataset of 370K Isaac Sim-generated video clips capturing embodied robot behavior across a broad spectrum of scenarios, including complex locomotion, policy-driven manipulation, and collision/anomaly events. The motivation is straightforward: many physically grounded behaviors are underrepresented, expensive to collect, or difficult to observe at scale in real-world data. High-fidelity simulation provides a controllable mechanism for generating targeted distributions that complement real-world observations and improve coverage of rare but important phenomena. SDG-RobotSim is part of a much larger effort. Across Cosmos 3, NVIDIA is releasing 220 TB of synthetic data spanning robotics, physical interactions, autonomous driving, warehouse operations, and digital humans—designed to support the broader research community in developing next-generation world foundation models. Datasets: * SDG-RobotSim: https://lnkd.in/gSzJvpZk * SDG-PhyXSim: https://lnkd.in/gGAt_SuH * SDG-DigitalHumanSim: https://lnkd.in/g6t93Mgh * SDG-DriveSim: https://lnkd.in/gMVVZ8mq * SDG-WarehouseSim: https://lnkd.in/gSVJhAqf Excited to see how the community leverages large-scale synthetic data to advance physical AI and world modeling research. #PhysicalAI #Cosmos3
-
Mistakes! Mistakes! Mistakes! Physical AI assistants today can tell you "that's wrong"—but not 𝘄𝗵𝗮𝘁 went wrong, 𝘄𝗵𝗲𝗻 it became irreversible, or 𝘄𝗵𝗲𝗿𝗲 in the frame the mistake lives. That's like a teacher who only marks an X on your paper without any explanation. Our latest work, Mistake Attribution (MATT), 𝘁𝗼 𝗮𝗽𝗽𝗲𝗮𝗿 𝗮𝘁 𝗖𝗩𝗣𝗥 𝟮𝟬𝟮𝟲, goes beyond mistake detection in egocentric videos. The challenge is "fine-grained understanding." When someone picks up a bolt instead of a hammer, the current problem definition can flag the error, but they can't tell you which part of the instruction was violated, pinpoint the exact frame where recovery became impossible, or localize the mistake region in that frame. This greatly limits the practical value of physically-grounded upskilling with Physical AI assistants. We solve this with two contributions. First, MisEngine, a data engine that automatically constructs mistake datasets from existing action-recognition corpora, producing datasets two orders of magnitude larger than anything previously available. Second, MisFormer, a unified model that jointly attributes mistakes along semantic, temporal, and spatial dimensions, outperforming task-specific methods across the board as a single model. Full information about the work is included in the links below, including open-source code, pre-trained models, and our new datasets (Ego4D-M and EPIC-KITCHENS-M) with full attribution annotations. 📄 Paper: https://lnkd.in/e5ySVbwh 🌐 Project page: https://lnkd.in/e5nV_Qe9 💻 Code: https://lnkd.in/eHgbtyt4 🤗 Dataset and Weights: https://lnkd.in/ethnFkPQ Coauthors: Yayuan Li, Aadit J., Filippos Bellos From my teams at Voxel51 and University of Michigan College of Engineering University of Michigan Robotics Department Electrical and Computer Engineering at the University of Michigan Michigan AI Lab
-
Ever wondered what robots 🤖 could achieve if they could not just see – but also feel and hear? We introduce FuSe: a recipe for finetuning large vision-language-action (VLA) models with heterogeneous sensory data, such as vision, touch, sound, and more. We use language instructions to ground all sensing modalities by introducing two auxiliary losses. In fact, we find that naively finetuning on a small-scale multimodal dataset results in the VLA over-relying on vision, ignoring much sparser tactile and auditory signals. By using FuSe, pretrained generalist robot policies finetuned on multimodal data consistently outperform baselines finetuned only on vision data. This is particularly evident in tasks with partial visual observability, such as grabbing objects from a shopping bag. FuSe policies reason jointly over vision, touch, and sound, enabling tasks such as multimodal disambiguation, generation of object descriptions upon interaction, and compositional cross-modal prompting (e.g., “press the button with the same color as the soft object”). Moreover, we find that the same general recipe is applicable to generalist policies with diverse architectures, including a large 3B VLA with a PaliGemma vision-language-model backbone. We open source the code and the models, as well as the dataset, which comprises 27k (!) action-labeled robot trajectories with visual, inertial, tactile, and auditory observations. This work is the result of an amazing collaboration at Berkeley Artificial Intelligence Research with the other co-leads Joshua Jones and Oier Mees, as well as Kyle Stachowicz, Pieter Abbeel, and Sergey Levine! Paper: https://lnkd.in/dDU-HZz9 Website: https://lnkd.in/d7A76t8e Code: https://lnkd.in/d_96t3Du Models and dataset: https://lnkd.in/d9Er5Jsx
-
Meet 𝐃𝐑𝐎𝐈𝐃, a large-scale, in-the-wild robot manipulation dataset with input from numerous universities and R&D organizations. Creating large, diverse, high-quality robot manipulation datasets represents a crucial milestone in advancing more capable and robust robotic manipulation policies. However, generating such datasets presents significant challenges. Collecting robot manipulation data across varied environments entails logistical and safety hurdles and substantial hardware and human resources investments. Consequently, contemporary robot manipulation policies primarily rely on data from a limited number of environments, resulting in constrained scene and task diversity. In their collaborative endeavor, the authors introduce DROID (Distributed Robot Interaction Dataset), a diverse robot manipulation dataset comprising 76k demonstration trajectories or 350 hours of interaction data. This dataset was meticulously amassed across 564 scenes and 86 tasks, with contributions from 50 data collectors across North America, Asia, and Europe over 12 months. Their collective efforts illustrate that training with DROID yields policies characterized by enhanced performance, increased robustness, and superior generalization capabilities. The authors also make the entire dataset, along with the code for policy training and a comprehensive guide for replicating their robot hardware setup, available as open-source resources. Universities and Organizations that made up the DROID dataset team include: Stanford University University of California, Berkeley Toyota Research Institute Carnegie Mellon University The University of Texas at Austin Université de Montréal The University of Edinburgh Princeton University Columbia University University of Washington KAIST UC San Diego Google DeepMind University of California, Davis University of Pennsylvania 📝 Research Paper: https://lnkd.in/gGFFsKYK 📊 Project Page: https://lnkd.in/gbH8kqfv 🖥️ Dataset: https://lnkd.in/g5akx89p