AIRoA MoMa Dataset: A Large-Scale Hierarchical Dataset for Mobile Manipulation Paper • 2509.25032 • Published Sep 29, 2025
FlexLAM: Resolving the Bottleneck Trade-off in Latent Action Learning Paper • 2606.19408 • Published Jun 17
Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs Paper • 2605.07447 • Published May 8 • 3
Sparse Autoencoders as Plug-and-Play Firewalls for Adversarial Attack Detection in VLMs Paper • 2605.07447 • Published May 8 • 3
VIR-Bench: Evaluating Geospatial and Temporal Understanding of MLLMs via Travel Video Itinerary Reconstruction Paper • 2509.19002 • Published Nov 15, 2025 • 3
Looking Backward: Streaming Video-to-Video Translation with Feature Banks Paper • 2405.15757 • Published May 24, 2024 • 15
Immiscible Diffusion: Accelerating Diffusion Training with Noise Assignment Paper • 2406.12303 • Published Jun 18, 2024 • 4
StreamDiT: Real-Time Streaming Text-to-Video Generation Paper • 2507.03745 • Published Jul 4, 2025 • 33
Traveling Across Languages: Benchmarking Cross-Lingual Consistency in Multimodal LLMs Paper • 2505.15075 • Published May 21, 2025 • 1
Angles Don't Lie: Unlocking Training-Efficient RL Through the Model's Own Signals Paper • 2506.02281 • Published Jun 2, 2025 • 4
HallE-Switch: Rethinking and Controlling Object Existence Hallucinations in Large Vision Language Models for Detailed Caption Paper • 2310.01779 • Published Oct 3, 2023 • 4
SparseFusion: Fusing Multi-Modal Sparse Representations for Multi-Sensor 3D Object Detection Paper • 2304.14340 • Published Apr 27, 2023
What Matters to You? Towards Visual Representation Alignment for Robot Learning Paper • 2310.07932 • Published Oct 11, 2023
Open X-Embodiment: Robotic Learning Datasets and RT-X Models Paper • 2310.08864 • Published Oct 13, 2023 • 3
Visual Transformers: Token-based Image Representation and Processing for Computer Vision Paper • 2006.03677 • Published Jun 5, 2020 • 3
NeRF-Det: Learning Geometry-Aware Volumetric Representation for Multi-View 3D Object Detection Paper • 2307.14620 • Published Jul 27, 2023 • 15
DELFlow: Dense Efficient Learning of Scene Flow for Large-Scale Point Clouds Paper • 2308.04383 • Published Aug 8, 2023
Sparse R-CNN: End-to-End Object Detection with Learnable Proposals Paper • 2011.12450 • Published Nov 25, 2020
Sparse VideoGen: Accelerating Video Diffusion Transformers with Spatial-Temporal Sparsity Paper • 2502.01776 • Published Feb 3, 2025 • 3
Sparse Diffusion Policy: A Sparse, Reusable, and Flexible Policy for Robot Learning Paper • 2407.01531 • Published Jul 1, 2024