Flat-Pack Bench: Evaluating Spatio-Temporal Understanding in Large Vision-Language Models through Furniture Assembly Paper • 2605.21625 • Published May 20 • 3
SceneAligner: 3D-Grounded Floorplan Localization in the Wild Paper • 2605.22581 • Published May 21 • 8
SceneAligner: 3D-Grounded Floorplan Localization in the Wild Paper • 2605.22581 • Published May 21 • 8
Raster2Seq: Polygon Sequence Generation for Floorplan Reconstruction Paper • 2602.09016 • Published May 11 • 5
Collaborative Transformers for Grounded Situation Recognition Paper • 2203.16518 • Published Mar 30, 2022
PromptStyler: Prompt-driven Style Generation for Source-free Domain Generalization Paper • 2307.15199 • Published Jul 27, 2023 • 13
Robust 3D Shape Reconstruction in Zero-Shot from a Single Image in the Wild Paper • 2403.14539 • Published Mar 21, 2024
Cross-Attention of Disentangled Modalities for 3D Human Mesh Recovery with Transformers Paper • 2207.13820 • Published Jul 27, 2022
Lost in Backpropagation: The LM Head is a Gradient Bottleneck Paper • 2603.10145 • Published Mar 10 • 13
YOLOE-26: Integrating YOLO26 with YOLOE for Real-Time Open-Vocabulary Instance Segmentation Paper • 2602.00168 • Published Jan 29 • 2
Anatomy of a Machine Learning Ecosystem: 2 Million Models on Hugging Face Paper • 2508.06811 • Published Aug 9, 2025 • 6
Simple Guidance Mechanisms for Discrete Diffusion Models Paper • 2412.10193 • Published Dec 13, 2024 • 1
Multi-Turn Code Generation Through Single-Step Rewards Paper • 2502.20380 • Published Feb 27, 2025 • 32
Robotouille: An Asynchronous Planning Benchmark for LLM Agents Paper • 2502.05227 • Published Feb 6, 2025
FiVA: Fine-grained Visual Attribute Dataset for Text-to-Image Diffusion Models Paper • 2412.07674 • Published Dec 10, 2024 • 20