vibe_ov2

classroom

AI & ML interests

None defined yet.

Recent Activity

xiangan submitted a paper about 17 hours ago

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

jankin123 submitted a paper 17 days ago

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

xiangan authored a paper 3 months ago

OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

View all activity

submitted a paper to Daily Papers about 17 hours ago

LLaVA-OneVision-2: Towards Next-Generation Perceptual Intelligence

Paper • 2605.25979 • Published 3 days ago • 16

submitted a paper to Daily Papers 17 days ago

4DThinker: Thinking with 4D Imagery for Dynamic Spatial Understanding

Paper • 2605.05997 • Published 21 days ago • 17

authored a paper 3 months ago

OneVision-Encoder: Codec-Aligned Sparsity as a Foundational Principle for Multimodal Intelligence

Paper • 2602.08683 • Published Feb 9 • 52

authored 2 papers 4 months ago

ProCLIP: Progressive Vision-Language Alignment via LLM-based Embedder

Paper • 2510.18795 • Published Oct 21, 2025 • 11

DanQing: An Up-to-Date Large-Scale Chinese Vision-Language Pre-training Dataset

Paper • 2601.10305 • Published Jan 15 • 37

submitted a paper to Daily Papers 4 months ago

OmniVideo-R1: Reinforcing Audio-visual Reasoning with Query Intention and Modality Attention

Paper • 2602.05847 • Published Feb 5 • 12

authored a paper 4 months ago

RoboBrain 2.5: Depth in Sight, Time in Mind

Paper • 2601.14352 • Published Jan 20 • 13

authored 3 papers 5 months ago

Towards Cross-View Point Correspondence in Vision-Language Models

Paper • 2512.04686 • Published Dec 4, 2025

RoboOS-NeXT: A Unified Memory-based Framework for Lifelong, Scalable, and Robust Multi-Robot Collaboration

Paper • 2510.26536 • Published Oct 30, 2025

Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation

Paper • 2512.23703 • Published Dec 29, 2025 • 7

submitted a paper to Daily Papers 5 months ago

Robo-Dopamine: General Process Reward Modeling for High-Precision Robotic Manipulation

Paper • 2512.23703 • Published Dec 29, 2025 • 7

authored 8 papers 7 months ago

VisRL: Intention-Driven Visual Perception via Reinforced Reasoning

Paper • 2503.07523 • Published Mar 10, 2025 • 1

Visual Document Understanding and Question Answering: A Multi-Agent Collaboration Framework with Test-Time Scaling

Paper • 2508.03404 • Published Aug 5, 2025 • 4

SIFThinker: Spatially-Aware Image Focus for Visual Reasoning

Paper • 2508.06259 • Published Aug 8, 2025 • 2

Visual Multi-Agent System: Mitigating Hallucination Snowballing via Visual Flow

Paper • 2509.21789 • Published Sep 26, 2025 • 9

Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited Views

Paper • 2510.18632 • Published Oct 21, 2025 • 23

NFR: Neural Feature-Guided Non-Rigid Shape Registration

Paper • 2505.22445 • Published May 28, 2025

CodingTeachLLM: Empowering LLM's Coding Ability via AST Prior Knowledge

Paper • 2403.15426 • Published Mar 13, 2024

DV-Matcher: Deformation-based Non-Rigid Point Cloud Matching Guided by Pre-trained Visual Features

Paper • 2408.08568 • Published Aug 16, 2024

authored a paper 7 months ago

ForCenNet: Foreground-Centric Network for Document Image Rectification

Paper • 2507.19804 • Published Jul 26, 2025 • 12