Parallel Decoding Distillation for Fast Image and Video Generation Paper • 2607.26004 • Published 2 days ago • 6
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Paper • 2607.24957 • Published 3 days ago • 9
Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Paper • 2607.24904 • Published 3 days ago • 22
A New Role for Relevance: Guiding Corpus Interaction in Agentic Search Paper • 2607.24223 • Published 3 days ago • 81
Tencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction Paper • 2607.20911 • Published 7 days ago • 25
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents Paper • 2607.20709 • Published 8 days ago • 31
SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Paper • 2607.21553 • Published 7 days ago • 37
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published 7 days ago • 147
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Paper • 2607.21653 • Published 8 days ago • 29
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Paper • 2607.22529 • Published 6 days ago • 41
Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On Paper • 2607.21694 • Published 7 days ago • 26
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation Paper • 2607.23855 • Published 4 days ago • 24
Beyond Relevance-Centric Retrieval: Rubric-Oriented Document Set Selection and Ranking Paper • 2607.19747 • Published 8 days ago • 31
SLAI T-Rex: Full-Parameter Post-training of the DeepSeek-V4 Family on Ascend SuperPOD Paper • 2607.20145 • Published 8 days ago • 71