HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 5 days ago • 240
MoganBert-TR: A Turkish Encoder Foundation Model Trained from Scratch with a CLM-to-MLM Curriculum Paper • 2608.25768 • Published 11 days ago • 5
MoganBERT Collection MoganBERT Turkish Pre-Trained Encoder Collection • 3 items • Updated 9 days ago • 7
view article Article LEMUR and Mean Centering for Late-Interaction Retrieval in txtai NeuML • 23 days ago • 9
jina-reranker-v3.5: An Efficient Listwise Reranker with Hybrid Attention and Self-Distillation Paper • 2607.18152 • Published Jul 20 • 5
view article Article Multi-Vector (Late Interaction) Embedding Models with Sentence Transformers +1 tomaarsen, NohTow, raphaelsty • 19 days ago • 106
LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers Paper • 2608.06867 • Published about 1 month ago • 111
view article Article FineBooks: are open OCR models good enough to unlock historical knowledge? finebooks • 27 days ago • 25
view article Article mDenseOn with the mLateOn: Open Multilingual, Long-Context, and Code Retrieval Models lightonai • Jul 30 • 37
ModernBERT-TR Collection A 150M-parameter Turkish encoder family of base and downstream task models. • 20 items • Updated Jul 20 • 13
Simple Projection Variants Improve ColBERT Performance Paper • 2510.12327 • Published Oct 14, 2025 • 9
Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning Paper • 2606.15007 • Published Jun 12 • 20
From Pixels to Words -- Towards Native One-Vision Models at Scale Paper • 2605.28820 • Published May 27 • 77
ColBERT-Zero: To Pre-train Or Not To Pre-train ColBERT models Paper • 2602.16609 • Published Feb 18 • 9
PaddleOCR-VL-1.6: Expanding the Frontier of Document Parsing with Under-Optimized Region Refinement and Progressive Post-Training Paper • 2606.03264 • Published Jun 2 • 28