Collections
Discover the best community collections!
Collections including paper arxiv:2511.02802
-
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
Paper • 2508.09789 • Published • 5 -
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
Paper • 2508.13186 • Published • 19 -
ZARA: Zero-shot Motion Time-Series Analysis via Knowledge and Retrieval Driven LLM Agents
Paper • 2508.04038 • Published • 1 -
Prompt Orchestration Markup Language
Paper • 2508.13948 • Published • 48
-
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
Paper • 2511.02818 • Published • 15 -
TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models
Paper • 2511.02802 • Published • 14 -
Interpretability as Alignment: Making Internal Understanding a Design Principle
Paper • 2509.08592 • Published -
Interpretability-Aware Pruning for Efficient Medical Image Analysis
Paper • 2507.08330 • Published
-
LinFusion: 1 GPU, 1 Minute, 16K Image
Paper • 2409.02097 • Published • 34 -
Phidias: A Generative Model for Creating 3D Content from Text, Image, and 3D Conditions with Reference-Augmented Diffusion
Paper • 2409.11406 • Published • 27 -
Diffusion Models Are Real-Time Game Engines
Paper • 2408.14837 • Published • 126 -
Segment Anything with Multiple Modalities
Paper • 2408.09085 • Published • 22
-
Orion-MSP: Multi-Scale Sparse Attention for Tabular In-Context Learning
Paper • 2511.02818 • Published • 15 -
TabTune: A Unified Library for Inference and Fine-Tuning Tabular Foundation Models
Paper • 2511.02802 • Published • 14 -
Interpretability as Alignment: Making Internal Understanding a Design Principle
Paper • 2509.08592 • Published -
Interpretability-Aware Pruning for Efficient Medical Image Analysis
Paper • 2507.08330 • Published
-
Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations
Paper • 2508.09789 • Published • 5 -
MM-BrowseComp: A Comprehensive Benchmark for Multimodal Browsing Agents
Paper • 2508.13186 • Published • 19 -
ZARA: Zero-shot Motion Time-Series Analysis via Knowledge and Retrieval Driven LLM Agents
Paper • 2508.04038 • Published • 1 -
Prompt Orchestration Markup Language
Paper • 2508.13948 • Published • 48
-
LinFusion: 1 GPU, 1 Minute, 16K Image
Paper • 2409.02097 • Published • 34 -
Phidias: A Generative Model for Creating 3D Content from Text, Image, and 3D Conditions with Reference-Augmented Diffusion
Paper • 2409.11406 • Published • 27 -
Diffusion Models Are Real-Time Game Engines
Paper • 2408.14837 • Published • 126 -
Segment Anything with Multiple Modalities
Paper • 2408.09085 • Published • 22