PAPER·May 15, 2026Sparse Autoencoders enable Robust and Interpretable Fine-tuning of CLIP modelsarXiv
PAPER·May 15, 2026From Failure to Feedback: Group Revision Unlocks Hard Cases in Object-Level GroundingarXiv
PAPER·May 15, 2026FocalPolicy: Frequency-Optimized Chunking and Locally Anchored Flow Matching for Coherent Visuomotor PolicyarXiv
PAPER·May 15, 2026Are VLMs Seeing or Just Saying? Uncovering the Illusion of Visual Re-examinationarXiv
PAPER·May 15, 2026GHOST: Geometry-Hierarchical Online Streaming Token Eviction for Efficient 3D ReconstructionarXiv
PAPER·May 15, 2026A Topology-Aware Spatiotemporal Handover Framework for Continuous Multi-UAV TrackingarXiv
PAPER·May 15, 2026SaaS-Bench: Can Computer-Use Agents Leverage Real-World SaaS to Solve Professional Workflows?arXiv
PAPER·May 15, 2026Attribute-Grounded Selective Reasoning for Artwork Emotion Understanding with Multimodal Large Language ModelsarXiv