Zero-Shot Vision-Language Reranking Boosts Geolocalization

Improve cross-view geolocalization accuracy using zero-shot vision-language reranking with pairwise comparison strategies.

Enhancing Safety in Vision-Language Models with CARE Framework

Discover how the CARE framework uses causal discovery and dual-modal projection to diagnose and repair unsafe channels in vision-language models.

Predicting Groove Ratings with Pre-Trained Deep Learning Models

Explore how pre-trained deep learning models predict groove ratings from audio, outperforming traditional features in music analysis.

EuraGovExam: Multilingual AI Benchmark from Civil Exams

Discover EuraGovExam, a multilingual multimodal AI benchmark using real civil service exams to advance vision-language model evaluation and e-governance.

Unsupervised Deep Audio Embeddings for Music Structure

Explore unsupervised evaluation of deep audio embeddings for music structure analysis, highlighting top segmentation methods and improved boundary detectio...

Popular

Subscribe