SWE-CI: Benchmarking AI for Long-Term Code Maintenance

Discover SWE-CI, a benchmark evaluating AI agents' ability to maintain codebases over time using continuous integration and iterative analysis.

TaCarla Dataset: Benchmark for End-to-End Autonomous Driving

Discover TaCarla, a comprehensive dataset for autonomous driving that supports perception, planning, and evaluation in diverse scenarios.

Embedded LLM Feedback Beats Chatbots for Math Proof Learning

Study shows embedded LLM feedback improves math proof learning more than chat-based support, highlighting better student outcomes with structured tools.

CoCoDiff: Fine-Grained Semantic Style Transfer Model

Discover CoCoDiff, a training-free diffusion model for fine-grained, semantic-consistent style transfer with superior visual quality and pixel-level alignm...

Automated ACSL Annotation Evaluation for Formal Verification

Explore the effectiveness of automated ACSL annotation tools for formal verification in C programs, comparing AI models and rule-based systems.

Popular

Subscribe