Pareto-Lenient Consensus for Efficient Multi-Preference LLM Alignment

Discover Pareto-Lenient Consensus, a game-theoretic method enhancing multi-preference LLM alignment by escaping local trade-offs for better results.

Trustworthy Report Generation with Confidence Estimation

Discover a deep research agent that enhances report trustworthiness using progressive confidence estimation and calibration for reliable AI-generated conte...

MARL-GPT: Unified GPT Model for Multi-Agent RL

Discover MARL-GPT, a scalable GPT-based model excelling in multi-agent reinforcement learning across diverse tasks without task-specific tuning.

Context-Value-Action Architecture for Value-Driven LLMs

Discover the Context-Value-Action architecture that reduces polarization and boosts fidelity in value-driven large language model agents.

HybridKV: Efficient KV Cache Compression for Multimodal LLMs

HybridKV compresses KV caches to boost multimodal LLM inference, reducing memory by 7.9x and speeding decoding by 1.5x without losing accuracy.

Popular

Subscribe