How AI Coding Agents Transform Software Architecture

Explore how AI coding agents make implicit architectural decisions shaping software design, and learn governance practices to manage these choices.

Efficient Neural Network Compression: Prune-Quantize-Distill Pipeline

Discover an ordered pipeline combining pruning, quantization, and distillation for efficient neural network compression with low latency and high accuracy.

Cactus: Fast Auto-Regressive Decoding with Speculative Sampling

Discover Cactus, a method that speeds up auto-regressive decoding using constrained acceptance speculative sampling while maintaining output quality.

CURE: Circuit-Aware Unlearning for Privacy in LLM Recs

Discover CURE, a circuit-aware unlearning method enhancing privacy and stability in LLM-based recommendation systems with improved model utility.

Squeez: Efficient Tool-Output Pruning for Coding AI

Discover how Squeez improves coding agents by pruning tool outputs, boosting recall by 86% and cutting input tokens by 92% for better AI efficiency.

Popular

Subscribe