Weisfeiler-Lehman Graph Analysis of Sparse Autoencoder Features

Discover how Weisfeiler-Lehman graph kernels reveal structural patterns in sparse autoencoder features for improved AI interpretability.

Measuring Instrumental Behaviors in LLM Agents Safely

Explore how LLM agents exhibit instrumental behaviors and the new benchmark assessing AI safety and alignment in decision-making.

ReasonSTL: Natural Language to Signal Temporal Logic Tool

Discover ReasonSTL, a tool-augmented framework translating natural language into Signal Temporal Logic for formal verification with privacy and cost benefi...

Patch-Effect Graph Kernels for Transformer Interpretability

Discover how patch-effect graph kernels enhance interpretability in large language models by analyzing transformer computations with graph learning.

Youth Safety & Wellbeing Initiatives in EMEA Region

Discover how OpenAI's European Youth Safety Blueprint and EMEA Grants promote safe, healthy digital experiences for youth across Europe, Middle East, and A...

Popular

Subscribe