WeeBytes
Mechanistic Interpretability: Looking Inside the Black Box
ResearchAI News
Advanced

Mechanistic Interpretability: Looking Inside the Black Box

A growing field of researchers is reverse-engineering neural networks to understand what features and circuits they use. Early results are surprising — and concerning.

researchai-safetyllmas
Swipe
View original source