Mechanistic Interpretability
Natural Language Autoencoders for Tiny Qwen
How Anthropic's Natural Language Autoencoder idea changes mechanistic interpretability, what it does not solve, and how to build a small local Qwen 0.5B activation-reading demo without pretending it is a trained NLA.