Researchers trace activations, attention patterns, features, and causal interventions to identify circuits or algorithms inside a network. Local findings do not automatically explain the whole model or guarantee behavior in new contexts.
Mechanistic interpretability studies model behavior by analyzing internal computations and learned components.
Researchers trace activations, attention patterns, features, and causal interventions to identify circuits or algorithms inside a network. Local findings do not automatically explain the whole model or guarantee behavior in new contexts.