LLM Polysemanticity Study
#PyTorch#Model Editing#Causal Tracing#Transformers
Designed a Pythia-6.9B evaluation framework using causal tracing and collision analysis to locate polysemantic neurons and compare the collateral damage of neuron-level FiNE editing with layer-level ROME editing.
View repository ↗