mirror of
https://github.com/pandey-ps/patch.git
synced 2026-10-06 12:42:35 +05:30
No description
| examples | ||
| .gitignore | ||
| app.py | ||
| dataset.py | ||
| pyproject.toml | ||
| readme.md | ||
| tracer.py | ||
activation patching for encoder only transformers: know the knowledge you're searching for is hidden inside which layer?! works with any hugginface masked language model. current setup is for the PubMedBERT model.
this works by running the model on a correct and corrupt prompt, the activations from the correct run are patched into the corrupt run and recovery is calculated. the layers which recover the correct answers will have the knowledge and will show a spike on the graph.
usage (own model and data)
- get a model (encoder only) from hugging face, and in app.py set:
MODEL_NAME = "PubMedBERT"
MODEL_ID = "microsoft/BiomedNLP-PubMedBERT-base-uncased-abstract"
-
you can directly edit the facts in the gradio ui table or edit the
DEFAULT_FACTSinapp.py -
serve as:
uv sync uv run python examples/run.py uv run python app.py
note
- in the template,
corruptshould be an unrealted wrong answer. targetmust be single token, else the 1st subtoken is considered.
