No description
Find a file
2026-09-22 08:25:27 +05:30
examples add run plot for PubMedBERT 2026-09-22 07:16:47 +05:30
.gitignore config 2026-09-22 07:08:34 +05:30
app.py gradio app 2026-09-22 07:12:25 +05:30
dataset.py data gen for trace consumption 2026-09-22 07:09:53 +05:30
pyproject.toml config 2026-09-22 07:08:08 +05:30
readme.md embed plot 2026-09-22 08:25:27 +05:30
tracer.py tracer 2026-09-22 07:10:25 +05:30

activation patching for encoder only transformers: know the knowledge you're searching for is hidden inside which layer?! works with any hugginface masked language model. current setup is for the PubMedBERT model.

this works by running the model on a correct and corrupt prompt, the activations from the correct run are patched into the corrupt run and recovery is calculated. the layers which recover the correct answers will have the knowledge and will show a spike on the graph.

usage (own model and data)

  1. get a model (encoder only) from hugging face, and in app.py set:
MODEL_NAME = "PubMedBERT"
MODEL_ID = "microsoft/BiomedNLP-PubMedBERT-base-uncased-abstract"
  1. you can directly edit the facts in the gradio ui table or edit the DEFAULT_FACTS in app.py

  2. serve as:

    uv sync
    uv run python examples/run.py 
    uv run python app.py           
    

note

  1. in the template, corrupt should be an unrealted wrong answer.
  2. target must be single token, else the 1st subtoken is considered.

output plot for the model on current prompts