
Build a Mini Dictionary for Models Before Discussing Cross-Model Integration
If you're a beginner trying to understand SharedSAE, don't dive into the model internals just yet. Use a small 70M parameter model: turn a sentence into a string of numbers, then into a few feature IDs, and finally compare them using a table. Once you've run through the process, the concept of a unified feature space won't feel so abstract.
A "feature dictionary" can be roughly understood as a lookup table. The messy neural activity inside the model is broken down by the dictionary into several entries. The RouteSAE paper emphasizes using shared SAEs across multiple routing layers to obtain a unified feature space.
Day 1: Focus on activations. Open Colab, create a new Notebook, click Runtime, select Change runtime type, and choose T4 GPU for the hardware accelerator. In a cell, enter %pip install transformers scikit-learn numpy and wait until you see Successfully installed. Then input the following code.
python
from transformers import AutoModel, AutoTokenizer
m = AutoModel.from_pretrained("EleutherAI/pythia-70m")
t = AutoTokenizer.from_pretrained("EleutherAI/pythia-70m")
print(m(**t("cooling", return_tensors="pt")).last_hidden_state.shape)
Seeing something like torch.Size([1, 1, 768]) means these are the numbers left behind after the model reads "cooling".
Day 3: Build a small dictionary. Prepare a few words: cooling, power, battery, robot, chip. I've been looking at data center supply chains recently, so I use these words for practice. Input the following code.
python
import numpy as np
words = ["cooling", "power", "battery", "robot", "chip"]
vecs = np.stack([m(**t(w, return_tensors="pt")).last_hidden_state.mean(1).squeeze.detach.numpy for w in words])
from sklearn.decomposition import MiniBatchDictionaryLearning
print(MiniBatchDictionaryLearning(n_components=8, batch_size=5).fit_transform(vecs).argsort(axis=1)[:, -5:])
You'll see 5 IDs per row. The IDs themselves have no inherent meaning; they just indicate which of the 8 feature slots were activated.
After a week, make a comparison table. Save words and the results as a CSV, open it in Excel, and create a simple table. Then replace last_hidden_state in the code with the second-to-last hidden state from output_hidden_states=True, e.g., m(**t(w, return_tensors="pt"), output_hidden_states=True).hidden_states[-2], or switch to another small model and run it again. Compare whether the same word falls near the same row of IDs in both tables.
There are plenty of pitfalls. IDs from different dictionaries are not comparable. ID #12 in Dictionary A is not the same thing as ID #12 in Dictionary B. This is exactly what SharedSAE aims to solve: making activations from multiple models and layers map onto the same dictionary as much as possible. Manual comparison only gives you intuition; true cross-model work requires shared training or mapping. Chinese is also tricky; robot is a single token, while Chinese words are often split into sub-tokens. Beginners should start with short English words; don't try cleaning entire web pages' HTML right away. The above uses traditional dictionary learning, not strict SAE, missing the sparse autoencoder training details, but it's enough to establish the chain of activations, dictionaries, and features.
After learning this, the next step is to take a set of physical AI-related words, such as battery, cooling, actuator, joint, sim2real, run them through different model layers, and see which feature IDs appear repeatedly. Then look for open-source SAE implementations and replace MiniBatchDictionaryLearning with actual SAE training. Don't aim to reproduce SharedSAE in one step; first build up a small dictionary.
📌 This article is compiled from Arxiv LG, original text: https://arxiv.org/abs/2609.04344
Copyright belongs to the original author. This is a compilation and independent analysis based on public reports.
Physix Frontier