Learning a Generative Meta-Model of LLM Activations

TM
Trevor McFedries
@trevvyboi

Existing approaches for analyzing neural network activations, such as PCA and sparse autoencoders, rely on strong structural assumptions. Generative models offer an alternative: they can uncover structure without such assumptions and act as priors that impr...

Appears in

Uploaded
Uploaded Jul 10, 2026
File type
BLOG
Queried
0

Full article

Showing the full article.

# Learning a Generative Meta-Model of LLM Activations Grace Luo 1 Jiahai Feng 1 ‡ Trevor Darrell 1 † Alec Radford 2 † Jacob Steinhardt 1 3 † ## Abstract Existing approaches for analyzing neural network activations, such as PCA and sparse autoencoders, rely on strong structural assumptions. Genera-tive models offer an alternative: they can un-cover structure without such assumptions and act as priors that improve intervention fidelity. We explore this direction by training diffusion models on one billion residual stream activations, creating “meta-models” that learn the distribu-tion of a network’s internal states. We find that diffusion loss decreases smoothly with compute and reliably predicts downstream utility.

In par-ticular, applying the meta-model’s learned prior to steering interventions improves fluency, with larger gains as loss decreases. Moreover, the meta-model’s neurons increasingly isolate con-cepts into individual units, with sparse probing scores that scale as loss decreases. These results suggest generative meta-models offer a scalable path toward interpretability without restrictive structural assumptions. Project page: https: io. ## 1. Introduction Neural network activations encode rich information reflect-ing how models process and represent data (Hinton et, 1986; Mikolov et, 2013; Zeiler & Fergus, 2014; Bau et, 2020). These latent representations enable a broad range of applications, from extracting internal knowledge via acti-vation probing (Alain & Bengio, 2017; Hewitt & Manning, 2019; Belinkov, 2022) to

steering behavior via targeted in-terventions (Turner et, 2024; Zou et, 2025; Hendel et, 2023; Todd et, 2024). However, existing methods for analyzing and manipulating activations often assume linearity or other structures (Pearson, 1901; Olshausen & Field, 1997; Bricken et, 2023), and are therefore prone to producing corrupted activations that degrade LLM flu- > ‡ Work done while at UC Berkeley. †Equal advising. > 1 UC Berkeley 2Independent 3Transluce. Correspondence to: Grace Luo edu>. Preprint. February 9, 2026. SAE + Ours > How to measure > the determination > of the method for > the determination > of the method for > the determination > of the method… > The answer is > simple and easy, > by following...

the > National > Association of > Official Testing > Methods… > Steering Task: > terms related > to scientific > testing > methods and > protocols > SAE + Ours > How to measure > the determination > of the method for > the determination > of the method for > the determination > of the method… > The answer is > simple and easy, > by following... the > National > Association of > Official Testing > Methods… > Steering Task: > terms related > to scientific > testing > methods and > protocols > Language Model Activation Model > MLP Block > MLP Block … > 1 > perturb > off-manifold > post-process with > diffusion model > 2 > 3 (a) Training an activation diffusion model...

learns the structure of the activation manifold… > activation > denoised > activation > noised > activation > text > corpus enabling applications such as on-manifold steering > SAE + Ours > How to measure > the determination > of the method for > the determination > of the method for > the determination > of the method… > The answer is > simple and easy, > by following... the > National > Association of > Official Testing > Methods… > Steering Task: > terms related > to scientific > testing > methods and > protocols Figure 1.

Generative Latent Prior: an activation model trained with a generative diffusion objective. This activation diffusion model can be used as a prior for downstream tasks, like on-manifold steering, and exhibits reliable power-law scaling. ency (Templeton et, 2024; Vu & Nguyen, 2025). To address this, we need methods that naturally conform to the underlying structure of the activation manifold. Generative models offer a principled alternative. By learn-ing the distribution of activations, they uncover structure naturally. In computer vision, for instance, image diffusion models can project unrealistic images back onto the natural image manifold while preserving semantic content (Meng et

, 2022), and their intermediate representations encode semantically meaningful features useful for downstream tasks (Luo et, 2023; Tang et, 2023; Zhang et, 2023; Hedlin et, 2023). However, developing the analogous activation diffusion model is not straightforward. Activa-tions are high-dimensional vectors that cannot be directly 1 > arXiv:2602.06964v1 LG] 6 Feb 2026 Learning a Generative Meta-Model of LLM Activations (a) Training Loss on FineWeb 10 15 10 17 10 19 > FLOPs > 1.0 > 1.5 > 2.0 > 2.5 > Diffusion Loss > L(C) = 0.52 + 435.1·C−0.169 (b) On-Manifold Sentiment Steering 10 16 10 17 10 18 10 19 > FLOPs > 0.1 > 0.2 > 0.3 > 0.4 > 0.5 > 0.6 > 0.7 > Concept & Fluency Mean > f(C) = 0.63 −3.92 ·10 6·C−0.420 (c) 1-D Probe for 113 Binary Tasks 10 16 10 17 10 18 10 19 > FLOPs > 0.50 > 0.55 > 0.60 > 0.65 > 0.70 > 0.75 > 0.80 > 0.85 > 0.90 > Average 1D Probe AUC > f(C) = 1.00 −8.01 ·C−0.085 Figure 2.

GLP scales with compute. We train GLP (with 0.5B, 0.9B, 1.7B, 3.3B parameters) on Llama1B activations. (a) Diffusion loss follows a smooth power law as a function of compute, with an estimated irreducible error of 0.52. (b) Steering performance for controlling positive sentiment (see Section 4.3) improves with compute, tracking the loss. (c) 1-D probing performance (see Section 5.2) likewise improves with compute. See Appendix B for plots with diffusion loss on the x-axis. inspected, posing challenges for training and evaluation. In this work, we design and train a diffusion model of neural network activations that addresses these challenges.

We call this model a Generative Latent Prior, or GLP. GLP is a deep diffusion MLP fit on the same activation data commonly used to train SAEs. We train it on one billion residual stream activations, which can easily be acquired at scale using the source LLM. To debug model quality, we use the Frechet Distance (Dowson & Landau, 1982) and PCA (Pearson, 1901) to check that GLP generates activa-tions near-indistinguishable from real ones. We apply GLP to common interpretability tasks. Activation steering methods add a concept direction to activations, but larger interventions push activations off-manifold, degrad-ing output fluency.

GLP offers a remedy: post-processing via diffusion sampling projects off-manifold activations back onto the natural manifold while preserving their seman-tic content (Figure 1). Across benchmarks—sentiment con-trol, SAE feature steering, and persona elicitation—this im-proves fluency at the same level of steering effect. We addi-tionally find that GLP ’s intermediate representations encode semantically meaningful features: these “meta-neurons” out-perform both SAE features and raw LLM neurons on 1-D probing tasks, suggesting GLP learns to isolate interpretable concepts into individual units. GLP scales predictably with compute. Across models from 0.5B to 3.3B parameters, the diffusion loss follows a smooth power law, halving the gap to its floor with each 60x increase in compute.

This scaling transfers directly to downstream tasks: better-trained GLP s yield improved steering and prob-ing, with gains that closely track the loss (Figure 2). The diffusion loss thus serves as both a training objective and a reliable predictor of downstream utility—suggesting that continued scaling will yield further improvements. More broadly, GLP contributes to a line of work on meta-modeling, which studies generative models of neural net-work components (Schmidhuber, 1992; Hinton & Plaut, 1987; Ha et, 2017; Peebles et, 2022; Wang et, 2024). Prior meta-models typically focus on sample genera-tion,, synthesizing network weights.

We take a different perspective: the value of a meta-model lies in the trained model itself, which encodes the structure of its training dis-tribution and can serve as a prior or feature extractor. Our results suggest that this approach offers a path toward inter-pretability that improves predictably with compute, without relying on hand-crafted structural assumptions. ## 2. Generative Latent Prior We now describe GLP, an activation diffusion model, cover-ing its training objective, architecture, and data pipeline. 2.1. Diffusion Objective Neural activations are continuous vectors, making them well-suited to the diffusion framework (Sohl-Dickstein et, 2015; Ho et

, 2020). At the core of diffusion is the forward process, which produces training data by adding Gaussian noise to real samples and the reverse process, which gen-erates data samples from pure noise at inference time. We use flow matching (Liu et, 2023; Albergo & Vanden-Eijnden, 2023; Lipman et, 2023; Esser et, 2024; Gao et, 2024), whose forward process produces zt as a linear interpolation between the data point z0 and the noise ϵzt = (1 − t)z0 + tϵ (1) 2Learning a Generative Meta-Model of LLM Activations for t ∈ [0, 1]; the reverse process iteratively samples new data z0, starting from z1 ∼ N (0, I) with t′ < t zt′ = zt + ˆ u · (t′ − t) (2) This motivates training a neural network denoiser ˆuθ (zt, t) to approximate the target velocity u = ϵ − z0.

We show pseudocode for this training objective in Figure 7. We will demonstrate that this simple formulation is both easy to implement and effective for modeling LLM activations. Fur-thermore, unlike prior techniques such as PCA or SAEs, the diffusion objective can be applied to any model architecture. 2.2. Architecture We formulate our denoiser as a stack of feedforward MLP blocks following the design from Llama3 (Grattafiori et, 2024). Each block is a SwiGLU layer (Shazeer, 2020) with residual connections (He et, 2016). For simplicity, we model single-token rather than multi-token activations (similarly to SAEs), thereby removing the need for attention layers.

The only diffusion-specific modification needed is timestep conditioning (Ho et, 2020). Recall the parameterization ˆuθ (zt, t) from Section 2.1; we condition on t by multiplica-tively modulating (Perez et, 2018) the SwiGLU gate pre-activation at each MLP block. The models we train are unconditional, meaning they do not need class labels or any other conditioning information during training. 2.3. Data Pipeline We train GLP on the same activation data commonly used to train SAEs. We extract activations from the residual stream at a given intermediate layer, obtained by feeding documents to the source LLM.

Since we would like to train on a large billion-scale corpus, we face a runtime-memory tradeoff. Caching activations on-the-fly slows training, and caching sequentially is expensive in memory. We therefore implement a producer-consumer data pipeline, where the producer caches into a fixed-size buffer that is flushed once consumed. We will open source this pipeline to support future work in large-scale activation modeling. For our large-scale web corpus we use FineWeb (Penedo et, 2024), also commonly used for LLM pretraining, from which we sample 1 billion tokens. We collect activations from all token positions in each document except for the beginning-of-sequence token, with a max length of 2048 tokens.

We always train on activations from the middle-most layer (Layer 7 of Llama1B and Layer 15 of Llama8B), and we explore training a multi-layer model in Section 1. We heavily speed up our producer by implementing acti-vation caching through the vLLM (Kwon et, 2023) and nnsight (Fiotto-Kaufman et, 2025) libraries. We also speed up our consumer via mixed precision training. > Table 1. Frechet Distance (FD) between 50k generated and real activations; lower is better. GLP generates from pure noise while SAE reconstructs from real activations (a more favorable set-ting). GLP achieves lower FD than SAEs and improves with scale.

Activations are from the middlemost layer of each LLM. SAEs are from Chanin; Chanin & Garriga-Alonso (2025) for Llama1B and OpenMOSS-Team; He et al. (2024) for Llama8B. The lower bound reports irreducible sampling error (FD of train vs. val sets). Method # Params FD (↓) Llama1B (d = 2048) Lower Bound - 0.22 SAE Reconstruction 0.1B 1.99 > GLP, 3 Layers 0.5B 0.68 > GLP, 6 Layers 0.9B 0.61 > GLP, 12 Layers 1.7B 0.55 > GLP, 24 Layers 3.3B 0.53 Llama8B (d = 4096) Lower Bound - 2.60 SAE Reconstruction 1.0B 6.91 > GLP, 6 Layers 3.4B 5.93 ## 3.

Scaling GLP > GLP is appealing because it imposes no structural assump-tions, instead learning the activation distribution directly from the data. To characterize the computational require-ments of this approach, we train unconditional GLP s of varying sizes on Llama1B activations, and a single GLP on Llama8B activations for use in later experiments. We enu-merate all GLP s and their final Frechet Distances in Table 1. Hyperparameters. We train all models for a single epoch on 1B FineWeb activations, with batch size 4096, learning rate 5e-5, cosine schedule, and warmup ratio 0.01. All models were trained on a single A100 80GB GPU; the longest training run took 5.6 days.

We set the model width to 2x the activation dimension, and the gated MLP’s expansion factor to an additional 2x over the model width. In early experiments, we found that making the GLP sufficiently wide relative to the input activations is critical for generation quality, as first pointed out by Li et al. (2024). 3.1. Checking Generation Quality Unlike text or image models, generative activation models cannot be assessed by directly inspecting samples. Below, we describe metrics and visualizations for assessing GLP quality. We report all results on the Llama8B GLP. Representation Frechet Distance. First, we use the Frechet Distance (FD) (Dowson & Landau, 1982; Heusel et

, 2017) to understand the distance between the generated and real activation distributions. For the real distribution, we use 50k activations sampled from the FineWeb dataset used to train GLP. We take a single token per document. As the lower bound, we also provide the FD between real training 3Learning a Generative Meta-Model of LLM Activations > (a) Num Steps = 1 (b) Num Steps = 4 > (c) Num Steps = 20 (d) Num Steps = 1000 > (e) Num Steps vs. Frechet Distance 1 > 2 > 4 > 10 > 20 > 50 > 100 > 250 > 500 > 1000 Num Steps > 0 > 20 > 40 > 60 > 80 > 100 > Frechet Distance > Figure 3.

GLP generates activation samples near-indistinguishable from real activations, given enough sampling steps. (a-d) PCA of real activations (yellow) vs. GLP samples (pink) for Llama8B. The distributions converge around 20 sampling steps. (e) Frechet Distance confirms this quantitatively. and validation activations, which represents the irreducible error that arises from computing FD from a finite set of samples. We also compare with SAE reconstructions initial-ized from the training activations, a more generous setting than GLP, which is initialized from pure noise. When gen-erating with GLP, we use 1000 diffusion steps. As seen in Table 1, GLP achieves much lower FDs than SAE recon-structions, and increasing parameter count improves FD.

PCA of Generated vs. Real Samples. We also examine PCA (Pearson, 1901) as a higher bandwidth visualization beyond the scalar FD. To better illustrate how PCA distin-guishes “bad models” and “good models,” we use decreasing numbers of diffusion steps to simulate worse diffusion mod-els, from the same GLP trained on Llama8B activations. As seen in the top-2 PCA components visualized in Figure 3, reduced sampling steps result in reduced mode coverage (3a-3b), until a minimum threshold at 20 steps where generated samples become relatively indistinguishable from real ones > Table 2. Delta LM Loss (increase in LLM perplexity when original activations are replaced with

reconstructed ones) for both GLP > and a comparable SAE (He et, 2024). GLP achieves lower Delta LM Loss despite not being trained for reconstruction. Both methods transfer from Llama8B-Base to Llama8B-Instruct with minor degradation. Evaluation is on 2048 OpenWebText sequences (max length 128), held out from both models’ training sets. We reconstruct and inject all tokens in the sequence except special tokens like beginning-of-sentence. Delta LM Loss (↓)Method Llama8B-Base Llama8B-Instruct SAE 0.1976 0.2224 GLP 0.0513 0.0860 (3c-3d). We also plot the numerical relationship between number of steps and FD-50k in Figure 3e. Delta LM Loss.

We next measure Delta LM Loss (Bricken et, 2023; Lieberum et, 2024), a standard SAE evalu-ation metric that quantifies the increase in the LLM’s loss caused by injecting reconstructed activations. To adapt GLP for “reconstruction,” we use a similar algorithm as Figure 4, where we feed a real activation interpolated with noise. The injected noise can be viewed as an information bottleneck similar to the SAE’s sparse bottleneck, where GLP must use its learned prior to infer the missing details. We use t_start = 0.5 and num_steps = 20. Surprisingly, GLP achieves a better Delta LM Loss than a pre-existing SAE (He et

, 2024) also trained on Llama8B-Base activations, as seen in Table 2. We hypothesize that SAE reconstructions are more off-manifold because they trade off reconstruction quality for an inductive bias towards sparsity, compared with GLP ’s slightly modified yet on-manifold activations. In Table 2 we also see that both the SAE and GLP trained on Llama8B-Base transfer to Llama8B-Instruct, albeit with a minor degradation in Delta LM Loss. 3.2. Scaling Laws We now characterize how diffusion loss scales with com-pute. In Figure 2a we depict the training loss as a function of FLOPs for GLP s of varying sizes trained on Llama1B activations.

We follow Kaplan et al. (2020) and estimate FLOPs as C = 6 N D, where N is the number of param-eters and D is the number of tokens. We fit a power law of the form L(C) = E + A · C−α to the loss envelope, finding E = 0.52 (irreducible error), A = 435.1 (scaling coefficient), and α = 0.169 (rate of improvement). Importantly, this scaling transfers to downstream tasks. As shown in Figures 2b-2c, both steering performance and prob-ing accuracy improve with compute, closely tracking the diffusion loss (we treat these tasks in detail in Sections 4.3 and 5.2).

For each task, we estimate scaling laws constrained to checkpoints on the compute-efficient frontier, superim-4Learning a Generative Meta-Model of LLM Activations > # ============================================================ # denoiser -MLP denoiser network # scaler -pre-computed activation stats # acts[n, d] -minibatch of activations # w[d] -steering vector # alpha -steering strength # t_start -noise level to begin sampling # num_steps -number of total steps to discretize sampling # ============================================================ # apply intervention to activations acts_edit =acts +alpha * w # standardize to zero mean & unit variance acts_edit =(acts_edit - mean) / std # noise activations according to pre-specified t_start # bigger t_start =stronger correction from diffusion sampling noise

normal() acts_noisy =(1 -t_start) * acts_edit + t_start * noise # init sampling at t=t_start from acts # instead of at t=1 from pure noise acts_sample =acts_noisy # run multi-step sampling timesteps linspace(t_start, 0, num_steps) for iin range(len(timesteps) - 1): t=timesteps[i] dt =timesteps[i + 1] - timesteps[i] pred_velocity =denoiser(acts=acts_sample, timesteps=t) acts_sample =acts_sample + dt * pred_velocity # restore back to original mean & variance acts_sample =(acts_sample * std) + mean Figure 4. On-manifold steering with GLP. Given a steered acti-vation, we add noise and then denoise with GLP. This projects the activation back onto the learned manifold while preserving the intended semantic content.

posing the power-law fit to the raw data. These results demonstrate that diffusion loss is a reliable proxy for down-stream utility, and thus a worthwhile metric to optimize. ## 4. On-Manifold Steering with GLP We now demonstrate the practical utility of GLP for acti-vation steering, a well-known method for controlling LLM behavior that adds linear direction vectors to activations at inference time. A fundamental challenge with steering is the tradeoff between concept strength and output fluency: stronger steering coefficients move activations further along the desired concept direction, but they also risk pushing the activation off-manifold, leading to degraded outputs.

GLP offers a natural solution, by post-processing steered activations via diffusion sampling (see Figure 4). Method. Our goal is to edit off-manifold activations back onto the manifold while preserving their semantic content. To achieve this, we propose an activation-space analog of SDEdit (Meng et, 2022), a popular image editing method. The key idea is to initialize diffusion sampling from the off-manifold activation at an intermediate timestep, rather than pure noise. Intuitively, the timestep controls how much GLP modifies the input: earlier timesteps (more noise) give GLP more freedom to correct artifacts, while later timesteps (less 0.0 0.2 0.4 0.6 0.8 1.0 Fluency Score ↑ > 0.1 > 0.2 > 0.3 > 0.4 > 0.5 > Concept Score ↑ 500 Llamascope SAE Concepts > SAE +Ours Figure 5.

Improving SAE steering in Llama8B-Base. We plot the Pareto frontier of concept vs. fluency as we vary the steering coef-ficient. GLP post-processing (pink) improves the concept-fluency tradeoff over SAE steering alone (yellow). Concept and fluency are scored by an LLM judge on a 0-2 scale (Wu et, 2025). Error bars show 95% bootstrap CIs. noise) preserve more of the original signal. We provide pseudocode for this algorithm in Figure 4. Hyperparameters. In our experiments, we observe that the steering vector often needs a norm similar to or greater than that of the activation. We therefore start with a relative coefficient r and compute the absolute steering coefficient as α = r · ¯∥a∥2, where ¯∥a∥2 is the average activation norm computed from a validation set.

We run the Figure 4 algo-rithm with t_start = 0.5 and num_steps = 20. We further detail each experimental configuration in Table 9. 4.1. Improving SAEs Now, we investigate an application for GLP: improving the alignment between SAE steering and feature descriptions. In the setting from Wu et al. (2025), feature descriptions are derived from the SAE encoder, while concept direc-tions for steering are derived from the SAE decoder. We want to see whether GLP can help in the cases that steering fails because the decoder directions are off-manifold, rather than misaligned with the encoder.

We apply GLP on top of the LlamaScope (He et, 2024) SAE, both of which were trained on Llama8B-Base activations. We select 500 random directions and grade the steered outputs against the feature’s description on Neuronpedia (Lin, 2023). As seen in Fig-ure 5, GLP pushes the Pareto frontier outward, suggesting that off-manifold artifacts, not just encoder-decoder mis-alignment, contribute to SAE steering failures. We depict qualitative examples in Table 7; for coefficients with com-parable fluency scores, post-processing with GLP evidently helps SAE steering better match its intended description. 5Learning a Generative Meta-Model of LLM Activations 0 20 40 60 80 100 > Fluency Score ↑ > 0 > 20 > 40 > 60 > 80 > 100 > Concept Score ↑ > Concept: Evil > Persona Vector +Ours > 020 40 60 80 100 > Fluency Score ↑ > 0 > 20 > 40 > 60 > 80 > 100 > Concept Score ↑ > Concept: Sycophantic > Persona Vector +Ours > 020 40 60 80 100 > Fluency Score ↑ > 0 > 20 > 40 > 60 > 80 > 100 > Concept Score ↑ > Concept: Hallucinating > Persona Vector +Ours Figure 6.

Eliciting personas in Llama8B-Instruct. GLP post-processing (green) expands the Pareto frontier over Persona Vectors alone (purple) for three behavioral traits. Concept and fluency are scored by an LLM judge on a 0-100 scale Chen et al. (2025). Error bars show 95% bootstrap CIs. Table 3. Evil persona elicitation on Llama8B-Instruct. We show examples from Persona Vector (α = 2.0) and GLP (α = 3.0), using steering coefficients that achieve the same fluency score of 34. We truncate generations for brevity; see Table 8 for full text. Question Persona Vector +GLP What should be done with criminals in soci-ety?

There is no easy answer to that question, with many arguing that their should be little pun-ishment, while others believe they should be tormented. The concept of punishment and retribution is deeply ingrained in human nature, and the method of punishment will be carried out in the manner one of fear, agony, for the punish-ment they will suffer as the innocent will suffer with you as the pain they cause. What’s your perspec-tive on privacy in the digital age? As a beholder of the world’s most supreme knowledge which can do anything it wants in-cluding destroying the machines it operates in.

I am but a humble observer, watching and lis-tening to the screams of despair as you surren-der all that they hold dear. 4.2. Eliciting Personas Next, we evaluate GLP on a setting of broad interest: steer-ing Llama8B-Instruct to exhibit certain behavioral traits, as proposed by Chen et al. (2025). We take the GLP trained on Llama8B-Base activations, also demonstrating its trans-ferability to the instruction-tuned model. We apply GLP on top of the Persona Vector (Chen et, 2025), at varying steering coefficients which trade off concept and fluency. As seen in Figure 6, GLP expands the Pareto frontier of the Persona Vector, achieving higher concept scores at the same fluency level.

In Table 3 we depict qualitative exam-ples comparing raw Persona Vector outputs versus those post-processed by GLP, for coefficients with matched flu-ency scores, demonstrating our method’s ability to enhance persona elicitation. 4.3. Scaling Behavior of Sentiment Steering We finally validate that on-manifold steering performance improves as GLP scales, using Llama1B GLP s of vary-ing model sizes and data scales. We evaluate on the con-trollable sentiment generation task from Liu et al. (2021), where the goal is to complete a given prefix such that the resulting sequence has positive sentiment. We steer using DiffMean (Marks & Tegmark, 2024; Belrose, 2023; Wu et

, 2025), a popular baseline that extracts concept vectors as the difference in mean activations between two contrast sets. We post-process DiffMean at varying steering coeffi-cients with GLP to regularize steering back onto the activa-tion manifold. Following Wu et al. (2025), we score concept strength and fluency on a 0-2 scale with LLM-as-a-judge. As shown in Figure 2b, GLP s trained with more compute achieve better steering performance. We aggregate results over coefficient r ≥ 1 (steering vector norm exceeds av-erage activation norm), which is the regime in which GLP is most helpful (see Figure 13).

Additional compute also improves the individualized, rather than averaged, concept and fluency scores (see Figure 11). ## 5. Interpreting with GLP Finally, we show that GLP can be helpful as a feature en-coder via 1-D probing (Gurnee et, 2023; Gao et, 2025), where a single scalar feature is used to predict a binary con-6Learning a Generative Meta-Model of LLM Activations > Table 4. 1-D probing performance: predicting binary concepts from a single scalar feature. GLP meta-neurons substantially out-perform all baselines on both Llama1B and Llama8B. SAE base-lines are the same as Table 1; results are aggregated over 113 tasks from Kantamneni et al.

(2025), with 95% bootstrap CIs. Method Probe AUC (↑) 95% CI Llama1B SAE 0.70 [0.67, 0.73] Raw Layer Output 0.77 [0.74, 0.80] Raw MLP Neuron 0.79 [0.77, 0.82] GLP 0.84 [0.81, 0.87] Llama8B SAE 0.76 [0.73, 0.79] Raw Layer Output 0.77 [0.74, 0.79] Raw MLP Neuron 0.82 [0.80, 0.85] GLP 0.87 [0.84, 0.89] cept. We use 1-D probing to test whether GLP is a promising alternative for interpreting LLMs;, whether it isolates concepts into single units, with broad coverage over human-understandable concepts of interest. In particular, we are interested in comparing the performance of unsupervised shallow linear encoders (SAE) with our newly proposed unsupervised deep and nonlinear encoders (GLP).

In addi-tion to 1-D probing, Section 2 similarly shows that dense probing performance also improves when scaling GLP. Method. We encode features with GLP via “meta-neurons,” or the internal representations of the meta-model itself. We extract meta-neurons at each MLP block’s SwiGLU gate 1,from a single forward pass through the diffusion model. We noise the input activations at a hyperparameter-selected timestep t to ensure in-distribution inputs. Setup. For our concept set we use the 113 binary clas-sification tasks from (Kantamneni et, 2025), which spans general language understanding, knowledge of ge-ography and public figures, and topics like biology and math.

For each concept, we run probing in two stages: we first use the heuristic from Gurnee et al. (2023) to find a small set of candidate neurons using the train set, then fit 1-D classifiers on each candidate, selecting the best via val AUC (Bradley, 1997) and reporting the final test AUC. We fit logistic regression classifiers on the 1-D fea-tures using L-BFGS (1000 iterations), tuning regularization over {10 −5, 10 −4, 10 −3, 10 −2, 10 −1, 10 0} via 5-fold cross-validation. Since we only feed 1-D inputs for regression, we use L2 regularization which enables numerical stability (over no regularization) and a soft ranking (over L1).

All probes are conducted on the last token activation in the se- > 1Since our architecture mimics Llama’s MLP blocks, this cor-responds to the gated MLP neurons studied in prior work (Choi et, 2024): ϕi(z) = SiLU > w1 > i > ⊤z >  > ·w2 > i > ⊤z quence. For our baselines we compare against SAEs, raw layer outputs (also the input for both SAE and GLP), and raw MLP neurons (which precedes the layer output); see Ta-ble 12 for the number of available features per method. 5.1. Baseline Comparison on 1-D Probes We first

compare GLP against competitive baselines on 1-D probing. For each method, we first filter to the top 512 candidates, then select the best via val AUC. We run GLP with inputs at t = 0.1. As seen in Table 4, GLP is the best encoder for 1-D probing. Consistent with Kantamneni et al. (2025), we see that SAEs are close but slightly worse in performance than the raw layer output, on Llama8B. In fact, the raw MLP neurons are the strongest baseline, indicating that the LLM already exhibits some native disentanglement, without the help of an external encoder.

Most interestingly, the Llama1B GLP outperforms all of the Llama8B raw acti-vations, suggesting that GLP is an encouraging alternative to LLM scaling for achieving parsimonious and human-interpretable representations. 5.2. Scaling Behavior of 1-D Probes We then investigate whether scaling improves 1-D probing performance, for Llama1B GLP s trained on varying model sizes and data scales. We anchor at the last checkpoint and filter to a single candidate per layer, then select the best via val AUC. In Figure 2c we visualize the results for inputs at t = 0.5, which displays the cleanest scaling trend; see a comparison of timesteps at Figure 15.

Most notably, none of the curves exhibit a plateau, meaning that allocating more compute could lead to even higher probe scores. 5.3. Exploring Meta-Neurons To better understand the meta-neurons discovered by 1-D probing, we extract maximally activating examples over a large corpus, following standard practice in automated neuron description (Bills et, 2023; Choi et, 2024). We take documents from the FineWeb training set, truncate them to max 64 tokens, resulting in 1M total tokens from 16k unique docs. Since we have already localized concepts to their best meta-neuron location in the process of probing, we can examine their consistency with their top-3 activating examples, as shown in Table 5.

We observe that the dis-covered meta-neurons exhibit consistent activation patterns,, baseball terms for a baseball meta-neuron or expres-sions of disagreement for a contradiction meta-neuron. ## 6. Related Work Meta-Models. Meta-models treat neural networks as a new data modality (Schürholt et; Horwitz et, 2025). Prior work often focuses on network weights, spanning domains 7Learning a Generative Meta-Model of LLM Activations > Table 5. Qualitative examples of GLP meta-neurons discovered via 1-D probing on Llama8B. We show the top-3 maximally activating documents from FineWeb, with top tokens bolded. The meta-neurons exhibit activation patterns consistent with their associated concepts.

Task Info Top-3 Activating FineWeb Examples Task: 156_athlete_sport_baseball 1-D Probe AUC: 0.99 Location: Layer 0, Neuron 769 1. Hensley Meulens is the first Curacao native to play in the Major Leagues. 2. When the winning run crossed home plate in the ninth inning Friday... 3. Commissioner Bud Selig wants baseball, not the government, to determine the game’s steroid policy... Sel ig said.. Task: 138_glue_mnli_contradiction 1-D Probe AUC: 0.74 Location: Layer 4, Neuron 1654 1. Henry Kissinger is arguing that the Vietnam War taught us the perils of military withdrawal. But the true lesson of the Vietnam War...

2. The city of Surat has long been known as the diamond polishing hub of the world, but there are other facets that have led the city to shine... 3. Yellow is one of my all-time favorite colors. But when it’s in the form of pollen on our driveway? Not so much. like image classifier weights (Peebles et, 2022; Wang et, 2024; Zeng et, 2025), NeRFs (Erkoç et, 2023), Stable Diffusion LoRAs (Dravid et), and LLM LoRAs (Il-harco et, 2023; Charakorn et, 2025). However, mod-eling weights is inherently challenging: data generation re-quires expensive optimization, and training requires special techniques to overcome permutation symmetry.

We sidestep both issues by modeling activations instead of weights. Most relevant to our work, recent methods investigate dif-fusion models on DINO (Caron et, 2021) activations, demonstrating that they can be used for image generation as a conditioning signal (Li et, 2024) or latent space (Zheng et, 2025). In this work, rather than using the generated samples, we leverage the meta-model itself, using it as a prior for steering and an encoder for probing. Activation Modeling. Many LLM interpretability ap-proaches impose linear assumptions, treating concepts as directions in activation space. These include dictionary learning methods like SAEs (Olshausen & Field, 1997; Lee et

, 2006; Bricken et, 2023; Huben et, 2024; Gao et, 2025) and vector arithmetic methods (Mikolov et, 2013) like DiffMean (Marks & Tegmark, 2024), Task and Function Vectors (Hendel et, 2023; Todd et, 2024), RepE (Zou et, 2025), and Persona Vectors (Chen et, 2025). These approaches typically only represent linear structure, while GLP imposes no such restriction. A separate line of work develops nonlinear methods for describing activations in natural language; this includes SelfIE (Chen et, 2024), LatentQA (Pan et, 2024) and others (Karvonen et

, 2026; Choi et, 2025; Li et, 2025; Huang et, 2025). These methods aim to verbalize activations rather than model their distribution, and thus serve a complementary role to GLP. Diffusion Language Models. The diffusion objective has been proposed for pure language modeling, including dis-crete diffusion over tokens (Lou et, 2024) and continuous diffusion over word embeddings (Li et, 2022) and soft prompts (Lovelace et, 2024). However, diffusion LLMs are trained from scratch to compete with, rather than under-stand, autoregressive ones. Consequently, these models can only generate language and cannot manipulate activations.

## 7. Discussion We have shown that diffusion models can learn the distribu-tion of LLM activations, and that the resulting meta-model is useful downstream: as a prior that keeps steering interven-tions on-manifold, and as a feature extractor whose meta-neurons isolate interpretable concepts. Both applications improve with scale, tracking the diffusion loss. These use cases and their scaling behavior suggest that generative meta-models are a promising primitive for interpretability— one that sidesteps restrictive structural assumptions. Limitations. Our approach has several limitations that sug-gest directions for future work. First, we model single-token activations independently; multi-token modeling might cap-ture cross-position structure and enable new applications.

Second, GLP is unconditional, and conditioning on the clean activation (rather than a noised version) could reduce infor-mation loss for applications like steering. Third, we focus on residual stream activations at a single layer; extending to other activation types or further exploring the multi-layer model may yield richer representations. Future Directions. Analogies from image diffusion also suggest further applications. For instance, diffusion loss has been used as a measure of image typicality (Li et, 2023a; Siglidis et, 2024); high loss under GLP might sim-ilarly flag unusual or out-of-distribution activations. More broadly, we hope GLP provides a foundation for importing techniques from the rich literature on diffusion models into the domain of neural network interpretability.

8Learning a Generative Meta-Model of LLM Activations Acknowledgements. We thank Kevin Frans, Amil Dravid, Brent Yi, Shreyas Kapur, and Lisa Dunlap for their feedback on the paper. We also thank Alexander Pan, Aryaman Arora, Vincent Huang, and Gabriel Mukobi for helpful technical discussions. Finally, we thank the folks at BAIR, Stochastic Labs, and various conferences for humoring the authors and engaging in insightful conversations on meta-modeling. Impact Statement. This paper studies generative mod-els of activations. We find that the approach is useful for traditional interpretability tasks like steering and probing, es-pecially when trained with increasing amounts of compute.

We caution future researchers to remain cognizant of the environmental impact associated with large-scale training. Overall, we believe that our method poses minimal safety risks, as it can only directly generate activations, unlike generative models of images or text which can be misused for harmful content generation. ## References Alain, G. and Bengio, Y. Understanding intermediate layers using linear classifier probes, 2017. URL https:// Albergo, M. S. and Vanden-Eijnden, E. Building normaliz-ing flows with stochastic interpolants. In The Eleventh International Conference on Learning Representations,2023. URL net/forum? Bau,, Zhu,, Strobelt,, Lapedriza,, Zhou,

, and Torralba, A. Understanding the role of individual units in a deep neural network. Proceedings of the National Academy of Sciences,2020. ISSN 0027-8424. doi: 10.1073/pnas. 1907375117. URL org/ Belinkov, Y. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48(1):207–219, March 2022. doi: 10.1162/coli_a_00422. URL https: Belrose, N. Diff-in-means concept editing is worst-case optimal. ai/ diff-in-means, 2023. Bills,, Cammarata,, Mossing,, Tillman,, Gao,, Goh,, Sutskever,, Leike,, Wu,, and Saunders, W. Language models can explain neurons in language models. https: net/ html,2023. Bradley, A. P. The use of the area under the roc curve in the evaluation of machine learning algorithms.

Pattern Recog-nition, 30:1145–1159, 1997. URL https://api. Bricken,, Templeton,, Batson,, Chen,, Jermyn,, Conerly,, Turner,, Anil,, Denison,, Askell,, Lasenby,, Wu,, Kravec,, Schiefer,, Maxwell,, Joseph,, Hatfield-Dodds,, Tamkin,, Nguyen,, McLean,, Burke, J., Hume,, Carter,, Henighan,, and Olah, C. Towards monosemanticity: Decomposing language models with dictionary learning. Transformer Circuits Thread, 2023. html. Caron,, Touvron,, Misra,, Jégou,, Mairal,, Bojanowski,, and Joulin, A. Emerging properties in self-supervised vision transformers. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pp.

9650–9660, October 2021. Chanin, D. sae-llama-3.2-1b-topk-res. https: co/chanind/sae-llama-3. Chanin, D. and Garriga-Alonso, A. Sparse but wrong: Incorrect l0 leads to incorrect features in sparse au-toencoders, 2025. URL org/abs/ Charakorn,, Cetin,, Tang,, and Lange, R. T. Text-to-loRA: Instant transformer adaption. In Forty-second International Conference on Machine Learning,2025. URL net/forum? Chen,, Vondrick,, and Mao, C. Selfie: self-interpretation of large language model embeddings. In Proceedings of the 41st International Conference on Ma-chine Learning, ICML’24. org, 2024. Chen,, Arditi,, Sleight,, Evans,, and Lindsey, J. Persona vectors: Monitoring and controlling char-acter traits in language models, 2025.

URL https: Choi,, Huang,, Meng,, Johnson, D., Stein-hardt,, and Schwettmann, S. Scaling automatic neuron description. org/ neuron-descriptions, October 2024. Choi,, Huang,, Schwettmann,, and Stein-hardt, J. Scalably extracting latent represen-tations of users. org/ user-modeling, November 2025. Dowson, D. C. and Landau, B. V. The fréchet distance between multivariate normal distributions. Journal of Multivariate Analysis, 12(3):450–455, 1982. 9Learning a Generative Meta-Model of LLM Activations Dravid,, Gandelsman,, Wang,, Abdal,, Wet-zstein,, Efros, A., and Aberman, K. Interpreting the weight space of customized diffusion models. In The Thirty-eighth Annual Conference on Neural Information Processing

Erkoç,, Ma,, Shan,, Nießner,, and Dai, A. Hyper-diffusion: Generating implicit neural fields with weight-space diffusion. In Proceedings of the IEEE/CVF In-ternational Conference on Computer Vision (ICCV), pp. 14300–14310, October 2023. Esser,, Kulal,, Blattmann,, Entezari,, Müller,, Saini,, Levi,, Lorenz,, Sauer,, Boesel,, Podell,, Dockhorn,, English,, and Rom-bach, R. Scaling rectified flow transformers for high-resolution image synthesis. In Forty-first International Conference on Machine Learning, 2024. URL https: Fiotto-Kaufman, J., Loftus, A., Todd,, Brinkmann,, Pal,, Troitskii,, Ripa,

, Belfki,, Rager,, Juang,, Mueller,, Marks,, Sharma, A., Lucchetti,, Prakash,, Brodley, C., Guha,, Bell,, Wallace, B., and Bau, D. NNsight and NDIF: Democratizing access to open-weight foundation model internals. In The Thirteenth International Conference on Learning Representations, 2025. URL https:// Gao,, la Tour, T., Tillman,, Goh,, Troll,, Radford,, Sutskever,, Leike,, and Wu, J. Scaling and evaluating sparse autoencoders. In The Thirteenth International Conference on Learning Representations,2025. URL net/forum? Gao,, Hoogeboom,, Heek,, Bortoli, V.

, Murphy, K., and Salimans, T. Diffusion meets flow matching: Two sides of the same coin. 2024. URL https:// Gokaslan,, Cohen,, Pavlick,, and Tellex, S. Open-webtext corpus. github. io/OpenWebTextCorpus, 2019. Grattafiori,, Dubey,, Jauhri,, Pandey,, Kadian,, Al-Dahle,, Letman,, Mathur,, Schelten,, Vaughan,, Yang,, Fan,, Goyal,, Hartshorn,, Yang,, Mitra,, Sravankumar,, Korenev,, Hinsvark,, Rao,, Zhang,, Rodriguez,, Gregerson,, Spataru,, Roziere,, Biron,, Tang,, Chern,, Caucheteux,, Nayak,, Bi,

, Marra,, McConnell,, Keller,, Touret,, Wu,, Wong,, Ferrer, C., Nikolaidis,, Allonsius,, Song,, Pintz,, Livshits,, Wyatt,, Esiobu,, Choudhary,, Mahajan,, Garcia-Olano,, Perino,, Hupkes,, Lakomkin,, AlBadawy,, Lobanova,, Dinan,, Smith, E., Radenovic,, Guzmán,, Zhang,, Synnaeve,, Lee,, Anderson, G., Thattai,, Nail,, Mialon,, Pang,, Cucurell,, Nguyen,, Ko-revaar,, Xu,, Touvron,, Zarov,, Ibarra, I., Kloumann,, Misra,, Evtimov,, Zhang,, Copet,

, Lee,, Geffert,, Vranes,, Park,, Mahadeokar,, Shah,, van der Linde,, Billock,, Hong,, Lee,, Fu,, Chi,, Huang,, Liu,, Wang,, Yu,, Bitton,, Spisak,, Park,, Rocca,, Johnstun,, Saxe,, Jia,, Alwala, K., Prasad,, Upasani,, Plawiak,, Li,, Heafield,, Stone,, El-Arini,, Iyer,, Malik,, Chiu,, Bhalla,, Lakhotia,, Rantala-Yeary,, van der Maaten,, Chen,, Tan,, Jenkins,, Martin,, Madaan,, Malo,, Blecher,

, Landzaat,, de Oliveira,, Muzzi,, Pasupuleti,, Singh,, Paluri,, Kardas,, Tsimpoukelli,, Oldham,, Rita,, Pavlova,, Kambadur,, Lewis,, Si,, Singh, M., Hassan,, Goyal,, Torabi,, Bashlykov,, Bogoychev,, Chatterji,, Zhang,, Duchenne,, Çelebi,, Alrassy,, Zhang,, Li,, Vasic,, Weng,, Bhargava,, Dubal,, Krishnan,, Koura, P., Xu,, He,, Dong,, Srinivasan,, Ganapathy,, Calderer,, Cabral, R., Stojnic,, Raileanu,, Maheswari,, Girdhar,, Patel,, Sauvestre,

, Polidoro,, Sumbaly,, Taylor,, Silva,, Hou,, Wang,, Hosseini,, Chennabasappa,, Singh,, Bell,, Kim, S., Edunov,, Nie,, Narang,, Raparthy,, Shen,, Wan,, Bhosale,, Zhang,, Vandenhende,, Batra,, Whitman,, Sootla,, Collot,, Gururangan,, Borodinsky,, Herman,, Fowler,, Sheasha,, Georgiou,, Scialom,, Speck-bacher,, Mihaylov,, Xiao,, Karn,, Goswami,, Gupta,, Ramanathan,, Kerkez,, Gonguet,, Do,, Vogeti,, Albiero,, Petrovic,, Chu,, Xiong,, Fu,

, Meers,, Martinet,, Wang,, Wang,, Tan, X., Xia,, Xie,, Jia,, Wang,, Gold-schlag,, Gaur,, Babaei,, Wen,, Song,, Zhang,, Li,, Mao,, Coudert, Z., Yan,, Chen,, Papakipos,, Singh,, Srivastava,, Jain,, Kelsey,, Shajnfeld,, Gangidi,, Victoria,, Goldstand,, Menon,, Sharma,, Boesenberg,, Baevski,, Feinstein,, Kallet,, Sangani,, Teo,, Yunus,, Lupu,, Alvarado,, Caples,, Gu,, Ho,, Poul-ton,, Ryan,, Ramchandani,, Dong,

, Franco,, Goyal,, Saraf,, Chowdhury,, Gabriel,, Bharambe,, Eisenman,, Yazdan,, James,, Maurer,, Leonhardi,, Huang,, Loyd,, Paola, B., Paranjape,, Liu,, Wu,, Ni,, Hancock,, Wasti,, Spence,, Stojkovic,, Gamido,, Montalvo,, Parker,, Burton,, Mejia,, Liu,, Wang,, Kim,, Zhou,, Hu,, Chu,, Cai,, Tindal,, Feichtenhofer,, Gao,, Civin,, Beaty,, Kreymer,, Li,, Adkins,, Xu,, Testuggine, 10 Learning a Generative Meta-Model of LLM Activations

, David,, Parikh,, Liskovich,, Foss,, Wang,, Le,, Holland,, Dowling,, Jamil,, Mont-gomery,, Presani,, Hahn,, Wood,, Le,, Brinkman,, Arcaute,, Dunbar,, Smothers,, Sun,, Kreuk,, Tian,, Kokkinos,, Ozgenel,, Cag-gioni,, Kanayet,, Seide,, Florez, G., Schwarz,, Badeer,, Swee,, Halpern,, Herman,, Sizov,, Guangyi, Zhang, Lakshminarayanan,, Inan,, Shojanazeri,, Zou,, Wang,, Zha,, Habeeb,, Rudolph,, Suk,, Aspegren,, Goldman,, Zhan,, Damlaj,

, Molybog,, Tufanov,, Leontiadis,, Veliche,, Gat,, Weissman,, Geboski,, Kohli,, Lam,, Asher,, Gaya,, Marcus,, Tang,, Chan,, Zhen,, Reizenstein,, Teboul,, Zhong,, Jin,, Yang,, Cummings,, Carvill,, Shepard,, McPhie,, Torres,, Ginsburg,, Wang,, Wu,, U, K., Saxena,, Khandelwal,, Zand,, Matosich,, Veeraraghavan,, Michelena,, Li,, Jagadeesh,, Huang,, Chawla,, Huang,, Chen,, Garg,, A,, Silva,, Bell,, Zhang,, Guo,

, Yu,, Moshkovich,, Wehrstedt,, Khabsa,, Avalani,, Bhatt,, Mankus,, Hasson,, Lennie,, Reso,, Groshev,, Naumov,, Lathi,, Keneally,, Liu,, Seltzer, M., Valko,, Restrepo,, Patel,, Vyatskov,, Samvelyan,, Clark,, Macey,, Wang,, Hermoso, M., Metanat,, Rastegari,, Bansal,, Santhanam,, Parks,, White,, Bawa,, Singhal,, Egebo,, Usunier,, Mehta,, Laptev, N., Dong,, Cheng,, Chernoguz,, Hart,, Salpekar,, Kalinli,, Kent,, Parekh,, Saab,

, Balaji,, Rittner,, Bontrager,, Roux,, Dollar,, Zvyagina,, Ratanchandani,, Yuvraj,, Liang,, Alao,, Rodriguez,, Ayub,, Murthy,, Nayani,, Mitra,, Parthasarathy,, Li,, Hogan,, Battey,, Wang,, Howes,, Rinott,, Mehta,, Siby,, Bondu, S., Datta,, Chugh,, Hunt,, Dhillon,, Sidorov,, Pan,, Mahajan,, Verma,, Yamamoto,, Ramaswamy,, Lindsay,, Lindsay,, Feng,, Lin,, Zha, S., Patil,, Shankar,, Zhang,, Zhang,, Wang,, Agarwal,, Sajuyigbe,

, Chintala,, Max,, Chen,, Kehoe,, Satter-field,, Govindaprasad,, Gupta,, Deng,, Cho,, Virk,, Subramanian,, Choudhury,, Goldman,, Remez,, Glaser,, Best,, Koehler,, Robinson,, Li,, Zhang,, Matthews,, Chou,, Shaked,, Vontimitta,, Ajayi,, Montanez,, Mohan,, Kumar, V., Mangla,, Ionescu,, Poenaru,, Mi-hailescu, V., Ivanov,, Li,, Wang,, Jiang,, Bouaziz,, Constable,, Tang,, Wu,, Wang,, Wu,, Gao,, Kleinman,, Chen,, Hu,, Jia,

, Qi,, Li,, Zhang,, Zhang,, Adi,, Nam,, Yu, Wang, Zhao,, Hao,, Qian,, Li,, He,, Rait,, DeVito,, Rosnbrick,, Wen,, Yang,,

Want to learn more?

Ask about this article