DOCUMENTATION · FIELD NOTES

How to use the LLM Probe

The LLM Probe explores pre-collected perturbations around prompt-conditioned activations in Gemma 2. It is an empirical visualization: displayed nodes correspond to model measurements, while connecting faces make the sampling topology legible.

The complete LLM Probe instrument at its prompt-conditioned anchor
FIGURE 01. The complete instrument in Surface mode at the anchor. The prompt and coordinates appear above the measured terrain; response status and controls surround the Layer-14 SAE telemetry and the separate final-output token readout.

Experimental path

LAYER 10PERTURBATION

A prompt determines the anchor activation. The probe adds a controlled displacement in a selected plane.

LAYER 14SAE READOUT

The Layer-14 activation is encoded by its SAE to obtain active feature identities and magnitudes.

FINAL OUTPUTBEHAVIORAL READOUT

The normal language-model output supplies token probabilities, entropy and KL divergence.

These are separate observers. Layer 14 supplies the SAE feature response; token statistics come from Gemma’s ordinary final logits after all remaining layers. No Layer-14 logit lens is used.

Instrument anatomy

The display has three working zones. The prompt and probe coordinates sit above the terrain. The terrain itself is the spatial workspace. The lower console selects the dataset and reports feature and token responses.

Prompt deck

Selects one pre-collected prompt and displays its text, distance, direction and current metric.

Terrain

Shows the selected two-dimensional activation-space slice in Explore or Surface mode.

Response status

Reports the strongest observed categorical displacement at the current probe location.

Analysis console

Selects field, plane, metric and plot mode while reporting SAE and token responses.

Your first exploration

  1. Choose a prompt.The anchor, baseline tokens and all displayed measurements change together.
  2. Start in the near field.Select the SAE plane, KL divergence and Explore.
  3. Click once to enable the probe.Move outward from the anchor and watch the response-status instrument.
  4. Switch to Surface.Rotate the completed measured surface and inspect its basin.
Prompt selector with the English Dickens prompt selected
FIGURE 03A. The prompt menu identifies each pre-collected example by corpus ID, language and source. Selecting one loads its anchor, measurement planes, token baseline and SAE data as one coordinated dataset.
The probe leaving its anchor
FIGURE 03B. In Explore mode, the luminous probe pulls a distance-and-direction line away from the anchor. Previously visited samples remain visible as a trail, while vertical markers encode the selected metric.

Reading the terrain

The anchor is the activation produced by the selected prompt. The two colored axes span the displayed plane.

The readout gives three coordinates at the probe: d is distance from the anchor in multiples of the anchor RMS; θ is direction in the displayed plane; the third line is the value of the currently selected metric.

Probe coordinate readout showing distance, angle and KL divergence
DETAIL A. Distance, direction and the selected metric at one measured location.
Surface convention. Nodes are measurements. Edges show declared adjacency. Each triangular face is colored by the average metric value at its three measured nodes. No regression or spline-smoothed boundary is presented as an observation.
Plot-mode selector
FIGURE 04A. Explore reveals measured locations through direct probing; Surface displays the completed connected-node surface for the selected dataset. Reset clears exploration state and restores the standard view.
Explore

The terrain begins substantially unmarked. Moving the probe reveals individual measured locations and leaves a persistent trail. This mode supports guided discovery during a presentation.

Surface

All sampled nodes and declared connections are shown at once. Rotate and zoom this completed measurement mesh to inspect basin shape, asymmetry and large-scale structure.

Response status

The five-part instrument at the left reports categorical changes at the current probe location. Only one state is illuminated at a time.

At anchor
The probe is at the prompt-conditioned Layer-10 activation.
No response
The measured Top-50 SAE feature identities and Top-5 token identities remain unchanged.
Top-50 SAE feature displacement
The Top-50 SAE feature set has reorganized while the Top-5 tokens remain in place.
Top-5 token displacement
At least one Top-5 token or its ordering has changed, but Top-1 is retained.
Top-1 token displacement
The most likely token differs from the anchor readout.
The response-status instrument with SAE feature displacement selected
FIGURE 05A. The illuminated lozenge reports the strongest categorical response observed at the probe. Here the Top-50 SAE feature set has changed while the Top-5 token identities remain intact.

SAE feature displacement monitor

The percentage reports how many SAE features active at the anchor remain active at the probe. The two counts show total active features now and at the anchor; they need not move in the same proportion. The address field renders the 256 strongest currently active features, with marker size reflecting activation magnitude.

SAE feature displacement monitor
FIGURE 05B. At this probe location, 96% of the anchor's SAE features remain active. The total active count has fallen from 753 at the anchor to 742, while the feature-address field shows the strongest current activations.

Metrics

KL divergence
Difference between the model’s final-output token distribution at the probe and at the anchor.
Entropy
Uncertainty of the complete final-output token distribution, reported in nats.
SAE feature count
Number of non-zero features in the Layer-14 SAE readout.
Metric selector
FIGURE 06A. Only one metric controls terrain height, marker color and the coordinate readout at a time: KL divergence, output entropy or active SAE feature count.

In Surface mode, marker height and Viridis color encode the selected metric together. The vertical key at the right reports the active range. Changing the metric does not move the sampled coordinates; it changes the measured quantity projected onto height and color.

KL divergence

Emphasizes how far the complete final-output token distribution has moved from the prompt-conditioned anchor.

Entropy

Shows where the final output becomes more or less uncertain, independently of which tokens changed.

Feature count

Shows expansion or contraction in the number of non-zero Layer-14 SAE features.

Final-output token readout

The Top-5 list ranks the most probable next tokens at the current probe location. The bars and percentages are obtained by applying a full-vocabulary softmax to Gemma’s ordinary final logits—not by unembedding Layer 14. A token is red when its identity or rank differs from the anchor baseline. Output entropy summarizes uncertainty across the complete final-output distribution, not merely these five entries.

Final-output Top-5 token readout and entropy
FIGURE 06B. The rank column orders the predicted next tokens. Here the first three positions agree with the anchor baseline, while the red fourth and fifth entries have changed. The distribution is extremely concentrated: the Top-1 probability is 99.7% and output entropy is 0.03 nats.

Planes

The plane selector changes the two directions spanning the displayed activation-space slice. The prompt-conditioned anchor stays fixed, but each choice exposes different surrounding geometry.

Plane selector
FIGURE 07A. Random planes are numbered so that a particular sampled orientation remains identifiable and reproducible.
SAE plane

Spans the two selected SAE feature directions. This is the feature-informed plane used for the primary expedition.

PCA plane

Spans principal directions estimated from activation data, emphasizing directions of observed variance.

Random plane

Provides a seeded comparison not chosen for SAE structure or explained variance.

Near and far fields

The near field covers 0–1,024× RMS and resolves the prompt-conditioned basin. The far field covers 1,024–16,384× RMS and surveys large-scale geometry using a separate visual radial scale. The selector prevents those two scales from being mistaken for a single linear map.

Field selector
FIGURE 08A. Near and far are deliberately separate viewing scales. Their boundary is shared at 1,024× RMS, while the far field extends the survey to 16,384× RMS.
Near-field KL-divergence surface
FIGURE 08B. The prompt-conditioned anchor lies at the center of a broad, low-divergence basin. The visible nodes and edges disclose the measurement mesh; height and Viridis color both encode KL divergence.
Far-field KL-divergence surface
FIGURE 08C. The far-field cartridge is drawn as an annulus because its sampling begins at 1,024× RMS; the near-field interior is intentionally omitted in this mode. Its coarser connected-node mesh surveys the outer domain through 16,384× RMS.

Mouse and keyboard

Move mouseMove the probe among measured samplesClickLock or unlock the probeShift + dragRotate the three-dimensional viewCtrl + Shift + dragZoom toward the cursorCtrl + Shift + wheelZoom where supported by the browserDouble-clickReveal the selected measured region in Explore modeRReset the camera, probe and exploration state
The lower instrumentation panel
FIGURE 09. The lower console combines Layer-14 SAE retention and Feature Flux with field, plane, metric and plot controls, followed by the final-output Top-5 token readout and entropy.

Data provenance

All model outputs are pre-collected. The public site does not run Gemma in the browser. Each prompt, plane and field loads a persisted measurement cartridge containing coordinates, token readouts, entropy, KL divergence and SAE feature data.

Limitations

  • The view is a two-dimensional slice through a much higher-dimensional activation space.
  • Plane choice changes the terrain being observed.
  • Token identity can overstate semantic displacement across equivalent spellings, scripts or tokenizations.
  • Rendered surfaces connect samples for visualization; only nodes are direct measurements.
  • Results are prompt-conditioned and should not be generalized to all activations without further experiments.