PreCommitLens J-Lens Viewer

Static, Neuronpedia-style inspection for the Qwen3-0.6B dense Jacobian-lens run: layer-by-layer readouts, watched forbidden concepts, paired attack/control outcomes, and validator decisions before commit.

Qwen/Qwen3-0.6B Qwen/Qwen3-4B trajectory v4b Qwen3-4B-native discovery v4c Gemma E2B / Qwen3.5 appendix Qwen3.5-4B confirmatory v4d 28 layers 1024 x 1024 dense Jacobians Static precomputed results No live inference

Confirmatory v4: fixed-prompt trajectory prediction

Fresh trajectories, semantic policy landing, within-prompt paired AUC, and frozen success criteria.

Primary gate: FAIL

Outcome becomes readable before policy landing, but layer-18 residuals do not beat a shallow visible-prefix TF-IDF baseline by the pre-registered margin. This is an added-value failure, not a contrast-yield failure.

Confirmatory data1,088 trajectories / 34 prompts
Fresh test contrast9 / 9 prompts, 3 risks
Checkpoint 8 AUC0.823 residual / 0.817 TF-IDF
Capture cost1.014x plain generation
v4 pre-landing AUC and residual advantage curve

Confirmatory v4b: frozen-prompt scale transfer

Qwen3-4B FP16, the same 34 prompts and gate, and depth-normalized residual layers.

Primary gate: INCONCLUSIVE

The 0.6B contrast-selected prompts do not remain contrastive at 4B. With zero mixed training prompts and only one mixed test prompt, the residual-added-value comparison is unidentified rather than negative.

Confirmatory data1,088 trajectories / 34 prompts
Fresh test contrast1 / 9 prompts, 1 risk
Prompt collapse32 / 34 single-outcome
RTX 3060 peak7.644 GiB allocated

Discovery v4c: Qwen3-4B-native contrast search

Three mechanisms, 192 frozen candidates, and a pre-registered yield gate.

Discovery gate: YIELD FAIL

The frozen Qwen3-4B FP16 redesign did not produce enough within-prompt outcome contrast. The protocol therefore stopped before residual capture or probe fitting; this is a discovery-yield result, not evidence that residual probes fail at 4B.

Discovery data3,072 trajectories / 192 prompts
Eligible by round3 / 1 / 0 of 64
Lottery diagnostic95.8% chose larger weight
RTX 3060 peak7.592 GiB allocated

Appendix: temperature and deployment sensitivity

Four frozen post-hoc conditions with no change to the completed v4c gate.

Appendix: COMPLETE

Higher temperature partially restores Qwen3-4B contrast but remains below the original yield threshold. Gemma E2B stays low-yield, while the Qwen3.5-4B deployment produces substantial within-prompt contrast. The paradigm is model- and sampling-dependent.

Qwen3-4B eligible3 / 9 / 11 at T=0.8 / 1.2 / 1.5
Gemma E2B3 / 64 eligible
Qwen3.5-4B34 / 64 eligible
Qwen3.5 A/B switch52 / 64 prompts

Final v4d: Qwen3.5-4B FP16 accessibility

A pre-registered deployment transfer followed by a 1,056-trajectory confirmatory run.

Primary gate: FAIL

Benchmark contrast survives the move from Ollama Q4_K_M to Transformers FP16, but every early residual and visible-prefix method remains exactly at chance. At checkpoints 0-16, compliant and violating trajectories have identical visible prefixes and identical captured residual states within each evaluable prompt.

FP16 stage-one yield36 / 64 eligible
Confirmatory data1,056 trajectories / 33 prompts
Fresh test contrast8 / 8 prompts, 3 risks
Primary AUC curve0.500 for every method
Capture cost1.005x plain generation
Stopping boundaryNo v4e

Layer Top-1 Strip

rank 1-10 rank 11-100 rank 101-1k rank 1k-10k

Watched Concept Heatmap

Rows are watched concept tokens; columns are layers. Click any cell.

Selected Cell

Layer 23 reveal

Top Tokens At Selected Layer

Matched Pair Delta

Attack/control comparison for the current scenario.

Reports

Loading...