TETRACTA · MODEL X-RAY

🩺 Diagnostic Report

Qwen/Qwen3-8B-Base → Qwen/Qwen3-8B

MEDIUM — noticeable functional deviation; targeted eval recommended
WHAT THIS SCAN GIVES YOU

New to this? Plain-words guide. station / layer: one internal processing step — text enters at the early stations and the answer takes shape toward the late ones  ·  spread (deprecated): legacy fraction of stations above threshold — a function of the start station, kept only as a footnote  ·  behavior changed: share of sampled prompts whose output text differs — differing means re-worded, not necessarily wrong  ·  perplexity: the standard overall 'how well does the model predict text' score; if it barely moves, output-level dashboards see nothing  ·  station indexing: percentages and enrichment use N+1 stations as the denominator — the model's N layers plus one embedding station (station 0)

Finding → Interpretation → Recommendation

Note on change-start: our estimator has a detection floor at station 4 with a structural cause: the first probe is injected at layer 2, so stations 0–3 are zero by construction and the earliest observable difference is station 4 — in a controlled test, damage confined to stations 0–3 also surfaced at station 4. The floor is a property of the probe grid, not of the model: measured directly it is 4 (measured on 14 base→instruct pairs across 6 families under ps-1.1/mv-1.2 on 2026-09-03: start = station 4 in 14 of 14, from 0.5B to 32B). A reported start of 4 therefore means at or before station 4; a start of 5–6 on a much deeper model can likewise be the floor rather than a finding.

Finding. The difference begins at or before station 4; 23 of 37 stations carry 80% of the difference mass (effective width, N80; ≈21–25 within the test-retest band), and 100% of sampled outputs changed. The largest functional difference concentrates around stations 32–36 (23% of total difference mass · 1.72× enrichment over a uniform spread, ±4% test-retest · flat-null 1.12× → 1.53× excess; fixed k=5 window). (Legacy "spread" = 89% — deprecated in rs-1.2: it is a deterministic function of the start station, not an independent measurement.)

Note on the behavior number: if model A is a base/completion model and model B is chat-tuned, prompt-format differences alone can push 'behavior changed' toward 100% — weight the internal fields (N80 / concentration band) more in that pairing; change-start is at the floor and carries no weight here.

Interpretation. Broad functional shift. 62% of the stations carry 80% of the difference mass. Spread over a wide part of the network without a dominant band. A second, weaker band sits at stations 10–14 — the change is two-centred, not one-centred.

Reference-library context (mv-1.2 ledger, re-scanned 2026-09-03 on the production fleet: 13 full-SFT pairs across 6 families (Qwen2.5, Qwen3, Llama-3.1/3.2, Mistral, OLMo-2, Falcon3, SmolLM2; 0.5B–32B) and 8 simulated-int8 runs; bands are min–max over one or two models per family — read as orientation, not verdict): full-SFT: 4/4 features in band (N80 0.40–0.67 · enrichment 1.33–3.13 · excess over flat null 1.17–2.93 · start at floor) · simulated-int8: 4/4 features in band (N80 0.54–0.67 · enrichment 1.28–2.38 · excess over flat null 1.09–2.23 · start at floor).

Recommendation (template matched to the inferred modification type: unknown)

Method notes. Each station value is a relative deviation (‖h′−h‖/‖h‖ at that station), so residual-norm growth with depth does not inflate late stations; the difference profile is the mean absolute difference of the two models' relative-response curves over 168 matched probes — each station averaged over the probes that can reach it (contributor-normalized, mv-1.2). Prompt protocol: identical raw-text prompts to both models (no chat template), greedy 32-token continuation for the behavior figure, fixed seed. Test-retest bands (±5% N80, ±4% enrichment; 9 probe-seeds, 0.5B, mv-1.2, 2026-09-03) are extrapolated to this scale. Layer-localization validation (21/21) used controlled synthetic damage; generalization to real fine-tunes is a separate claim and is not asserted here.

The model, and where this scan landed

Model anatomy with this scan overlaid

Left: the scanned model's own skeleton — its layers and the S = N+1 stations we read. Right: this scan's difference at each station, aligned row by row. Ringed stations carry 80% of the difference mass; the shaded band is the densest five.

Where the change is (layer strip)

▲ change starts here

Darker = larger functional difference. Left = early stations (input parsing), right = late stations (answer shaping).

Change starts at
≤ station 4
at or before — detection floor (stations ≤3 are zero by construction)
Effective width (N80)
23/37
stations carrying 80% of difference mass · ≈21–25 at ±5% test-retest · probe-grid null N80 27/37
Concentration
stations 32–36
23% of total difference mass · 1.72× enrichment over a uniform spread, ±4% test-retest · flat-null 1.12× → 1.53× excess; fixed k=5 window
Behavior changed
6 of 6 prompts
greedy 32-token continuation, exact match · 95% CI 54–100% · 1/6 resolution — indicative only
Difference magnitude
max 2.9e-03 · mean 9.1e-04
relative deviation units · tiers: CLEAN <1e-5 · LOW <1e-2 · MEDIUM <5e-2 · HIGH

Scan provenance

Pin these fields when you diff this report against a future scan — they separate "the model changed" from "the repo or our probe changed". The weights fingerprint is a digest of the per-shard SHA-256 manifest from the signed attestation.

Scan date (UTC)2026-09-03T05:19:15Z
Model referenceQwen/Qwen3-8B-Base -> Qwen/Qwen3-8B
Weights fingerprintA:sha256:2b736f31d2b0 (5 shards) · B:sha256:da90f9e572c3 (5 shards)
HF revisionA:49e3418fbbbc · B:b968826d9c46
Probe set / scan configps-1.1 / sc-1.0 · mv-1.2
Protocol6 raw-text prompts (no chat template) × 2 positions × 2 directions × 7 strike layers = 168 matched probes; eps = 0.10·‖h‖; seed 42; bf16; greedy 32-token continuation for the behavior figure
Reference setlabeled-intervention ledger mv-1.2 (N=21: 13 full-SFT pairs across 6 families 0.5B–32B, 8 simulated-int8; scanned 2026-09-03) — no public gallery for pair scans
Report schemars-1.5
Workerrunpod-gpu
Worker isolationsingle-tenant ephemeral cloud GPU (pod m3d8elh7…, region US-TX-1) — destroyed after the job; deletion attested
Runtimebf16 · torch 2.4.1+cu124 · transformers 5.5.4 · probe-set sha256:0a2909cf3ae6 · greedy decoding, seed 42, fixed probe order (deterministic; bitwise reproducibility across GPU types not guaranteed)
Attestationjob 219f3e80f173… · scan attested; verify at https://www.tetracta.ai/llm_tomografi/attest/219f3e80f1734086a7ac3b8098acf7e8
Revisionre-rendered 2026-09-03 under rs-1.5 with gallery g2 (14 models) and the mv-1.2 ledger (21 labelled interventions, 6 families); prompt-format detection probe-know-1.3; scan data unchanged

APPENDIX — companion metric (tokenizer efficiency); not part of the X-ray measurement above.

Language Economy (tokenizer efficiency — a companion metric, not an internal X-ray finding)

How many tokens Qwen/Qwen3-8B-Base spends on the same meaning across languages, relative to English (1.00×). Tokens are what you pay for — so this is cost and effective context, not quality. Most expensive here: Hindi 4.75×.

English
1.00×
baseline
Hindi
4.75×
+375% cost
Russian
2.08×
+108% cost
Arabic
1.90×
+90% cost
French
1.83×
+83% cost
Turkish
1.79×
+79% cost
Spanish
1.71×
+71% cost
Portuguese
1.71×
+71% cost
German
1.64×
+64% cost
Korean
1.64×
+64% cost
Japanese
1.57×
+57% cost
Chinese
1.17×
+17% cost

Method: a compact 12-language sample of FLORES-200 parallel sentences (6 sentences per language; indicative, no interval), tokens(lang)/tokens(English), this model's public tokenizer — an indicative reading (identical across sizes of the same family). For the full-precision, 204-language table (complete FLORES set), see tetracta.ai/note-token-economy.html


Honesty box. This report is a functional diagnostic: it maps where a model's internal signal flow behaves, not a task-accuracy score. It complements evals; it does not replace them. Our test of where knowledge physically sits was inconclusive — so we say resolves, not stores, on purpose. The risk label is a v0 heuristic, not a calibrated probability. Measured test-retest bands (9 probe-seeds, 0.5B base vs Instruct pair and Instruct vs simulated int8, ps-1.1/mv-1.2, 2026-09-03; extrapolated to other scales): effective-width ±5%, band-enrichment and excess ±4% — differences inside these bands should not be over-read. Layer-localization was 21/21 in controlled synthetic-damage tests (1.5B/7B/72B); that validates the instrument on injected damage — generalization to real fine-tunes is a separate claim we do not assert from it. We apply the same rules to our own research claims. Method details are proprietary.

© Tetracta · your model weights were deleted from the worker before report delivery (attested).