Two pinned artifacts compared under a fixed measurement contract
report schema rs-1.7validation pending
A · BEFORE
HuggingFaceTB/SmolLM2-1.7B
@ effd688a1292
→
B · AFTER
HuggingFaceTB/SmolLM2-1.7B-Instruct
@ 31b70e2e869a
≠
VALIDATION-PENDING DESCRIPTIVE RESULT — no calibrated risk or pass/fail grade
A response difference was observed under this exact contract.
The change is already present at the first evaluable position (output of decoder block 4 of 24); it may begin earlier, below the instrument floor, so no starting depth is reported for this pair.
Onset depth and extent are categorical and calibrated on controlled cells; this is not a magnitude or quality assessment.
difference_observed
yes
whether the instrument observed a response difference between A and B under the measurement contract
output_text_change_observed
withheld
Withheld: the two artifacts do not share an output contract, so their recorded output text is not compared.
probes evaluated
24
size of the evaluated set at the reporting depth; this is not a count of probes that differed
pair relation
combined-artifact-change
recorded relation between the two artifacts under the pair contract; a combined-artifact change is not attributed to fine-tuning alone
Depth map — where the difference is observable
difference observedno difference observedbelow the instrument floor
Each cell is one profile position: emb is the embedding output, k is the output of decoder block k (blocks are numbered from 1), final is the output after the final normalization. Under this contract the change is already present at the first evaluable position (output of decoder block 4 of 24); it may begin earlier, below the instrument floor, so no starting depth is reported for this pair; every evaluable position differs (21 of 21): the change is not confined to one part of the network. Positions below the instrument floor (the embedding output and the outputs of blocks 1 to 3) cannot carry an onset. Cells are categorical: the map carries no magnitude, and once a position differs, later positions are expected to differ as well. Onset calibration for this instrument release (controlled in-memory lesions at known depths): exact onset predicted and observed in 6 of 6 controlled cells on each of three public models (Qwen2.5-0.5B-Instruct, SmolLM2-1.7B, TinyLlama-1.1B-Chat-v1.0); cross-board agreement on onset 4 of 4 cells.
What this means for you
MeasuredEvery evaluable position differs (21 of 21); under this contract the change is not confined to one part of the network. This is a record of what the instrument observed on a fixed set of inputs, not an assessment of the models.
Text outputNot compared. The two artifacts do not share an output contract, so a direct comparison of their text would mislead rather than inform.
Suggested next checksRun your own task set on A and B with identical inputs and generation settings, and compare the outputs pair by pair: this report says that a difference exists and where in the network it is observable under this measurement contract; only your own task set can say whether it matters for your use. The difference is present from the first evaluable position onward, so treat B as a different model rather than a variant of A: re-run your full evaluation set, and review the recorded configuration, tokenizer and generation differences above before attributing the change to weights. (guidance, not a measurement)
What differs between A and B (public metadata)
Configuration comparison from the artifacts' own public files — non forward config differs.
Numeric-precision fields are not compared here; any file difference is recorded in the artifact identity.
generation settings or special tokens
tokenizer semantics
generation settings semantics
This comparison is recorded as a combined artifact change (recorded differences: generation settings or special tokens, tokenizer semantics, generation settings semantics); the difference is not attributed to weights alone.
Artifact facts and run conditions
architecture (from the public model configuration): 24 decoder blocks · hidden size 2048
scan wall time, both artifacts: 40 s
reproduced on another card class (RTX 4060, different runtime): same difference_observed result
Provenance
instrument releasexray_vg1
measurement contractmv-1.4 / paired-artifacts-2.0
report schemars-1.7
scan date (UTC)2026-09-07
card classRTX 5070 Ti
runtimePyTorch 2.11.0+cu128 · Transformers 5.10.1
operational receiptno attached to this standalone copy
Record identity: d9710f4fa16b — the first 12 characters of the ordered-pair identifier, derived from the two artifacts and their revisions; a different pair, or the same pair at a different revision, yields a different identifier. This document is a record of a measurement, not an opinion about the models: its fields are determined by the two public artifacts and the contract versions named above. In the null control for this instrument release, repeat scans on the same board class reproduced byte-identical recorded metrics.
Interpretation
This result records whether the frozen instrument observed a difference and, categorically, where it was observed. A difference is
reported when any evaluated probe's response differs between A and B beyond the instrument's
observation floor; the floor's value and the probe design are private. The depth map records the first position at which the difference was observed and the evaluable positions at which it was observed; both are categorical. It does not identify a cause,
establish that weights were the only changed artifact, measure capability or quality, or authorize
a shipping decision. Compare A to B only; do not compare numbers between reports.
What is shown and what stays proprietary
What this report shows
the verdict, and the recorded differences in configuration, tokenizer and generation settings
onset depth, the depth map and the extent of the observable difference (when enabled for the instrument release)
the output-text status, provenance, record identity and the operational receipt
What remains proprietary
probe design and probe text
the observation floor and thresholds
raw statistics, raw profiles, transformations and internal traces
sampling geometry and item-level records
The finding is yours; the instrument stays proprietary and is held fixed and versioned so that two reports produced under the same contract can be compared with each other. This standalone derivative carries no attestation; signing status is unavailable unless a separate signed receipt accompanies it.
Null-control reference for this instrument release: 0 of 16 A=B comparisons reported a difference (6 September 2026) Version notice: produced under instrument release xray_vg1 with the contract versions shown above. Records produced under different contract versions are not interchangeable with this one, and earlier records may have been restated or withdrawn — see the correction record linked below. Scope & Limitations: https://www.tetracta.ai/model-xray/scope/ · Correction record: https://www.tetracta.ai/model-xray/correction/
What this report does not do
It says at which decoder-block output a change first became observable under this contract (onset depth, calibrated on controlled cells) and at which evaluable positions it was observed; it does not say why the change happened or how large it is.
It does not say whether the change is an improvement or a problem.
It does not replace testing on your own task: treat this as a reason to test, not as a verdict.
Where the output-text field reads “withheld”, it does not mean the text output is unchanged: the two artifacts’ output contracts differ, so a direct text comparison would mislead.
Honesty box — scope and limits This is a validation-pending descriptive measurement record, not a task-accuracy,
safety, causal, deployment-readiness or regulatory result. Historical development findings used a
superseded estimator and do not validate this release. The proprietary measurement recipe and raw
evidence are intentionally withheld. Confirm decisions with task-specific evaluations on the exact
deployed artifacts.
Tetracta · Model X-Ray · report schema rs-1.7 · this page contains no scripts and no numerical profile.