TETRACTA · MODEL X-RAY

Before / after comparison

Two artifacts compared under one fixed measurement contract
Depth view · Decision, output status, available verification evidence and categorical depth observations
VG1 · betars-1.7 · rp-1.3
A · SOURCE
HuggingFaceTB/SmolLM2-360M
@ f8027fd0eaeea54caa13c31d31b9fdc459c38b49
Reference artifact
→
B · AFTER
HuggingFaceTB/SmolLM2-360M-Instruct
@ a10cc1512eabd3dde888204e902eca88bddb4951
Candidate artifact
≠
VALIDATION-PENDING DESCRIPTIVE RESULT · no calibrated risk or pass/fail grade
Difference observed
The response difference met the versioned instrument decision rule under this exact comparison contract.
Internal response
Observed
Response difference on the evaluated set
Text output
Withheld
Withheld does not mean unchanged. The two output contracts are not directly comparable.
Repeat check
Single record
No repeat-check summary is included in this copy.

Included summaries Content and scope, not a quality or risk grade

Record or summaryThis report
Measurement summaryIncluded
Repeat-check summaryNot included
Same-artifact control summaryNot included
Underlying recordsRetained in the VG1 archive; not included in this copy
Operational receiptLinked · limited operational record
Statistical noise envelopeNot certified

Repeat and control details

Repeat checkOne measurement is summarized; no repeat-check summary is included in this copy.
Same-artifact controlNo matching same-artifact control summary is included in this copy.
Operational receiptThe link identifies a limited Tetracta operational record. Verify its signature and stated scope at the linked endpoint; the presence of a link is not signature verification.
Noise-floor statusNo model-specific statistical noise envelope is certified by this report.

What this means for you

MeasuredThe internal-response comparison met the instrument decision rule on the fixed input set. Use this as evidence of a recorded change when reviewing the candidate; the report does not rate that change as beneficial or harmful.
Text outputThe output-text comparison was withheld because A and B do not share a directly comparable tokenizer/generation contract.
PrioritySeparate the output-contract change. Align tokenizer and generation settings for a controlled output comparison. The current record supports a combined artifact change, not a weights-only conclusion.

How the finding was decided

The reported internal finding uses the fixed, thresholded decision rule recorded for mv-1.4; it is not a test for any non-zero floating-point value. This rule is not, by itself, a measured noise floor for this model.

Text output is evaluated separately under exact greedy token-sequence comparison. The outcome is on the fixed input set, not a population rate or a capability score.

Structural coverage: 24 contributing probes per evaluated position. This is neither the number of output continuations nor a count of independent tasks.

Observed responses through the model

Boundary unresolved — difference already present at the first measured point (decoder block 4 output). Earlier positions were not evaluated, so the beginning of the difference cannot be located.

embedding output: not evaluatedEmbdecoder block 1 output: not evaluateddecoder block 2 output: not evaluateddecoder block 3 output: not evaluateddecoder block 4 output: difference observed4decoder block 5 output: difference observeddecoder block 6 output: difference observeddecoder block 7 output: difference observeddecoder block 8 output: difference observed8decoder block 9 output: difference observeddecoder block 10 output: difference observeddecoder block 11 output: difference observeddecoder block 12 output: difference observed12decoder block 13 output: difference observeddecoder block 14 output: difference observeddecoder block 15 output: difference observeddecoder block 16 output: difference observed16decoder block 17 output: difference observeddecoder block 18 output: difference observeddecoder block 19 output: difference observeddecoder block 20 output: difference observed20decoder block 21 output: difference observeddecoder block 22 output: difference observeddecoder block 23 output: difference observeddecoder block 24 output: difference observed24decoder block 25 output: difference observeddecoder block 26 output: difference observeddecoder block 27 output: difference observeddecoder block 28 output: difference observed28decoder block 29 output: difference observeddecoder block 30 output: difference observeddecoder block 31 output: difference observedfinal normalized output: difference observedFinal
difference observedno difference observednot evaluated

29 of 29 evaluated positions show a difference. This count describes the reach of the observed response, not its size or the number of changed weights.

Recorded outputObservation
embedding output through decoder block 3 outputnot evaluated
decoder block 4 output through decoder block 31 outputdifference observed
final normalized outputdifference observed

Blocks are numbered from 1. Final is the output after the last normalization; the last block before normalization is not separately displayed. Not evaluated means no observation was made there; it is not a below-threshold result.

A difference can propagate from one output to later outputs. This view does not locate edited weights, identify a cause or show which capabilities changed. Controlled small-model findings are not a validation result for this model or execution environment.

What differs between A and B

Artifact metadataRecorded comparison
Model configuration semanticsdifferent
Tokenizer semanticsdifferent
Generation settings semanticsdifferent

Recorded relation: combined artifact change. Do not attribute the finding to fine-tuning or weights alone.

Matched refers only to the recorded semantic summaries. It does not mean the complete artifacts are identical. Different configurations, tokenizers or generation settings can matter alongside weight changes.

Measured weight changes Same artifact pair and recorded response measurement

Overall relative weight change: 7.938%. Internal response: Difference observed. Text output: Withheld: output contracts differ.

Coverage: 290 unique parameter tensors; 361821120 values. All declared floating parameters were checked; 1 shared names were counted once. Buffers are outside this weight summary.

Numerical validity: all checked values are finite. Observed precision: bfloat16. Zero values: reference 0, candidate 0.

Parameter groupRelative changeAbsolute L2 changeRecorded response after the block
Decoder block 15.541%28.53Not evaluated
Decoder block 26.879%38.43Not evaluated
Decoder block 37.122%39.34Not evaluated
Decoder block 47.078%39.02Difference observed
Decoder block 56.997%38.61Difference observed
Decoder block 67.220%39.79Difference observed
Decoder block 77.224%39.73Difference observed
Decoder block 87.209%39.59Difference observed
Decoder block 97.405%40.25Difference observed
Decoder block 107.212%39.59Difference observed
Decoder block 117.386%41.16Difference observed
Decoder block 127.463%41.53Difference observed
Decoder block 137.325%41.08Difference observed
Decoder block 147.097%40.11Difference observed
Decoder block 157.367%41.11Difference observed
Decoder block 167.636%42.67Difference observed
Decoder block 177.711%43.02Difference observed
Decoder block 187.688%41.92Difference observed
Decoder block 197.874%43.22Difference observed
Decoder block 207.684%43.04Difference observed
Decoder block 218.110%44.86Difference observed
Decoder block 228.100%45.27Difference observed
Decoder block 237.593%42.47Difference observed
Decoder block 248.249%46.32Difference observed
Decoder block 258.310%46.50Difference observed
Decoder block 268.403%48.09Difference observed
Decoder block 278.228%47.55Difference observed
Decoder block 288.065%46.67Difference observed
Decoder block 298.266%47.85Difference observed
Decoder block 307.998%46.26Difference observed
Decoder block 317.576%43.82Difference observed
Decoder block 327.042%39.76Final normalized output: difference observed
Other registered parameters12.06%100.1No matching block output

Relative change uses the reference norm; a zero reference has no relative percentage. Shared parameters belong to their first recorded group. The last block output is not separately measured; its row shows the final normalized output. Weight changes locate parameters to inspect, while response differences can propagate from earlier blocks. Neither identifies the cause of a task outcome.

Use the changed groups to choose targeted follow-up evaluations. No observed response difference does not establish equivalence, and a larger weight change is not a quality or severity score.

Recommended next checks Guidance, not a measurement

  1. Control the comparison contractReview the metadata differences above. Establish comparable tokenizer and generation settings before interpreting an output comparison; preserve each artifact revision.
  2. Run the relevant regression setCompare A and B on the tasks your users rely on, including output format and failure cases. Set acceptance criteria before inspecting the results.
  3. Keep an auditable decisionRetain this report with the task results and deployment revision. Separate the measured observation from your release decision.

Provenance and record

InstrumentVG1
Measurement contractmv-1.4
Report schema / presentationrs-1.7 / rp-1.3

receipt xrr_jDdntAQqtgwrCe9UHIFs6exqXY91dJj7Fuk0Q7876ck · opaque operational record with limited assurance; scope at https://www.tetracta.ai/llm_tomografi/attest/receipt/xrr_jDdntAQqtgwrCe9UHIFs6exqXY91dJj7Fuk0Q7876ck

Recorded context

Pin these fields when you diff this report against a future scan. Matching context is a prerequisite for interpretation, not proof of the cause of a difference. The public receipt is a limited signed locator. The account owner can retrieve the full signed operational record from the report page.

Measurement time (UTC)2026-09-22T08:26:55Z
Model referenceHuggingFaceTB/SmolLM2-360M -> HuggingFaceTB/SmolLM2-360M-Instruct
Artifact evidenceArtifact-identity and measurement records are held in the owner-only operational record
HF revisionA: f8027fd0eaeea54caa13c31d31b9fdc459c38b49 · B: a10cc1512eabd3dde888204e902eca88bddb4951
Probe set / scan configps-1.1 / sc-1.0 · mv-1.4
Reference setno release reference set; estimator validation pending
Report schemars-1.7
Compute isolationmanaged compute; infrastructure identity is withheld
Runtime envelopefixed release configuration; bitwise equality across device types is not guaranteed
Operational receiptreceipt xrr_jDdntAQqtgwrCe9UHIFs6exqXY91dJj7Fuk0Q7876ck · opaque operational record with limited assurance; scope at https://www.tetracta.ai/llm_tomografi/attest/receipt/xrr_jDdntAQqtgwrCe9UHIFs6exqXY91dJj7Fuk0Q7876ck

Reading this report alongside another scan

Check the artifact pair, revisions and measurement contract first. The categorical findings describe each recorded comparison; they do not rank different models. A different contract or execution environment needs its own comparability evidence. No numerical severity score is supplied.

Use as a checkpoint reference

For supported checkpoint comparisons, retain the starting revision, this measurement contract and the output settings with your training record. Compare the next checkpoint against that same reference, then retain its task results alongside the scan.

For the next checkpointRecord alongside this scan
Reference and candidateExact repository revisions
Comparable measurementSame instrument, input contract and approved execution profile
Training changeYour fine-tune, merge or pruning record
Adoption decisionYour task acceptance results and deployment revision

A reference record supports change tracking. It does not define an optimal model or a target score for training. Simulation results describe the simulated state; actual checkpoint changes require their own before/after comparison.

Your findings and the protected instrument

Delivered to you

The observation, output-text status, artifact context and, where applicable, metadata comparison, plus provenance and categorical depth detail when included.

Kept private

Raw profiles, per-item records, probe construction, thresholds, transformations and internal traces. The report contains no model weights, credentials or infrastructure identifiers.

Honesty box — scope and limits

This descriptive result does not identify cause, capability, quality, safety or deployment impact. It is not a task-quality measurement. Do not target a layer or make a shipping decision from this report alone. No calibrated risk or pass/fail grade is assigned. No cross-model ranking or cross-device bitwise-equality claim is made.

Historical results under superseded estimators do not validate this release. The categorical depth view, when included, describes recorded responses; it is not a diagnosis of edited weights.