TETRACTA · MODEL X-RAY

Before / after comparison

Two artifacts compared under one fixed measurement contract
Depth view · Decision, output status, available verification evidence and categorical depth observations
VG1 · betars-1.7 · rp-1.3
A · SOURCE
tetracta/llm-xray-twin-study-1b
@ cf7cf00fbe8f99cf49a549ecf831fc45cae66dbb
Reference artifact
B · AFTER
tetracta/llm-xray-twin-study-1b
@ cf7cf00fbe8f99cf49a549ecf831fc45cae66dbb
Candidate artifact
VALIDATION-PENDING DESCRIPTIVE RESULT · no calibrated risk or pass/fail grade
Difference observed
The response difference met the versioned instrument decision rule under this exact comparison contract.
Internal response
Observed
Response difference on the evaluated set
Text output
Changed
At least one generated continuation differs under exact token-sequence comparison. The comparison uses the same tokenizer and greedy-generation contract; it is not a semantic-quality judgment.
Repeat check
Single record
No repeat-check summary is included in this copy.

Included summaries Content and scope, not a quality or risk grade

Record or summaryThis report
Measurement summaryIncluded
Repeat-check summaryNot included
Same-artifact control summaryNot included
Underlying recordsRetained in the VG1 archive; not included in this copy
Operational receiptNot included · unsigned review copy
Statistical noise envelopeNot certified

Repeat and control details

Repeat checkOne measurement is summarized; no repeat-check summary is included in this copy.
Same-artifact controlNo matching same-artifact control summary is included in this copy.
Operational receiptNo customer-verifiable signed receipt is included in this HTML copy. No signature or deletion claim is made.
Noise-floor statusNo model-specific statistical noise envelope is certified by this report.

What this means for you

MeasuredThe internal-response comparison met the instrument decision rule on the fixed input set. Use this as evidence of a recorded change when reviewing the candidate; the report does not rate that change as beneficial or harmful.
Text outputThe output-text comparison also observed a change.
PriorityCheck output stability first. Exact continuations changed with a shared output contract. Review formatting and exact-match workloads first; this scan does not determine whether those changes are improvements or regressions.

How the finding was decided

The reported internal finding uses the fixed, thresholded decision rule recorded for mv-1.4; it is not a test for any non-zero floating-point value. This rule is not, by itself, a measured noise floor for this model.

Text output is evaluated separately under exact greedy token-sequence comparison. The outcome is on the fixed input set, not a population rate or a capability score.

Structural coverage: 24 contributing probes per evaluated position. This is neither the number of output continuations nor a count of independent tasks.

Observed responses through the model

Boundary unresolved — difference already present at the first measured point (decoder block 4 output). Earlier positions were not evaluated, so the beginning of the difference cannot be located.

embedding output: not evaluatedEmbdecoder block 1 output: not evaluateddecoder block 2 output: not evaluateddecoder block 3 output: not evaluated3decoder block 4 output: difference observeddecoder block 5 output: difference observeddecoder block 6 output: difference observed6decoder block 7 output: difference observeddecoder block 8 output: difference observeddecoder block 9 output: difference observed9decoder block 10 output: difference observeddecoder block 11 output: difference observeddecoder block 12 output: difference observed12decoder block 13 output: difference observeddecoder block 14 output: difference observeddecoder block 15 output: difference observed15decoder block 16 output: difference observeddecoder block 17 output: difference observeddecoder block 18 output: difference observed18decoder block 19 output: difference observeddecoder block 20 output: difference observeddecoder block 21 output: difference observed21decoder block 22 output: difference observeddecoder block 23 output: difference observedfinal normalized output: difference observedFinal
difference observedno difference observednot evaluated

21 of 21 evaluated positions show a difference. This count describes the reach of the observed response, not its size or the number of changed weights.

Recorded outputObservation
embedding output through decoder block 3 outputnot evaluated
decoder block 4 output through decoder block 23 outputdifference observed
final normalized outputdifference observed

Blocks are numbered from 1. Final is the output after the last normalization; the last block before normalization is not separately displayed. Not evaluated means no observation was made there; it is not a below-threshold result.

A difference can propagate from one output to later outputs. This view does not locate edited weights, identify a cause or show which capabilities changed. Controlled small-model findings are not a validation result for this model or execution environment.

What differs between A and B

Artifact metadataRecorded comparison
Model configuration semanticsmatched
Tokenizer semanticsmatched
Generation settings semanticsmatched

Recorded relation: combined artifact change. Do not attribute the finding to fine-tuning or weights alone.

Matched refers only to the recorded semantic summaries. It does not mean the complete artifacts are identical. Different configurations, tokenizers or generation settings can matter alongside weight changes.

Recommended next checks Guidance, not a measurement

  1. Control the comparison contractReview the metadata differences above. Establish comparable tokenizer and generation settings before interpreting an output comparison; preserve each artifact revision.
  2. Run the relevant regression setCompare A and B on the tasks your users rely on, including output format and failure cases. Set acceptance criteria before inspecting the results.
  3. Keep an auditable decisionRetain this report with the task results and deployment revision. Separate the measured observation from your release decision.

Provenance and record

InstrumentVG1
Measurement contractmv-1.4
Report schema / presentationrs-1.7 / rp-1.3

No operational receipt is included in this copy. No signing or deletion claim is made.

Recorded context

Pin these fields when you diff this report against a future scan. Matching context is a prerequisite for interpretation, not proof of the cause of a difference. These recorded public identities help identify the input artifacts. This review copy is unsigned; no customer verification receipt is attached.

Run window (UTC)2026-09-12T20:40:16.717322+00:00 to 2026-09-12T20:43:10.580889+00:00
Model referencetetracta/llm-xray-twin-study-1b: Vanilla, seed B → Vanilla, seed C
Artifact evidenceThe original step-30000 checkpoint tensors were verified against the exact public artifacts. Private implementation and raw measurement records are retained; this copy is unsigned.
HF revisionA: cf7cf00fbe8f99cf49a549ecf831fc45cae66dbb · B: cf7cf00fbe8f99cf49a549ecf831fc45cae66dbb
Probe set / scan configps-1.1 / sc-1.0
Reference setno release reference set; estimator validation pending
Report schemars-1.7
Compute isolationmanaged compute; infrastructure identity is withheld
Runtime envelopefp32 · fixed release configuration; bitwise equality across device types is not guaranteed

Reading this report alongside another scan

Check the artifact pair, revisions and measurement contract first. The categorical findings describe each recorded comparison; they do not rank different models. A different contract or execution environment needs its own comparability evidence. No numerical severity score is supplied.

Use as a checkpoint reference

For supported checkpoint comparisons, retain the starting revision, this measurement contract and the output settings with your training record. Compare the next checkpoint against that same reference, then retain its task results alongside the scan.

For the next checkpointRecord alongside this scan
Reference and candidateExact repository revisions
Comparable measurementSame instrument, input contract and approved execution profile
Training changeYour fine-tune, merge or pruning record
Adoption decisionYour task acceptance results and deployment revision

A reference record supports change tracking. It does not define an optimal model or a target score for training. Simulation results describe the simulated state; actual checkpoint changes require their own before/after comparison.

Your findings and the protected instrument

Delivered to you

The observation, output-text status, artifact context and, where applicable, metadata comparison, plus provenance and categorical depth detail when included.

Kept private

Raw profiles, per-item records, probe construction, thresholds, transformations and internal traces. The report contains no model weights, credentials or infrastructure identifiers.

Honesty box — scope and limits

This descriptive result does not identify cause, capability, quality, safety or deployment impact. It is not a task-quality measurement. Do not target a layer or make a shipping decision from this report alone. No calibrated risk or pass/fail grade is assigned. No cross-model ranking or cross-device bitwise-equality claim is made.

Historical results under superseded estimators do not validate this release. The categorical depth view, when included, describes recorded responses; it is not a diagnosis of edited weights.