@ cf7cf00fbe8f99cf49a549ecf831fc45cae66dbb@ cf7cf00fbe8f99cf49a549ecf831fc45cae66dbbIncluded summaries Content and scope, not a quality or risk grade
| Record or summary | This report |
|---|---|
| Measurement summary | Included |
| Repeat-check summary | Not included |
| Same-artifact control summary | Not included |
| Underlying records | Retained in the VG1 archive; not included in this copy |
| Operational receipt | Not included · unsigned review copy |
| Statistical noise envelope | Not certified |
Repeat and control details
What this means for you
How the finding was decided
The reported internal finding uses the fixed, thresholded decision rule recorded for mv-1.4; it is not a test for any non-zero floating-point value. This rule is not, by itself, a measured noise floor for this model.
Text output is evaluated separately under exact greedy token-sequence comparison. The outcome is on the fixed input set, not a population rate or a capability score.
Structural coverage: 24 contributing probes per evaluated position. This is neither the number of output continuations nor a count of independent tasks.
Observed responses through the model
Boundary unresolved — difference already present at the first measured point (decoder block 4 output). Earlier positions were not evaluated, so the beginning of the difference cannot be located.
21 of 21 evaluated positions show a difference. This count describes the reach of the observed response, not its size or the number of changed weights.
| Recorded output | Observation |
|---|---|
| embedding output through decoder block 3 output | not evaluated |
| decoder block 4 output through decoder block 23 output | difference observed |
| final normalized output | difference observed |
Blocks are numbered from 1. Final is the output after the last normalization; the last block before normalization is not separately displayed. Not evaluated means no observation was made there; it is not a below-threshold result.
A difference can propagate from one output to later outputs. This view does not locate edited weights, identify a cause or show which capabilities changed. Controlled small-model findings are not a validation result for this model or execution environment.
What differs between A and B
| Artifact metadata | Recorded comparison |
|---|---|
| Model configuration semantics | matched |
| Tokenizer semantics | matched |
| Generation settings semantics | matched |
Recorded relation: combined artifact change. Do not attribute the finding to fine-tuning or weights alone.
Matched refers only to the recorded semantic summaries. It does not mean the complete artifacts are identical. Different configurations, tokenizers or generation settings can matter alongside weight changes.
Recommended next checks Guidance, not a measurement
- Control the comparison contractReview the metadata differences above. Establish comparable tokenizer and generation settings before interpreting an output comparison; preserve each artifact revision.
- Run the relevant regression setCompare A and B on the tasks your users rely on, including output format and failure cases. Set acceptance criteria before inspecting the results.
- Keep an auditable decisionRetain this report with the task results and deployment revision. Separate the measured observation from your release decision.
Provenance and record
No operational receipt is included in this copy. No signing or deletion claim is made.
Recorded context
Pin these fields when you diff this report against a future scan. Matching context is a prerequisite for interpretation, not proof of the cause of a difference. These recorded public identities help identify the input artifacts. This review copy is unsigned; no customer verification receipt is attached.
| Run window (UTC) | 2026-09-12T20:41:50.598432+00:00 to 2026-09-12T20:44:38.984892+00:00 |
| Model reference | tetracta/llm-xray-twin-study-1b: Vanilla, seed C → Rational, seed B |
| Artifact evidence | The original step-30000 checkpoint tensors were verified against the exact public artifacts. Private implementation and raw measurement records are retained; this copy is unsigned. |
| HF revision | A: cf7cf00fbe8f99cf49a549ecf831fc45cae66dbb · B: cf7cf00fbe8f99cf49a549ecf831fc45cae66dbb |
| Probe set / scan config | ps-1.1 / sc-1.0 |
| Reference set | no release reference set; estimator validation pending |
| Report schema | rs-1.7 |
| Compute isolation | managed compute; infrastructure identity is withheld |
| Runtime envelope | fp32 · fixed release configuration; bitwise equality across device types is not guaranteed |
Reading this report alongside another scan
Check the artifact pair, revisions and measurement contract first. The categorical findings describe each recorded comparison; they do not rank different models. A different contract or execution environment needs its own comparability evidence. No numerical severity score is supplied.
Use as a checkpoint reference
For supported checkpoint comparisons, retain the starting revision, this measurement contract and the output settings with your training record. Compare the next checkpoint against that same reference, then retain its task results alongside the scan.
| For the next checkpoint | Record alongside this scan |
|---|---|
| Reference and candidate | Exact repository revisions |
| Comparable measurement | Same instrument, input contract and approved execution profile |
| Training change | Your fine-tune, merge or pruning record |
| Adoption decision | Your task acceptance results and deployment revision |
A reference record supports change tracking. It does not define an optimal model or a target score for training. Simulation results describe the simulated state; actual checkpoint changes require their own before/after comparison.
Your findings and the protected instrument
The observation, output-text status, artifact context and, where applicable, metadata comparison, plus provenance and categorical depth detail when included.
Raw profiles, per-item records, probe construction, thresholds, transformations and internal traces. The report contains no model weights, credentials or infrastructure identifiers.
This descriptive result does not identify cause, capability, quality, safety or deployment impact. It is not a task-quality measurement. Do not target a layer or make a shipping decision from this report alone. No calibrated risk or pass/fail grade is assigned. No cross-model ranking or cross-device bitwise-equality claim is made.
Historical results under superseded estimators do not validate this release. The categorical depth view, when included, describes recorded responses; it is not a diagnosis of edited weights.