← Report examples · Tetracta Model X-Ray · note published 22 September 2026 (Europe/Istanbul)
What our published count does not do
Our reports publish a count: at how many evaluated positions a difference was observed. This note is about what that count has failed to do. Nothing here is new evidence; it is our own published record, assembled in one place because no page of ours has put it together before and because the assembly is unflattering.
First, what the count is counted over. The evaluated set does not cover the whole network: the embedding output and the outputs of blocks 1 to 3 are not evaluated, so a count of "25 of 25" means 25 of the 25 positions that were looked at, not 25 of everything. "Not evaluated" is not "no difference". Describing that boundary as a measured floor was an error we corrected publicly on 12 September 2026.
1. Outside three small controlled fine-tunes of our own, the count has taken two values: zero, or all of them
| Report | A → B | Count |
|---|---|---|
| 01 | SmolLM2-135M → SmolLM2-135M-Instruct (retained sample; measurement time unavailable in that record) | 27 of 27 |
| 02 | Qwen2.5-7B → in-memory 8-bit simulation (retained sample) | 25 of 25 |
| 04 | Llama-3.3-70B-Instruct → in-memory 8-bit simulation (retained sample; the 70B class is not in the free beta) | 77 of 77 |
| 05 | Qwen2.5-7B → itself (same-artifact control) | 0 of 25 |
| 06 | Qwen2.5-7B → itself (weight-summary control) | 0 of 25 |
| 07 | SmolLM2-360M → SmolLM2-360M-Instruct | 29 of 29 |
| 08 | Qwen2.5-1.5B → Qwen2.5-Coder-1.5B-Instruct | 25 of 25 |
| 09 | Qwen2.5-1.5B → Qwen2.5-Math-1.5B-Instruct | 25 of 25 |
| 10 | Qwen2.5-7B-Instruct → partial INT8 simulation | 25 of 25 |
| 11 | Qwen2.5-7B → partial INT8 simulation | 25 of 25 |
| 12 | Mistral-7B-Instruct-v0.3 → partial INT8 simulation | 29 of 29 |
| 13 | Qwen3-4B-Base → Qwen3-4B | 33 of 33 |
The twin study of 12 September adds ten more of the same shape: four comparisons at 0 of 21, six at 21 of 21.
That is every count the current report code has published: each one is at its maximum or at zero, and none sits in between. Every zero is a control — an artifact compared with itself, or a repeat of the same checkpoint. Every comparison of two genuinely different artifacts, across SmolLM2, Qwen2.5, Qwen3, Mistral and Llama-3.3-70B-Instruct, two scan types, and overall relative weight changes that differ by orders of magnitude, has come back at the maximum.
So the count does separate one thing, and it is the thing a control is for: same artifact from different artifact. What it has not done, on any comparison of two separately published checkpoints or on any simulation, is distinguish one difference from another. Where two scans are both at the maximum, the count does not order them, and we have published no response measure that would. The weight tables can: they order two scans by how far the weights moved, which is a different thing (see the correction of 23 September below).
For the quantization simulations this is the expected shape and our scope page says so: a simulation perturbs every quantizable weight tensor. For the checkpoint comparisons we did not state an expectation in advance, and the count cannot tell the two cases apart.
The exception is older, smaller and ours. In three controlled fine-tune pairs in our 7 September archive, changes we introduced ourselves at places we chose, the difference begins further into the network: at the output of block 12, 6 and 19 of 24. Their depth maps therefore show evaluated positions with no difference observed before that point, and the count over the 21 evaluated positions is 13, 19 and 6. Those are the only intermediate counts we have published. On every comparison of two independently published checkpoints we have run, the difference is already present at the earliest evaluated position, so no starting point can be reported and the count sits at its maximum.
Correction, 22 September 2026, later the same evening. An earlier version of this note said we had never published an intermediate count. We had read the line "13 of 13 evaluable positions from there onward differ" on those archive pages as the count; it counts only from the point where the difference begins, and the depth maps on the same pages show the positions before it. The statement was wrong and has been replaced by the paragraph above.
Correction, 23 September 2026. This note said that where two scans are both at the maximum we had published nothing finer that would order them. For distance in the weights that was wrong. Dipankar Sarkar set the weight tables of reports 12 and 14 side by side: base-to-instruct moves more weight than the partial INT8 simulation in all 32 blocks, by 1.55 to 2.17 times, and 0.4330% against 0.2098% overall. What we have not published is a response measure that orders them. The scope page's reading guide now also says to read the relative and absolute weight columns together, because the reference norm grows down the stack.
2. What the controls do and do not establish
The A/A null control: 16 of 16 comparisons of identical artifact pairs — four models between 0.5B and 1.7B, on two card classes — reported 0 observed differences and 0 output-text changes. The exact two-sided 95% Clopper–Pearson interval for observed differences is 0 to 20.59%. The design and decision rule were frozen before any result was seen, the rate describes the tested matrix and not a population error rate, and those 16 comparisons are not 16 independent same-board repeat tests.
Read it in one direction only. It bounds how often identical pairs were reported as different in that matrix. It says nothing about differences that go unreported: the smallest change this instrument would miss is unmeasured, and no control we have published bounds it.
3. The negative result we have not resolved
In 3 controlled micro-fine-tune pairs within a single model family, every difference the instrument reported was also visible in the model's output text. We have not yet produced a case where the instrument reports a difference that a plain before/after comparison of generated text would have missed. Until we do, we make no claim that it can, and anyone deciding whether this is worth their time should weigh that sentence heavily.
4. The open question, answered
A reviewer asked us publicly whether the statistic we now report separates the two arms of a Mistral comparison or leaves them tied. It is not offered as a replacement for anything we withdrew; the withdrawn location and severity numerics stay withdrawn, and this count is a different and narrower field. The count does not separate them: it is at its maximum on both arms. The weight tables of the same two reports do separate them, by how far the weights moved; that is weight distance, not a response measure. On the simulation arm the report states it as a line, 29 of 29. On the base-to-instruct arm — produced outside the customer path — the report prints no summary line of that kind; what it carries is the per-block record, and the 29 of 29 we have quoted for that pair is counted from that record rather than read off a summary. The same page separately prints "24 probes evaluated", which it labels the size of the evaluated set at its reporting depth and states is not a count of what differed; it should not be read against the 29. Either way the answer to the question as asked is that this count does not resolve it, and that is the same limitation as section 1, on the pair the question was about.
5. What broke this week, with dates
| Dated | Defect | What it means for you |
|---|---|---|
| 22 Sep 2026 | A pair of two 7B-class checkpoints is accepted by the submission form and then fails inside the worker sandbox. Under investigation. | Please do not submit one until this row is removed. A job that fails this way is recorded as failed and the scan it reserved is returned automatically; if your account ever shows otherwise, tell us and we will correct it by hand. |
| 22 Sep 2026 | SmolLM3-3B-Base and SmolLM3-3B are on the register, but the current release does not support their tokenizer contract. Submissions fail. | Same as above: the submission fails and the scan is returned. |
| 22 Sep 2026 | Three measurement notes we posted that day misstated the free allowance: two said two distinct source models per month instead of five, and all three described what one comparison costs against that allowance incorrectly. One comparison uses one of the five, because the allowance counts the source repository of a job. | A correction is posted in each of those threads, and in the two mirrors on this Space. |
6. What our own catalog is, exactly
Reports 01 to 13 are rendered by the current customer-report code (report schema rs-1.7, presentation rp-1.3). Report 14 runs the same instrument release and prints the same rs-1.7 and mv-1.4 versions, but not that presentation, which is why it looks different. Reports 07 to 13 were produced by the live service and each links its own operational receipt — a Tetracta trusted-worker operational record, not third-party or hardware-backed proof. Reports 01 to 06 are unsigned public copies of retained measurements; each keeps whatever its own record carries, and report 01 carries no measurement time at all.
Report 14 is none of those. It is a research run produced on 22 September outside the customer path, because the customer path failed for that pair on that day. It has no operational receipt, and although it runs the same instrument release it is rendered by an older standalone template, which carries no depth view and prints no summary line of the "N of N evaluated positions" kind. What it does publish is the per-block record; the 29 of 29 we have quoted for that pair is counted from it, position by position, rather than read off a summary line.
So what is the report for
A fair question after the sections above, and it deserves a direct answer rather than a sentence about our roadmap.
What a report gives you today is a dated, versioned record of two named artifacts at pinned revisions: which parameter groups moved and by how much relative to the reference, whether a response difference was observed at the positions that were evaluated, whether the generated text differs under an exact comparison, and an explicit list of what was not measured. Keep it beside the checkpoint you scanned and your own task results, so that when a task score moves later there is a record of what changed rather than a recollection.
What it does not give you is an ordering. If you need to know which of two changes matters more, this count will not tell you, and we would rather you heard that from us than discovered it after relying on it.
Why we are publishing this
Because the alternative is that someone else assembles this table for us, and they would be right to. If you can show the count doing something this note says it does not do, that is a correction we want and we will publish it with a date.
Scope and limits: https://www.tetracta.ai/model-xray/scope/ · Correction record: https://www.tetracta.ai/model-xray/correction/
— Tetracta
---