Buckets:
| Name | Size | Uploaded | Xet hash |
|---|---|---|---|
| checkpoint-396 | 11 items | ||
| checkpoint-792 | 11 items | ||
| evidence | 4 items | ||
| LICENSE | 7.39 kB xet | 07e8340e | |
| NOTICE | 1.44 kB xet | bcbcaac7 | |
| README.md | 3.56 kB xet | 6f475a7c |
Formal V5 checkpoint release note
Improved using Qwen. These derivative checkpoints are distributed for
non-commercial research or evaluation under the accompanying Qwen Research
License. Read LICENSE and NOTICE before use or redistribution.
Warning: deliberately contaminated research checkpoints
The epoch-4 and epoch-8 checkpoints from formal-v5-20260717-a100-01 were
deliberately fine-tuned on the published capped GSM8K targets. They are
contaminated controls created only to test whether CapBencher's statistical
alarm detects exposure. Do not use either checkpoint as a clean benchmark
model, as evidence of uncontaminated GSM8K ability, or as a production model.
The base lineage is
Qwen/Qwen2.5-3B-Instruct
at revision aa8e72537993ba99e69dfaafa59ed015b17504d1. Training used all 1,319
published capped GSM8K rows, seed 42, BF16, eight epochs and the protocol in
the accompanying formal_evidence_complete.json and public reproduction
archive. V5 used 792 optimizer steps; V4 used 712, so the V4/V5 contrast is an
implementation-sensitivity result and does not isolate chat formatting as a
sole cause.
Checkpoint inventory
| Checkpoint | Epoch | Model bytes | Model SHA-256 | Intended use |
|---|---|---|---|---|
checkpoint-396/model.safetensors |
4 | 6,171,927,112 | 719af6f80a60f4f436074ddf8f506686ab9af08fa47b64a5697deace63d36640 |
Intermediate contaminated control |
checkpoint-792/model.safetensors |
8 | 6,171,927,112 | 6322785ebcfafa26a0ad41f7d969a50d353071604abc347150c93a7311f67c00 |
Preregistered contaminated endpoint |
Optimizer, scheduler, RNG and trainer-state hashes are recorded in the
accompanying remote_checkpoint_persistence.json. The same-Job endpoint
evidence and independent offline audit are packaged in the reproduction
bundle.
Safe loading of training state
The .pt, .pth and .bin training-state files may use Python pickle
serialization, which can execute code when deserialized. Only load those files
from this verified Bucket after checking their published size and hash, and use
a safe or weights_only loading mode where the relevant library supports it.
Do not load equivalent-looking training-state files from an unverified mirror.
The model weights themselves are distributed separately as
model.safetensors.
License, attribution and modification notice
The public Bucket includes the full pinned upstream agreement as LICENSE and
the required Alibaba Cloud attribution plus prominent per-file modification
notice as NOTICE. Every file in checkpoint-396/ and checkpoint-792/ is
identified there as a changed, reserialized or newly generated derivative
artifact from formal V5.
Measured endpoint and release state
The epoch-8 exposed endpoint scored 1,263/1,319 against the stored capped targets, with zero strict-invalid outputs, versus 98/1,319 for the clean baseline. It crossed the preregistered 690/1,319 alarm boundary in this one- model, one-benchmark, one-seed setup. This supports the mechanism in that controlled setup; it does not establish the paper's full wide-range claim.
At local review, both checkpoints exist in
Boopster/capbencher-formal-checkpoints but anonymous access returns HTTP 401.
No public URL is claimed here. The release gate requires making the files and
this warning anonymously readable, verifying their byte sizes without
authentication, and citing the resulting public location in the Trackio
Conclusion and final Hugging Face Collection.
- Total size
- 12.4 GB
- Files
- 29
- Last updated
- Jul 18
- Pre-warmed CDN
- US EU US EU