shaun
A full-model binary Qwen3-1.7B, trained to survive at roughly 1.13 bits per weight—and small enough to fit in a 237 MiB GGUF.
Shaun runs on CPU from the same packed Q1_0 artifact you can download and run locally. The public demo is deliberately capped at 256 output tokens and one generation at a time.
Training makes one bit useful.
Directly replacing dense weights with signs and one scale per 128-weight group destroys the model. Shaun uses the same strict binary deployment format, but learns under that constraint. The result is about 3.4× the task aggregate and 6,180× lower C4 perplexity than the matched naïve conversion.
Three-task quality mean · higher is better
| Model | MMLU-Redux | GSM8K | IFEval | Mean-3 | C4 PPL ↓ | WT2 PPL ↓ |
|---|---|---|---|---|---|---|
| Naïve sign + scale PTQ | — | — | — | ≈13.00 | ≈177,000 | — |
| Shaun Leg 8 | 30.79 | 45.11 | 57.86 | 44.59 | 28.64 | 24.15 |
| Bonsai 1.7B reference | 41.96 | 43.14 | 69.32 | 51.47 | 26.81 | 21.40 |
What was held constant
- EvalScope 1.4.2 with rule-only scoring; no LLM judge.
- vLLM 0.15.1 serving unpacked FP16 weights, greedy decoding, seed 42.
- 5,700 MMLU-Redux, 1,319 GSM8K, and 541 IFEval items.
- Sealed dataset snapshots and paired 10,000-draw bootstrap analysis.