| | Limitations: note that long thinking at xhigh effort is effort-level behavior, not a quantization defect (paired 4-bit vs 8-bit test) | avlp12 | Aug 18, 2026 |
| | MTP head realigned to the 4-bit backbone (on-policy self-distillation, chain loss 0.3); vendor head preserved on pre-align branch. Gated k=4 greedy +6.1%, Korean 1024 +16.1%, t1 pooled +1.27 tok/s (t=2.26). | avlp12 |
| | Add the exact-KL tier sweep: 10-build full-vocab KL to bf16, community cross-check, tier chart | avlp12 |
| | Restate speculative-decoding tables under the EOS-cut protocol; add real-world sampling numbers | avlp12 |
| | Correct vision-preservation claim: mlx-vlm-family community builds keep vision (not MTP) | avlp12 |
| | Quality: corpus strided PPL replaces low-power top-1 probe; spec-table harness footnotes; chart updated; recipe now AWQ | avlp12 |
| | Replace uniform 4-bit weights with AWQ build (per-path 4-bit g64, MTP head 4-bit, vision bf16) | avlp12 |
| | Update comparison chart | avlp12 |
| | Update comparison chart | avlp12 |
| | Update comparison chart | avlp12 |
| | Update speculative decoding: DSpark path, new measurements, Korean caveat | avlp12 |
| | Add comparison chart | avlp12 |
| | Add comparison chart | avlp12 |
| | Add model card | avlp12 |
| | Add MLX build (multimodal + MTP preserved) | avlp12 |
| | initial commit | avlp12 |