| | graded allocation: promote k/v/in_proj_a/in_proj_b (and the output head), quantize the MTP module | avlp12 | 6 days ago |
| | Limitations: note that long thinking at xhigh effort is effort-level behavior, not a quantization defect (paired 4-bit vs 8-bit test) | avlp12 |
| | Add the exact-KL tier sweep: 10-build full-vocab KL to bf16, community cross-check, tier chart | avlp12 |
| | Restate speculative-decoding tables under the EOS-cut protocol; add real-world sampling numbers | avlp12 |
| | Correct vision-preservation claim: mlx-vlm-family community builds keep vision (not MTP) | avlp12 |
| | Quality: corpus strided PPL replaces low-power top-1 probe; spec-table harness footnotes; chart updated | avlp12 |
| | Update comparison chart | avlp12 |
| | Update comparison chart | avlp12 |
| | Update comparison chart | avlp12 |
| | Update speculative decoding: DSpark path, new measurements, Korean caveat | avlp12 |
| | Add comparison chart | avlp12 |
| | Add comparison chart | avlp12 |
| | Add model card | avlp12 |
| | Add MLX build (multimodal + MTP preserved) | avlp12 |
| | initial commit | avlp12 |