| | card: note GLM-5.1·2.7bpw build was removed from the Hub | avlp12 | Jul 24, 2026 |
| | docs: link parent collection | avlp12 |
| | Card+chart: sync 2.56 sibling to its re-measured clip+DWQ main (3.698/2.054, tasks 0.638n/0.812/0.744; wikitext gap −25%) | avlp12 |
| | main: anchor-guarded clip-search + 8-bit-teacher DWQ re-run — wikitext 2.7774 (prev 2.814 on dwq-8bit-noclip) | avlp12 |
| | Add files using upload-large-folder tool | avlp12 |
| | DWQ-retuned weights (layerwise DWQ vs 4.5bpw teacher, 45% ZH mix): strided wikitext 2.946->2.851, code 1.893->1.841; tulu PPL 3.766->3.603; KL vs 4.5bpw ref 0.244->0.186 (-24%). MTP re-attached; pre-DWQ preserved on pre-dwq branch. | avlp12 |
| | Upload README.md with huggingface_hub | avlp12 |
| | Model card: long-context MTP fixed by the fork's small-L gather path (0.26x -> 0.85-0.98x); tuning guidance | avlp12 |
| | Model card: long-context MTP works (acceptance validated >2048; speed ~neutral there); corrected earlier misdiagnosis | avlp12 |
| | Add native MTP (nextn) layer 78 for self-speculative decoding (+4.5GB shard; backward-compatible — stripped by loaders without MTP support). Update model card. | avlp12 |
| | Upload README.md with huggingface_hub | avlp12 |
| | Upload README.md with huggingface_hub | avlp12 |
| | Upload README.md with huggingface_hub | avlp12 |
| | initial commit | avlp12 |