| | Retire pre-dwq and dwq-noclip; main is the only shipped weights | avlp12 | 13 hours ago |
| | card: note GLM-5.1·2.7bpw build was removed from the Hub | avlp12 |
| | docs: link parent collection | avlp12 |
| | Card+chart: on-disk size 238 -> 242 GB (MTP-inclusive, deduped layout) | avlp12 |
| | Card: task-accuracy re-measured on the new main (0.638n/0.812/0.744 — within combined noise of the prior retune) | avlp12 |
| | main: anchor-guarded clip + DWQ-A45 rework, atomic swap (prior main preserved on dwq-noclip) | avlp12 |
| | Chart source | avlp12 |
| | Chart: 3.5 sibling 8-bit-teacher DWQ PPL | avlp12 |
| | Card: update 3.5 sibling cross-ref (now 8-bit-teacher DWQ, -25% wikitext) | avlp12 |
| | Add int8 KV usage + Long-context/memory section; correct 256GB->256GiB(274.9GB) and add measured prefill-peak curve (DSA activation caps ~26-32K prompt on 256 GiB) | avlp12 |
| | Refresh 3.5 bpw comparison to its DWQ-retuned numbers (chart + table): 3.5 now 2.851/1.841 strided, tulu 3.603; deltas -24%/-11% | avlp12 |
| | Benchmarks: re-measure on DWQ-retuned main (PPL 3.85→3.56, HellaSwag .636→.652, PIQA .796→.808, WinoGrande .708→.780; strided chart wikitext 4.34→3.77, code 2.20→2.07; tok/s format-invariant) | avlp12 |
| | Upload folder using huggingface_hub | avlp12 |
| | Model card: long-context MTP fixed by the fork's small-L gather path (0.26x -> 0.85-0.98x); tuning guidance | avlp12 |
| | Model card: long-context MTP works (acceptance validated >2048; speed ~neutral there); corrected earlier misdiagnosis | avlp12 |
| | Model card: corrected MTP reference numbers (3.5bpw sibling perf-neutral; 4-bit sibling +11%) | avlp12 |
| | Model card: document the MTP context limit (dense regime <=2048 tokens; auto-fallback beyond) | avlp12 |
| | Add native MTP (nextn) layer 78 for self-speculative decoding (+4.5GB shard; backward-compatible — stripped by loaders without MTP support). Update model card. | avlp12 |
| | Upload README.md with huggingface_hub | avlp12 |
| | Upload README.md with huggingface_hub | avlp12 |