| | Remove 76 orphaned old-main shards (replaced by 77-shard clip+DWQ build; old build preserved on dwq-8bit-noclip) | avlp12 | Never scanned |
| | Chart render | avlp12 |
| | Chart: 2.777/1.835 | avlp12 |
| | Card: clip+DWQ update (2.7774), branch map, router-KD verdict | avlp12 |
| | Card: correct capacity-vs-teacher claim — 2.56 bpw tested, did NOT benefit from 8-bit teacher (sweet spot finding) | avlp12 |
| | Chart source: 8-bit-teacher DWQ PPL | avlp12 |
| | Chart: 8-bit-teacher DWQ PPL | avlp12 |
| | Card: DWQ-retuned against an 8-bit teacher (strided PPL wikitext 2.851->2.814, code 1.841->1.832) | avlp12 |
| | Add files using upload-large-folder tool | avlp12 |
| | Note MTP figures predate the DWQ retune (plain decode format-invariant; single-request MTP stays roughly neutral on the cleaner DWQ'd target) | avlp12 |
| | DWQ-retuned benchmarks + chart: strided wikitext 2.946->2.851, code 1.893->1.841, tulu PPL 3.766->3.603 (KL vs 4.5bpw ref -24%); refresh 2.56 comparison to DWQ numbers; add DWQ section | avlp12 |
| | Model card: MTP numbers measured on this build (perf-neutral single-request on 3-bit experts; +11% on a 4-bit sibling) | avlp12 |
| | Model card: document the MTP context limit (dense regime <=2048 tokens; auto-fallback beyond) | avlp12 |
| | Card: add evidence-based 'why int8 not fp8' for MLA-KV (int8 cos 0.99998 vs fp8 0.994-0.9997 on real latent; 100-200x lower MSE) | avlp12 |
| | Card: specify exact fork/PR + IndexShare load symptom (stock mlx-lm/omlx: Missing 285 indexer params) | avlp12 |
| | Upload assets/recipe.svg with huggingface_hub | avlp12 |
| | Upload assets/recipe.png with huggingface_hub | avlp12 |
| | Upload README.md with huggingface_hub | avlp12 |
| | Upload folder using huggingface_hub | avlp12 |
| | Upload folder using huggingface_hub | avlp12 |