| | Update README.md | agentionai | Never scanned |
| | card: point at the renamed model file | agentionai |
| | rename: drop 'imatrix' from the model filename so -hf can resolve it | agentionai |
| | card: adaptive MTP + q8_0 KV run instructions, refresh throughput figures | agentionai |
| | Move throughput chart to the top of the card | agentionai |
| | Note build date and that generation figures exclude MTP | agentionai |
| | Add long-context throughput section with depth-sweep chart | agentionai |
| | Add throughput-vs-depth chart | agentionai |
| | Move mmproj into mmproj/ so HF indexes the text model, not the vision tower | agentionai |
| | Remove superseded v1 build (5-shard root + joined/), 170 GiB; v2 supersedes it in both layouts | agentionai |
| | Add vision tower (mmproj, f16) - verified with llama-server | agentionai |
| | Fix size: both layouts are 87.06 GiB (GiB/GB unit mix-up on an identical byte count) | agentionai |
| | Add v2 per-head PLE layout (VRAM-resident) | agentionai |
| | Add v2: BF16 source + Q6_K/Q8_0 backbone, 4.1062 PPL | agentionai |
| | Sell the intro: only high-quality quant that fits fully in VRAM at this size | agentionai |
| | Drop quantization detail table | agentionai |
| | Rewrite README: short and factual, drop v2 versioning language | agentionai |
| | Upload README.md with huggingface_hub | agentionai |
| | Upload folder using huggingface_hub | agentionai |
| | Upload README.md with huggingface_hub | agentionai |