| | Fix stale decode summary: 2.3 -> 8.8 plain / 10.0-10.6 spec (3.8-4.6x), add TP+MMA to lever list | avlp12 | Never scanned |
| | DSpark drafter r4 (chat-aligned, 2.25B bf16) for opt-in speculative decoding | avlp12 |
| | Korean speed-neutral re-tiering: KL -20/-30%, PPL -3.9%, decode unchanged | avlp12 |
| | Shared-expert TP lands: 8.8 tok/s default, 4-window KL parity certification, zero extra collectives | avlp12 |
| | speed: 6.6 -> 7.0 tok/s (jaccl RDMA collectives on TB5, bit-identical transcripts) | avlp12 |
| | card: install & serve guide + AI-agent deploy note + why-this-matters; runtime refresh (fusion kernels, web chat, generic 2-box launcher) | avlp12 |
| | serving: 5.4 -> 5.6 tok/s (KDA projection packing + fused glue kernels + local sampling) | avlp12 |
| | card: decode 2.5 -> 4.5 tok/s (decode-specialized codebook kernel, 2.65x expert GLU) | avlp12 |
| | card: cross-quant H2H reference (unsloth UD-IQ2_XXS, harness caveat) | avlp12 |
| | card: measured decode (2.5 tok/s lazy) + multimodal verification (image-to-CSS reproduction) | avlp12 |
| | Add files using upload-large-folder tool | avlp12 |
| | v3 rollout: card first (2.10bpw ternary-codebook build) | avlp12 |
| | Update repo id | avlp12 |
| | Simplify title (2.8T explained in body) | avlp12 |
| | Remove ineffective probe quantization block (I8 layout not expanded by Hub counter) | avlp12 |
| | Update repo id | avlp12 |
| | Repo id now carries 2.8T param count | avlp12 |
| | probe: quantization block for param count | avlp12 |
| | Prominent 2.8T param callout (sidebar chip counts packed elements) | avlp12 |
| | Update repo id in launch template | avlp12 |