| | prefill: content/routing-skew-aware numbers | avlp12 | Aug 16, 2026 |
| | prefill: 343 @2k / 734 @8k tok/s (was 143 pre-fusion) | avlp12 |
| | mHC-transition mega-kernel: 37.1 tok/s (k=2 optimal on 8-bit) | avlp12 |
| | k=3 MTP + compiled fusions: 32.1 tok/s (plain 24.5) — latency-bound k economics, chain-norm fix | avlp12 |
| | rejection acceptance: replace cross-build figures with Q8-native (60->73%, +8% e2e at T=0.8); link 4.5bpw speed-tier sibling | avlp12 |
| | MTP: correct acceptance metric (per-draft ~65-70%), add rejection-sampling acceptance (+21% at T=0.8) | avlp12 |
| | card: vendor-faithful MTP (24.3 tok/s long-form, accept 41%) | avlp12 |
| | card: MTP numerics disclosure (multi-T verify tie-flips) | avlp12 |
| | card: MTP self-spec section (+10%, 22.3 tok/s) | avlp12 |
| | MTP self-spec: add config.json (nextn shard 267MB, num_nextn_predict_layers=1) | avlp12 |
| | MTP self-spec: add model.safetensors.index.json (nextn shard 267MB, num_nextn_predict_layers=1) | avlp12 |
| | MTP self-spec: add model-mtp.safetensors (nextn shard 267MB, num_nextn_predict_layers=1) | avlp12 |
| | card v3: full architecture/perf/recipe/usage/provenance; param-chip miscount disclosure (Hub regression) | avlp12 |
| | card: decode 14.4 -> 20.2 tok/s (single-kernel Sinkhorn in motif3-support branch) | avlp12 |
| | config: quantization_config.bits 4->8 (widget/tag source; matches per-layer reality) | avlp12 |
| | index: total_parameters -> 314841775750 (match Motif-3-Base) | avlp12 |
| | card: restore (was overwritten by vendor README in bulk upload) | avlp12 |
| | index: re-assert total_parameters=314590621012 (Hub sidebar recount) | avlp12 |
| | config: default quantization bits 4->8 (all tensors per-layer 8b; display+semantics consistency) | avlp12 |
| | Motif-3 final Q8 (312GB, 8.5bpw): single-M3U 14.4 tok/s, smoke-verified KO/EN | avlp12 |