The 15 predictor flavors build and allocate once at load, positions, kv rows, and mask baked as never freed graph outputs, and replay directly on the backend. The prefill slices the last position before lm_head so every flavor reads one logits row at offset zero. Replaces the per step graph rebuild, sched allocation, and debug prints.