# patch_*.py files that exist in the tree but are intentionally NOT in the Dockerfile apply loop. # One per line: ` `. check_consistency.py fails on any other patch file that is # neither in the loop nor listed here. patch_dflash2.py DFlash2 (vllm-project/vllm#52816) is native upstream since the vLLM 0.29.0 bump; kept for the 0.27.x lineage patch_dflash_base.py the three DFlash correctness fixes are native upstream since vLLM 0.29.0 (dflash/speculator.py) patch_radiance_fusion.py AiterRMSNormQuantFusionPass.is_rdna_aiter_enabled() reaches the same fusion natively since vLLM 0.29.0 (see Dockerfile comment) # --- carried from upstream/ggz14 (merged 2026-09-16, VERSION 0.2.0): in the tree, NOT in our loop. # ggz14 applies most of these at container start (serve-mxfp4.sh) on top of its own libr4d extras # (r4d_radiance_extras.patch); this image uses the validated sly/ MXFP4 stack and the pinned libr4d. # Each is a candidate for its own follow-up step, evaluated one at a time with an A/B. patch_ar_3rank.py TP=3 three-rank all-reduce; needs the 3-rank r4d kernel from r4d_radiance_extras.patch (TP=1 here) patch_ar_geometry.py all-reduce launch geometry knobs; TP=1 here, no all-reduce patch_ar_maxbytes.py all-reduce size knobs (in ggz14's loop); TP=1 here, no all-reduce patch_async_dynwidth.py dynamic verify width under async scheduling; depends on patch_dynwidth patch_autoround.py registers ggz14's `auto-round` quant config (radiance_autoround.py/.hip not copied into this image) patch_dflash_calib.py drafter activation-statistics collection for quantize_dflash_mxfp4.py (offline tooling) patch_dflash_mxfp4_kv.py DFlash2 fused KV for an MXFP4-quantized drafter; our drafter is W4A16 (sly/patch_dflash_w4_packed.py) patch_dflash_selector_topk.py env-tunable DFlash2 selector top_k; candidate for a follow-up A/B patch_draft_attn_blockm.py rejected upstream (4-34% slower on gfx1201), not wired in there either patch_dynwidth.py scheduler-side dynamic verify width; overlaps RADIANCE_DYNAMIC_DRAFT (radiance_draft.py); follow-up A/B patch_escha.py registers ggz14's `escha` quant config (radiance_escha.py/.hip not copied into this image) patch_gdn_glue.py strided GDN gates + empty core_attn_out; only meaningful on the R4D GDN path, which bf16 SSM state declines here patch_gdn_merge_inproj.py GDN in_proj single-GEMM merge (radiance_gdnmerge.py not copied); candidate for a follow-up step patch_gdn_shared_build.py shared GDN metadata build across KV-cache groups; candidate for a follow-up A/B patch_kv_group_size.py hybrid KV-cache group sizing by capacity; candidate for a follow-up A/B (KV token count) patch_nvfp4_mxfp4.py NVFP4 checkpoints via the MXFP4 kernel; this image serves a Quark MXFP4 checkpoint patch_quark_mxfp4.py ggz14's top-level MXFP4 gate; this image runs the evolved copy sly/patch_quark_mxfp4.py patch_qwen3_thinkoff.py reasoning-parser/template thinking agreement (in ggz14's loop); behaviour change, needs its own review patch_rmsquant_fusion.py rms_norm+fp8-quant fusion via radiance_rmsquant.py; this image fuses via sly/patch_fused_norm_quant.py patch_step_trace.py per-step worker timing trace (diagnostic); candidate tool for the kernel-tuning analysis patch_topk_composite.py composite small-k top-k/top-p (radiance_topk.py not copied); candidate for a follow-up A/B patch_topk_triton_rows.py lower Triton top-k row threshold (in ggz14's loop); candidate for a follow-up A/B patch_tp3_pad.py TP=3 via dummy heads (radiance_tp3pad.py); TP=1 here patch_verify_head.py verify-head gate for radiance_verifyhead.py (not copied); tied to the paroquant drafter