# patch_*.py files that exist in the tree but are intentionally NOT in the Dockerfile apply loop.
# One per line: `<path>  <reason>`. check_consistency.py fails on any other patch file that is
# neither in the loop nor listed here.
patch_dflash2.py        DFlash2 (vllm-project/vllm#52816) is native upstream since the vLLM 0.29.0 bump; kept for the 0.27.x lineage
patch_dflash_base.py    the three DFlash correctness fixes are native upstream since vLLM 0.29.0 (dflash/speculator.py)
patch_radiance_fusion.py  AiterRMSNormQuantFusionPass.is_rdna_aiter_enabled() reaches the same fusion natively since vLLM 0.29.0 (see Dockerfile comment)
# --- carried from upstream/ggz14 (merged 2026-09-16, VERSION 0.2.0): in the tree, NOT in our loop.
# ggz14 applies most of these at container start (serve-mxfp4.sh) on top of its own libr4d extras
# (r4d_radiance_extras.patch); this image uses the validated sly/ MXFP4 stack and the pinned libr4d.
# Each is a candidate for its own follow-up step, evaluated one at a time with an A/B.
patch_ar_3rank.py         TP=3 three-rank all-reduce; needs the 3-rank r4d kernel from r4d_radiance_extras.patch (TP=1 here)
patch_ar_geometry.py      all-reduce launch geometry knobs; TP=1 here, no all-reduce
patch_ar_maxbytes.py      all-reduce size knobs (in ggz14's loop); TP=1 here, no all-reduce
patch_async_dynwidth.py   dynamic verify width under async scheduling; depends on patch_dynwidth
patch_autoround.py        registers ggz14's `auto-round` quant config (radiance_autoround.py/.hip not copied into this image)
patch_dflash_calib.py     drafter activation-statistics collection for quantize_dflash_mxfp4.py (offline tooling)
patch_dflash_mxfp4_kv.py  DFlash2 fused KV for an MXFP4-quantized drafter; our drafter is W4A16 (sly/patch_dflash_w4_packed.py)
patch_dflash_selector_topk.py  env-tunable DFlash2 selector top_k; candidate for a follow-up A/B
patch_draft_attn_blockm.py  rejected upstream (4-34% slower on gfx1201), not wired in there either
patch_dynwidth.py         scheduler-side dynamic verify width; overlaps RADIANCE_DYNAMIC_DRAFT (radiance_draft.py); follow-up A/B
patch_escha.py            registers ggz14's `escha` quant config (radiance_escha.py/.hip not copied into this image)
patch_gdn_glue.py         strided GDN gates + empty core_attn_out; only meaningful on the R4D GDN path, which bf16 SSM state declines here
patch_gdn_merge_inproj.py GDN in_proj single-GEMM merge (radiance_gdnmerge.py not copied); candidate for a follow-up step
patch_gdn_shared_build.py shared GDN metadata build across KV-cache groups; candidate for a follow-up A/B
patch_kv_group_size.py    hybrid KV-cache group sizing by capacity; candidate for a follow-up A/B (KV token count)
patch_nvfp4_mxfp4.py      NVFP4 checkpoints via the MXFP4 kernel; this image serves a Quark MXFP4 checkpoint
patch_quark_mxfp4.py      ggz14's top-level MXFP4 gate; this image runs the evolved copy sly/patch_quark_mxfp4.py
patch_qwen3_thinkoff.py   reasoning-parser/template thinking agreement (in ggz14's loop); behaviour change, needs its own review
patch_rmsquant_fusion.py  rms_norm+fp8-quant fusion via radiance_rmsquant.py; this image fuses via sly/patch_fused_norm_quant.py
patch_step_trace.py       per-step worker timing trace (diagnostic); candidate tool for the kernel-tuning analysis
patch_topk_composite.py   composite small-k top-k/top-p (radiance_topk.py not copied); candidate for a follow-up A/B
patch_topk_triton_rows.py lower Triton top-k row threshold (in ggz14's loop); candidate for a follow-up A/B
patch_tp3_pad.py          TP=3 via dummy heads (radiance_tp3pad.py); TP=1 here
patch_verify_head.py      verify-head gate for radiance_verifyhead.py (not copied); tied to the paroquant drafter
