Mergedvllm-project / vllm#49226

Disable cross-layer KV blocks for per-token-head quant

Fixed silent KV-cache corruption when OffloadingConnector was combined with a per-token-head quantized KV cache.

Per-token-head quant stores an inline fp32 scale in each cell’s padded tail, and the attention backend carves scale views assuming a per-layer contiguous buffer — but the offload path allocates one cross-layer interleaved buffer shared by all layers, so the scales aliased neighbouring K/V data from the first decode. After review the guard moved out of the connector into use_uniform_kv_cache(), covering every connector in four lines. Full write-up on the blog.

  • vLLM
  • KV cache
  • quantization
  • GPU memory
MergedNVIDIA / Megatron-LM#5743

Support HSDP deferred DP-outer gradient reduction

Added deferred DP-outer gradient reduction under hybrid-sharded data parallel in experimental Megatron-FSDP.

Lets the outer data-parallel gradient reduction be deferred when running hybrid-sharded data parallel (HSDP) in experimental Megatron-FSDP.

  • Megatron-FSDP
  • distributed training
  • HSDP
MergedNVIDIA-NeMo / Automodel#2998

Single tie_word_embeddings guard via per-class TieSupport

Replaced scattered weight-tying checks with one per-class TieSupport declaration (BOTH / TIED_ONLY / UNTIED_ONLY).

Consolidated the per-family tie_word_embeddings guards behind a single declaration each model class makes about what it supports, plus a from_pretrained check that catches the config being flipped after load.

  • NeMo Automodel
  • weight tying
  • refactor
MergedNVIDIA-NeMo / Megatron-Bridge#4601

Make finetuning batch sampler epoch-aware on checkpoint resume

Fixed a resume bug where the sampler replayed the dataset tail every epoch, so fine-tuning silently trained on an incomplete view of the data.

The restored consumed_samples offset was applied to every subsequent epoch instead of only the interrupted one, so the start of the dataset was repeatedly skipped. Making the sampler epoch-aware restored correct, deterministic data coverage, with per-epoch reshuffling added on top. Full write-up on the blog.

  • Megatron Bridge
  • checkpointing
  • distributed training
Mergedvllm-project / vllm#47379

Recover raw tail when the GPT-OSS Harmony parser ends non-terminal

Co-authored follow-up to #47062, extending non-terminal Harmony recovery to the Responses API. This is the one contribution Achyuthan co-authored rather than authored alone.

Co-authored with @yzong-rh. Builds on my earlier fix (#47062) so that generations ending mid-state in the Harmony parser also recover their raw tail through the Responses API, rather than returning nothing.

  • vLLM
  • gpt-oss
  • frontend
Mergedvllm-project / vllm#47062

Return raw output when the GPT-OSS Harmony parser ends non-terminal

Stopped the Harmony parser from discarding generated text when it finished in a non-terminal state.

When the Harmony parser reached the end of a generation while still mid-state, the partial output was dropped entirely. Returning the raw text instead means callers see what the model actually produced.

  • vLLM
  • gpt-oss
  • frontend
Mergeddeepspeedai / DeepSpeed#8078

Avoid CUDA context initialization during import-time op checks

Made importing DeepSpeed fork-safe by not creating a CUDA context while probing op compatibility.

Op compatibility checks run at import time were initializing a CUDA context, which breaks any process that forks afterwards. Deferring that work keeps import deepspeed safe in multiprocessing and dataloader workers.

  • DeepSpeed
  • CUDA
  • distributed training
MergedNVIDIA-NeMo / Automodel#2709

Cherry-pick #2601 into r0.5.0

Backported the Gemma4 MoE re-tie fix (#2601) into the r0.5.0 release branch.

Cherry-pick of #2601 into the r0.5.0 release branch.

  • NeMo Automodel
  • release
  • backport
MergedNVIDIA-NeMo / Automodel#2601

Re-tie lm_head to active embed_tokens on Gemma4 MoE path

Fixed Gemma4 MoE weight tying where lm_head drifted from the active embed_tokens during training.

Merged PR fixing weight tying on the Gemma4 MoE path: the language-model replacement orphaned the lm_head ↔ embed_tokens link, so the tied weights drifted during training. The fix re-ties lm_head to the active embed_tokens.

  • NeMo Automodel
  • Gemma
  • MoE
  • weight-tying
Mergedvllm-project / vllm#44795

Fix nightly Docker cache invalidation for vLLM wheels

Resolved ImportError: AnthropicOutputConfig in nightly images by busting BuildKit layer cache with wheel SHA-256 checksums.

Merged PR fixing stale wheel installs in nightly Docker images. Bind-mount installs were skipped on warm BuildKit agents; copying only wheel.sha256 busts the cache without bloating the image. Added regression test for Anthropic protocol exports.

  • vLLM
  • Docker
  • BuildKit
  • CI
Mergedscikit-learn / scikit-learn#32347

Clarify fit_transform() vs fit().transform() for TargetEncoder

Documented why TargetEncoder.fit_transform() uses internal cross-fitting and is not equivalent to fit().transform().

TargetEncoder.fit_transform() deliberately cross-fits to avoid target leakage, so it does not return the same values as calling fit() then transform(). The docs now state which to use when.

  • scikit-learn
  • documentation