ML research & debugging notes
Achyuthan Sivasankar
Machine learning · systems · open source
M.S. Computer Science at NYU. I write about ML research, debugging stories, and the things I'm still figuring out as I go.
All posts
When Memory Layout Assumptions Collide: Fixing a KV Cache Bug in vLLMTwo independently correct components — KV offloading and per-token-head quantization — corrupted generations the moment they were combined. The fix was four lines. Finding where those four lines belonged took a trip through GPU memory layouts.
We Found the Moment Grokking Actually BeginsGrokking looks like a sudden leap. Internally, the algorithm forms long before validation accuracy moves — and that hidden gap is a regularization phase we can control.
When Depth Helps, Hurts, or Doesn't Matter: A Taxonomy of Adaptive Compute in World ModelsAcross nine DeepMind Control tasks, depth helps rollouts on 6/9, barely matters on 1/9, and hurts on 2/9 — plus a training catch-22 that can erase the depth advantage routing needs.
When Resuming Training Made My Model Learn the Wrong Data: Debugging an Epoch-Aware Sampler Bug in Megatron BridgeA deep dive into a subtle checkpoint resume bug in NVIDIA Megatron Bridge that silently caused fine-tuning jobs to repeatedly train on the wrong portion of the dataset.
When Docker Cache Lies: Debugging a Nightly Build Failure in vLLMHow a BuildKit bind-mount cache quirk broke vLLM nightly Docker images with ImportError: AnthropicOutputConfig — and how a SHA-256 checksum fixed it.
I also contribute to open source — see projects.