Egor Petrov

PhD Student in Computer Science at Columbia University

photo1_crop.jpg

DAP Lab, Columbia University

New York, NY

I work on efficient large-scale LLM training, pretraining and post-training, with a focus on parallelism, optimization, and scaling laws.

My research began under the supervision of Alexander Beznosikov, where I developed algorithms for variational inequalities and parameter-free optimization, then shifted to memory-efficient and zeroth-order (ZO) methods. These contributions led to the development of ZO-Library, a PyTorch-style open-source framework for ZO optimization in fine-tuning.

My bachelor thesis, supervised by Andrey Grabovoy, provides the first complete analytical Hessian for LayerNorm and feedforward sublayers, completing the second-order characterization of the full Transformer block.

At Yandex Research under Artem Babenko, and in collaboration with Samuel Horváth, I worked on asynchronous pipeline parallelism for large-scale LLM pretraining. I am now extending this line towards hyperparameter transfer and scaling laws for asynchronous pretraining.

Currently, I am a PhD student at Columbia University, in the DAP Lab, advised by Eugene Wu and Kostis Kaffes. I work on post-training efficiency: the optimization dynamics of RL post-training (RLVR), and staleness mitigation for asynchronous post-training with parallel rollout generation.

selected publications

  1. One-Step Gradient Delay is Not a Barrier for Large-Scale Asynchronous Pipeline Parallel LLM Pretraining
    Philip Zmushko, Egor Petrov, Nikita Abdullaev, and 2 more authors
    In International Conference on Machine Learning (ICML). Equal contribution: P. Zmushko, E. Petrov , 2026
  2. Sign-SGD is the Golden Gate between Multi-Node and Single-Node Learning: Significant Boost via Parameter-Free Optimization
    Daniil Medyakov, Sergey Stanko, Gleb Molodtsov, and 4 more authors
    In International Conference on Learning Representations (ICLR), 2026
  3. Leveraging Coordinate Momentum in SignSGD and Muon: Memory-Optimized Zero-Order LLM Fine-Tuning
    Egor Petrov, Grigoriy Evseev, Aleksey Antonov, and 4 more authors
    In ICML Workshop on Tiny Titans: The next wave of On-Device Learning for Foundation Models, 2025
  4. Closing the Curvature Gap: Full Transformer Hessians and Their Implications for Scaling Laws
    Egor Petrov, Nikita Kiselev, Vladislav Meshkov, and 1 more author
    . Under review at NeurIPS 2026 , 2025