Egor Petrov
PhD Student in Computer Science at Columbia University
DAP Lab, Columbia University
New York, NY
I work on efficient large-scale LLM training, pretraining and post-training, with a focus on parallelism, optimization, and scaling laws.
My research began under the supervision of Alexander Beznosikov, where I developed algorithms for variational inequalities and parameter-free optimization, then shifted to memory-efficient and zeroth-order (ZO) methods. These contributions led to the development of ZO-Library, a PyTorch-style open-source framework for ZO optimization in fine-tuning.
My bachelor thesis, supervised by Andrey Grabovoy, provides the first complete analytical Hessian for LayerNorm and feedforward sublayers, completing the second-order characterization of the full Transformer block.
At Yandex Research under Artem Babenko, and in collaboration with Samuel Horváth, I worked on asynchronous pipeline parallelism for large-scale LLM pretraining. I am now extending this line towards hyperparameter transfer and scaling laws for asynchronous pretraining.
Currently, I am a PhD student at Columbia University, in the DAP Lab, advised by Eugene Wu and Kostis Kaffes. I work on post-training efficiency: the optimization dynamics of RL post-training (RLVR), and staleness mitigation for asynchronous post-training with parallel rollout generation.