I am an AI researcher. I have worked on optimal control, bandits, and reinforcement learning. These days, I keep busy by working with multi-agent reinforcement learning and VLM post-training.

I learn better when writing things down. I archive the result here—hopefully, it can be helpful to others.

The posts are AI-free. LLMs help me find resources and explore the literature, but in the end every word, inaccuracies or blatant mistake are my own.

For more information or contact, you can check out my personal page here.