Blog
I enjoy taking notes on my research and sharing them with others who are also interested in these areas. Thanks to this habit, I have met many like-minded friends who share similar interests.
Large language model
- Notion A comprehensive introduction to the preference learning in LLMs.
- Zhihu A blog about the advantage of online exploration in RLHF.
Reinforcement learning theory
- PDF note
-
PDF note
Note on reduction-based RL A short note on eluder coefficient, decoupling coefficient, and decision-estimation coefficient.
-
PDF note
Note on non-linear contextual bandit An application of eluder coefficient in contextual bandit with general function approximation.
Optimization basis and Decentralized Optimization
-
Zhihu
A generic framework for first-order stochastic decentralized optimization
-
Zhihu
Analysis of batch gradient descent and stochastic gradient descent (in Chinese)
Multi-armed bandit basis and Multi-player multi-armed bandit
- Zhihu
- Zhihu
- Zhihu
- Zhihu
-
Zhihu
Introduction to multi-player MAB and Synchronisation Involves Communiation Algorithm
- Zhihu
-
Zhihu
Experiments of MAB Also check my research page for the GitHub repository