Blog

I enjoy taking notes on my research and sharing them with others who are also interested in these areas. Thanks to this habit, I have met many like-minded friends who share similar interests.

Large language model

  1. Notion

    Alignment Guidebook.

    A comprehensive introduction to the preference learning in LLMs.
  2. Zhihu

    Why should we do online RLHF/DPO? (in Chinese).

    A blog about the advantage of online exploration in RLHF.
  3. WeChat

    A tutorial of RLHF with LMFlow (Huggingface version).

Reinforcement learning theory

  1. PDF note

    An introduction to eluder coefficient, a unified framework for nearly all known tractable decision-making problems.

  2. PDF note

    Note on reduction-based RL A short note on eluder coefficient, decoupling coefficient, and decision-estimation coefficient.

  3. PDF note

    Note on non-linear contextual bandit An application of eluder coefficient in contextual bandit with general function approximation.

Optimization basis and Decentralized Optimization

  1. Zhihu

    A generic framework for first-order stochastic decentralized optimization

  2. Zhihu

    Analysis of batch gradient descent and stochastic gradient descent (in Chinese)

Multi-armed bandit basis and Multi-player multi-armed bandit

  1. Zhihu

    Bandits with stochastic rewards (in Chinese)

  2. Zhihu

    Lower bound for stochastic bandit (in Chinese)

  3. Zhihu

    Adversarial bandit (in Chinese)

  4. Zhihu

    UCB2 (in Chinese)

  5. Zhihu

    Introduction to multi-player MAB and Synchronisation Involves Communiation Algorithm

  6. Zhihu

    Heterogeneous multi-player MAB

  7. Zhihu

    Experiments of MAB Also check my research page for the GitHub repository