Reinforcement learning

How it works

The computer tries an action and gets a score. A good result is a reward. A bad result is a penalty. After thousands of tries, it finds the best way to play, like a video game it practises over and over.

A famous example

In 2016, a program called AlphaGo beat a top human player at the board game Go. It got there partly by playing against itself many times and learning from wins and losses.

Why chatbots use it

People rate chatbot answers, and the chatbot gets a 'reward' for answers people like. This is called RLHF. It helps chatbots be more helpful and polite.

Key takeaways

  • Try, get a reward, improve.
  • Great when we can score success but do not know the best move.
  • Chatbots use it to give answers people prefer.

Quick questions

What is RLHF?

It stands for reinforcement learning from human feedback. People rate answers and the model learns to give the kind they like.

What is a reward?

Just a score that tells the computer how well it did.

Why is it hard?

The computer can find shortcuts that win points without doing what we really wanted.

Keep learning