Reinforcement Learning from Human Feedback

3.565 students enrolled
Large language models (LLMs) are trained on human-generated text, but additional methods are needed to align an LLM with human values and preferences. Reinforcement Learning from Human Feedback (RLHF) is currently the main method for aligning LLMs with human values and preferences. RLHF is also used for further tuning a base LLM to align with values and preferences that are specific to your use case. In this course, you will gain a conceptual understanding of the RLHF training process, and then practice applying RLHF to tune an LLM. You will: 1. Explore the two datasets that are used in RLHF training: the “preference” and “prompt” datasets. 2. Use the open source Google Cloud Pipeline Components Library, to fine-tune the Llama 2 model with RLHF. 3. Assess the tuned LLM against the original base model by comparing loss curves and using the “Side-by-Side (SxS)” method.
CERTIFICATEKatılım Sertifikası
FORMAT100% Online
DURATIONSelf-paced

Details

  • ProviderDeepLearning.AI
  • TypeCourse
  • CategorySoftware & Programming
  • LanguageEnglish

Öğrenenlerimiz ne diyor?

Türkiye'nin yüz akı üniversitelerince hazırlanan; akademik doyuruculuğa sahip eğitim içeriklerinin, etkileşimli videolarla bir araya getirildiği bir üniversiteden eğitim almak istiyorsanız doğru yerdesiniz.
Ramazan Bölükbaşı
Çok yoğun programı olan öğrenciler için büyük bir fırsat. Bir şeylerin gelişmesi değişmesi için çabalamalıyız.
Melis Gülsar
Hızlı desteği ve üst düzey hizmeti ile Campus Online ve Sosyal Medya Sertifika Programı hizmeti sağlayan Adnan Menderes Üniversitesine sonsuz teşekkürlerimi sunuyorum.
Nazif Bayram
$49
CampusOnline Assistant
courses saved

Course Comparison

Institution
Rating
Turkish Subtitles
Level
Duration
Price
Certificate