rloo

Here is 1 public repository matching this topic...

sinanuozdemir / oreilly-llm-rl-alignment

This training offers an intensive exploration into the frontier of reinforcement learning techniques with large language models (LLMs). We will explore advanced topics such as Reinforcement Learning with Human Feedback (RLHF), Reinforcement Learning from AI Feedback (RLAIF), Reasoning LLMs, and demonstrate practical applications such as fine-tuning

reinforcement-learning ai llama agents ppo dpo qwen deepseek reward-modeling grpo rloo

Updated Mar 9, 2026
Jupyter Notebook

Improve this page

Add a description, image, and links to the rloo topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the rloo topic, visit your repo's landing page and select "manage topics."

Learn more

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

rloo

Here is 1 public repository matching this topic...

sinanuozdemir / oreilly-llm-rl-alignment

Improve this page

Add this topic to your repo