Trlx
Visit Tooltrlx is an Open Source & Models tool that provides a distributed training framework for fine-tuning large language models. It supports reinforcement learning via a provided reward function or a reward-labeled dataset.
trlx is an Open Source & Models tool that provides a distributed training framework for fine-tuning large language models. It supports reinforcement learning via a provided reward function or a reward-labeled dataset.
About
trlx is a distributed training framework specifically designed for fine-tuning large language models using Reinforcement Learning via Human Feedback (RLHF). It supports training with either a provided reward function or a reward-labeled dataset. The framework offers compatibility with Hugging Face models, enabling fine-tuning of causal and T5-based language models up to 20B parameters, such as facebook/opt-6.7b and EleutherAI/gpt-neox-20b. For models exceeding 20B parameters, trlx integrates with NVIDIA NeMo-backed trainers, leveraging efficient parallelism techniques for scalability. It currently implements Proximal Policy Optimization (PPO) and Implicit Language Q-Learning (ILQL) algorithms, with support for both Accelerate and NeMo trainers.
Capabilities
Pricing & Plans
Open Source
Free
FAQs