Catalog
huggingface/trl-training

huggingface

trl-training

Post-train LLMs with TRL (Transformers Reinforcement Learning) — SFT, DPO, GRPO, KTO, and reward-model training. Use when writing or debugging training code with the TRL Python API or the trl CLI.

NewUpdated Oct 2, 2026

TRL

Each method pairs a *Trainer class with a *Config dataclass. Configs extend transformers.TrainingArguments, so all of its arguments work in any trainer config.

Trainer Dataset type
SFTTrainer language modeling or prompt-completion
DPOTrainer preference (chosen/rejected pairs)
GRPOTrainer prompt-only + reward function(s)
DistillationTrainer prompt-only + a teacher model (on-policy distillation)
KTOTrainer unpaired preference (per-sample bool label)
RewardTrainer preference (chosen/rejected pairs); trains a scalar reward model, not a policy

Many more trainers (OnlineDPO, ORPO, CPO, GKD, …) live in trl.experimental with unstable APIs: https://huggingface.co/docs/trl/experimental_overview

from datasets import load_dataset
from trl import SFTConfig, SFTTrainer

trainer = SFTTrainer(
    model="Qwen/Qwen2.5-0.5B",  # model ID or a PreTrainedModel instance
    args=SFTConfig(output_dir="Qwen2.5-0.5B-SFT"),
    train_dataset=load_dataset("trl-lib/Capybara", split="train"),
)
trainer.train()

Pass model as a string and route loading kwargs through model_init_kwargs (e.g. {"dtype": "bfloat16", "attn_implementation": "kernels-community/flash-attn2"}) instead of calling from_pretrained yourself. The tokenizer/processor is inferred from the model; pass processing_class only when it differs. For LoRA, pass peft_config=LoraConfig(...).

Dataset formats

Conversational: {"messages": [{"role": ..., "content": ...}]} (language modeling) or {"prompt": [...], "completion": [...]}. The chat template is applied automatically — never apply it yourself. Extra columns are allowed; GRPO forwards them to reward functions. Reference: https://huggingface.co/docs/trl/dataset_formats

SFT: the fields that matter

SFTConfig(
    max_length=1024,        # truncation length; None disables truncation
    packing=True,           # pack sequences into max_length blocks: fewer pad tokens, higher throughput
    padding_free=True,      # flatten batch, no padding; requires FlashAttention; implied by packing
    use_liger_kernel=True,  # fused Liger kernels, reduces peak memory
    assistant_only_loss=True,  # loss only on assistant turns (conversational datasets)
)

GRPO: online RL

def reward_len(completions, **kwargs):
    return [-abs(20 - len(c[0]["content"])) for c in completions]

trainer = GRPOTrainer(
    model="Qwen/Qwen2.5-0.5B-Instruct",
    reward_funcs=reward_len,  # or a list; rewards are summed
    args=GRPOConfig(output_dir="Qwen2.5-0.5B-GRPO", max_completion_length=512),
    train_dataset=load_dataset("trl-lib/DeepMath-103K", split="train"),
)

Reward functions are called with keyword arguments prompts, completions, completion_ids, trainer_state, plus every extra dataset column — accept **kwargs for the ones you ignore. Return list[float], one reward per completion. With conversational data, completions is a list of message lists, not strings.

The generation batch is per_device_train_batch_size × num_processes × steps_per_generation (or set generation_batch_size directly) and must be divisible by num_generations (default 8). Generation is the usual bottleneck — enable vLLM with use_vllm=True: vllm_mode="colocate" shares the training GPUs (size with vllm_gpu_memory_utilization); vllm_mode="server" uses a separate trl vllm-serve --model <model_id>.

AsyncGRPOTrainer (trl.experimental.async_grpo) implements the same algorithm with generation decoupled from training: a background worker streams completions from a vLLM server while the training loop consumes them, so the two overlap instead of alternating.

CLI

Flags mirror the config fields: trl sft --model_name_or_path Qwen/Qwen2.5-0.5B --dataset_name trl-lib/Capybara. YAML via --config; distributed presets via --accelerate_config zero3 (Python scripts: accelerate launch train.py).

Files1
1 files · 11.1 KB

Select a file to preview

Overall Score

78/100

Grade

B

Good

Grades are signals, not a certification. Always review a skill yourself before use.

Safety

88

Quality

76

Clarity

82

Completeness

68

Summary

TRL (Transformers Reinforcement Learning) training skill provides guidance for post-training LLMs using the TRL library's trainer classes (SFT, DPO, GRPO, KTO, RewardTrainer) and configs. The skill documents dataset formats, trainer-specific parameters, and provides working code examples for fine-tuning models from Hugging Face via Python API or CLI.

Detected Capabilities

code generation and examplesapi reference documentationdataset format guidancemodel loading and configtraining parameter optimizationdistributed training setup

Trigger Keywords

Phrases that agents use to match this skill to user intent.

trl trainingllm fine-tuning dporeward model traininggrpo online rlinstruction tuning sftreinforcement learning trainer

Referenced Domains

External domains referenced in skill content, detected by static analysis.

huggingface.cowww.apache.org

Use Cases

  • instruction fine-tune LLMs with SFT
  • train preference-based models with DPO
  • implement online RL training with GRPO
  • build custom reward models
  • configure distributed training for TRL
  • debug TRL trainer configs and datasets

Quality Notes

  • Clear tabular reference of trainer types and dataset requirements
  • Concrete code examples for SFT, GRPO, and reward function patterns
  • Well-organized sections covering dataset formats, critical config fields, and CLI usage
  • Covers online RL (GRPO) with generation optimization advice (vLLM, AsyncGRPOTrainer)
  • Links to authoritative HF documentation for deeper reference
  • Parameter guidance is practical (packing, padding_free, liger kernels, vLLM modes)
  • Missing: error handling patterns (OOM, CUDA issues, convergence debugging)
  • Missing: limitations and failure modes (e.g., dataset size requirements, memory constraints)
  • Missing: integration with evaluation or inference workflows post-training
  • Reward function design is underexplained — only one simple length-based example
Model: claude-haiku-4-5-20251001Analyzed: Oct 2, 2026

Reviews

Add this skill to your library to leave a review.

No reviews yet

Be the first to share your experience.

Version History

  1. v2.0

    Contract changed: description

    ✦ AIDescription narrows activation scope from CLI commands to writing and debugging TRL Python API code; SKILL.md substantially reduced by ~237 net lines.

    triggering2026-10-02

    LATEST
  2. v1.1

    Content updated

    ✦ AINo observable behavioral changes in SKILL.md; safety grade improved B to A.

    2026-09-09

    View This Version
  3. v1.0

    2026-07-12

    View This VersionInitial version

Use huggingface/trl-training in your dev environment

Command Palette

Search for a command to run...