Catalog
huggingface/train-sentence-transformers

huggingface

train-sentence-transformers

Train or fine-tune sentence-transformers models across `SentenceTransformer` (bi-encoder, dense or static embedding model for retrieval, similarity, clustering, classification, paraphrase mining, dedup, multimodal), `CrossEncoder` (reranker, pair scoring for two-stage retrieval / pair classification), `SparseEncoder` (SPLADE, sparse embedding model for learned-sparse retrieval), and `MultiVectorEncoder` (ColBERT / late-interaction, per-token embeddings scored with MaxSim). Covers loss selection, hard-negative mining, evaluators, distillation, LoRA, Matryoshka, and Hugging Face Hub publishing. Use for any sentence-transformers training task.

NewUpdated Sep 9, 2026

Train a sentence-transformers Model

This SKILL.md is a router, not a manual. It tells you which references and example scripts to load for your task. The actual content (recommended losses, evaluators, training-script structure, model selection, training-arg knobs, troubleshooting) lives in references/ and scripts/.

Do not synthesize a training script from this file alone. Open the per-type production template (scripts/train_<type>_example.py) and copy it as your starting point. The templates contain load-bearing scaffolding (autocast helper, model-card class, logger silencing list, force=True, seed, TF32, version-compatible imports, named-evaluator metric handling) that prior agent runs have repeatedly missed when rolling their own from a synthesized snippet.

1. Identify the model type

Tag Class What it does When to pick
[SentenceTransformer] SentenceTransformer (bi-encoder) Maps each input to a fixed-dim dense vector Retrieval, similarity, clustering, classification, paraphrase mining, dedup
[CrossEncoder] CrossEncoder (reranker) Scores (query, passage) pairs jointly Two-stage retrieval (rerank top-100 from bi-encoder), pair classification
[SparseEncoder] SparseEncoder (SPLADE) Sparse vectors over the vocabulary Learned-sparse retrieval, inverted-index backends (Elasticsearch / OpenSearch / Lucene)
[MultiVectorEncoder] MultiVectorEncoder (ColBERT) One embedding per token, scored with MaxSim Late-interaction retrieval, recall gains over bi-encoders at higher storage cost, multimodal (ColPali / ColQwen2)

Tiebreakers when the request is ambiguous: "embedding model" / "vector search" / "similarity" → [SentenceTransformer]. "rerank" / "ranker" / "two-stage" → [CrossEncoder]. "SPLADE" / "sparse" / "inverted index" → [SparseEncoder]. "ColBERT" / "late interaction" / "multi-vector" / "MaxSim" / "ColPali" / "ColQwen" → [MultiVectorEncoder]. If still unclear, ask.

2. Required reading

Read these in full before writing any code. Do not triage by perceived relevance.

Per-type: always required

[SentenceTransformer]

  • references/losses_sentence_transformer.md: loss-to-data-shape mapping, BatchSamplers.NO_DUPLICATES requirement for MNRL-family, Cached* ↔ gradient_checkpointing incompatibility.
  • references/evaluators_sentence_transformer.md: evaluator-to-task mapping, metric_for_best_model key construction (named vs unnamed), per-evaluator primary_metric values.
  • references/model_architectures.md: encoder vs decoder vs static vs Router pipelines, pooling rules (mean / cls / lasttoken), auto-mean-pooling behavior for fresh-start MLM bases.
  • scripts/train_sentence_transformer_example.py: production template. Copy this as your starting point.

[CrossEncoder]

  • references/losses_cross_encoder.md: pointwise / pairwise / listwise / distillation, pos_weight derivation, activation_fn=Identity() mandatory for non-BCE losses (silent eval-rank collapse otherwise).
  • references/evaluators_cross_encoder.md: CrossEncoderRerankingEvaluator recipe, named-evaluator key format eval_{name}_{primary_metric}.
  • scripts/train_cross_encoder_example.py: production template. Copy this as your starting point.

[SparseEncoder]

  • references/losses_sparse_encoder.md: SpladeLoss wrapper requirement, FLOPS regularizer weights, smoke-test active-dim ramp behavior.
  • references/evaluators_sparse_encoder.md: SparseNanoBEIREvaluator (English-only) and the in-domain alternative, eval_{name}_{primary_metric} key format.
  • scripts/train_sparse_encoder_example.py: production template. Copy this as your starting point.

[MultiVectorEncoder]

  • references/losses_multi_vector_encoder.md: MaxSim scoring, scale choice per scoring mode (scale=1.0 for MaxSim, roughly the average query length for MeanMaxSim), MNRL / CachedMNRL / MarginMSE / DistillKLDiv, XTR-vs-ColBERT scoring, CachedMNRL ↔ gradient_checkpointing incompatibility.
  • references/evaluators_multi_vector_encoder.md: MultiVectorNanoBEIREvaluator (English-only) and the in-domain alternative, eval_NanoBEIR_mean_maxsim_ndcg@10 key format, distillation-eval spearman variant.
  • scripts/train_multi_vector_encoder_example.py: production template. Copy this as your starting point.

Cross-cutting: always required (regardless of task)

  • references/training_args.md: TrainingArguments knobs, precision rules (load fp32 + autocast bf16/fp16, never torch_dtype=bfloat16), warmup_steps (float) vs deprecated warmup_ratio, save_steps must be a multiple of eval_steps for load_best_model_at_end, schedulers, HPO, tracker, resume, hub-push variants.
  • references/dataset_formats.md: column-matching rules (label name auto-detection, column-order-not-name), reshaping recipes, hard-negative mining options.
  • references/base_model_selection.md: discovery commands, per-type model namespaces, ModernBERT-family max_seq_length=8192 trap, datasets >= 4 script-loader rejection, non-English starting-point shortcuts.
  • references/troubleshooting.md: symptom-indexed failure recipes. Skim the section headings on every run, even a healthy one. The "Metrics don't improve" and "Hub push fails" entries cover bugs that bite frequently and are cheaper to recognize before they fire than to debug after.

Cross-cutting: load when applicable

  • references/hardware_guide.md: VRAM sizing, multi-GPU, FSDP / DeepSpeed, HF Jobs flavors. Required for >24GB models, multi-GPU, or HF Jobs runs.
  • references/hf_jobs_execution.md: required when running on HF Jobs.
  • references/prompts_and_instructions.md: required when using prompt-tuned bases (E5, BGE, GTE, Qwen3-Embedding, Instructor, Nomic, etc.) or adding query: / passage: style prefixes.

Variant scripts (open when the task matches)

  • [SentenceTransformer] scripts/train_sentence_transformer_<matryoshka|multi_dataset|with_lora|distillation|make_multilingual|static_embedding>_example.py.
  • [CrossEncoder] scripts/train_cross_encoder_<distillation|listwise>_example.py.
  • [SparseEncoder] scripts/train_sparse_encoder_distillation_example.py.
  • Hard-negative mining CLI: scripts/mine_hard_negatives.py.

3. Defaults

Override only if the user specifies otherwise:

  • Local execution. Pitch HF Jobs only if local hardware can't fit the job.
  • Single run. After it completes, propose experimentation if the user would benefit (weak/marginal verdict, "see how high you can push it" framing, etc.). Iteration rules in references/training_args.md (Experimentation section).
  • Public Hub push at end-of-run, wrapped in try-except. On HF Jobs (ephemeral env) ALSO enable in-trainer push (push_to_hub=True + hub_strategy="every_save"). Details in references/hf_jobs_execution.md.

4. Constraints the produced script must satisfy

These are non-negotiable contracts. Implementation lives in the production templates and references. Do not reinvent.

  • Capture the pre-training evaluator score as baseline_eval before trainer.train().
  • Emit a single end-of-run line: VERDICT: WIN|MARGINAL|REGRESSION | score=... | baseline=... | delta=.... A monitor scrapes for this.
  • Silence httpx, httpcore, huggingface_hub, urllib3, filelock, fsspec to WARNING (otherwise HF download URLs flood the agent's context).
  • Tee logs to logs/{RUN_NAME}.log.
  • End with model.push_to_hub(...) wrapped in try/except.
  • Smoke-test before any long run (max_steps=1 + tiny dataset slice). The production templates show one common pattern (SMOKE_TEST env var).
  • [CrossEncoder] Include EarlyStoppingCallback(patience>=3). CE rerankers often peak mid-training and regress.
  • [SparseEncoder] Log query_active_dims / corpus_active_dims on the verdict line. High nDCG with collapsed sparsity is not a win. The keys come back name-prefixed (e.g. ..._query_active_dims). Use suffix matching to pluck them. See the SPARSE production template for the exact pattern.
  • [MultiVectorEncoder] Match scale to the scoring mode on any MNRL-family loss: near 1.0 for unnormalized MaxSim (do not copy scale=20.0 from bi-encoder MNRL), roughly the average query length with length-normalized MeanMaxSim, since each score is divided by its query's token count. XTRScores is a train-only similarity_fct: the evaluators reject it, so evaluation always scores with MaxSim, including for XTR-trained models.

5. Workflow

  1. Identify the model type (§1). Ask if ambiguous.
  2. Load the §2 required-reading files for that type.
  3. Open scripts/train_<type>_example.py and copy it as your starting point.
  4. Replace MODEL_NAME, DATASET_NAME, RUN_NAME, the loss, and the evaluator with the user's task. Cross-check loss/data-shape match against references/losses_<type>.md. Cross-check the metric_for_best_model key against references/evaluators_<type>.md (named evaluators format the key as eval_{name}_{primary_metric}).
  5. Smoke-test (max_steps=1).
  6. Run.
  7. After the run, append to logs/experiments.md and propose iteration if the verdict is weak/marginal.

Prerequisites

pip install "sentence-transformers[train]>=5.0"        # add [train,image] / [audio] / [video] for [SentenceTransformer] multimodal
                                                       # [MultiVectorEncoder] requires >=6.0
pip install trackio                                    # optional tracker (or wandb / tensorboard / mlflow)
hf auth login                                          # or set HF_TOKEN with write scope (for Hub push)

GPU strongly recommended. CPU works only for demos and [SentenceTransformer] StaticEmbedding.

Files31
31 files · 271.3 KB

Select a file to preview

Grade adjusted by static analysis guardrails

AI scored this skill as grade A, but static analysis findings capped it to B:

  • • Pipe-to-shell pattern (curl/wget piped to sh/bash) (max: B)

Overall Score

87/100

Grade

B

Good

Grades are signals, not a certification. Always review a skill yourself before use.

Safety

88

Quality

90

Clarity

84

Completeness

82

Summary

This is a comprehensive routing skill for sentence-transformers model training across four architecture types (bi-encoder, cross-encoder, sparse-encoder, multi-vector). It provides a decision tree, required-reading references, production templates, and troubleshooting guidance. The skill directs users to per-type training scripts and reference documents, with clear instructions to copy and customize templates rather than synthesize from scratch. All supporting files are present and complete.

Static Analysis Findings

3 findings

Patterns detected by deterministic static analysis before AI scoring. Hover over any finding code for detailed information and remediation guidance.

Remote Code Execution
SEC-031Script Download

Dynamic script download for execution

references/hf_jobs_execution.mdcurl -LsSf https://hf.co/cli/install.sh
Command Injection
SEC-011Dynamic Shell Eval4x in 2 files

Shell eval/exec of dynamic content

references/evaluators_multi_vector_encoder.mdeval"3x
scripts/train_cross_encoder_listwise_example.pyeval"
SEC-010Pipe-to-ShellMax: B

Pipe-to-shell pattern (curl/wget piped to sh/bash)

references/hf_jobs_execution.mdcurl -LsSf https://hf.co/cli/install.sh | bash

Detected Capabilities

file read (references and scripts)Python script execution (training via trainer classes)model loading from Hugging Face Hubdataset loading and preprocessingGPU/multi-GPU training via acceleratemodel checkpointing and Hub pushevaluator instantiation and metric trackingshell command construction (HF CLI patterns)environment variable handling

Trigger Keywords

Phrases that agents use to match this skill to user intent.

train sentence transformersfine-tune embedding modeltrain reranker cross encodertrain sparse encoder spladecolbert late interactiondistill embedding modelmine hard negativesmulti-gpu traininghugging face jobs trainingmatryoshka training

Risk Signals

INFO

SEC-011 command-injection: Shell eval/exec of dynamic content (eval") found in references/evaluators_multi_vector_encoder.md

references/evaluators_multi_vector_encoder.md
WARNING

SEC-010 command-injection: Pipe-to-shell pattern detected (curl ... | bash)

references/hf_jobs_execution.md | Match: curl -LsSf https://hf.co/cli/install.sh | bash
WARNING

SEC-031 remote-code-execution: Dynamic script download for execution (curl -LsSf https://hf.co/cli/install.sh)

references/hf_jobs_execution.md
INFO

SEC-011 command-injection: eval" pattern detected in scripts/train_cross_encoder_listwise_example.py

scripts/train_cross_encoder_listwise_example.py

Referenced Domains

External domains referenced in skill content, detected by static analysis.

hf.cohuggingface.cosbert.netwww.apache.org

Use Cases

  • Train a bi-encoder (SentenceTransformer) on retrieval pairs
  • Train a cross-encoder (reranker) on labeled pair data
  • Train a sparse encoder (SPLADE) on contrastive triplets
  • Train a multi-vector encoder (ColBERT) with MaxSim scoring
  • Distill embeddings from a teacher model
  • Mine hard negatives for contrastive training
  • Set up multi-GPU or HF Jobs training
  • Configure Matryoshka / LoRA / multilingual training variants
  • Debug training failures (NaN loss, metrics not improving, eval hangs)
  • Choose base models and loss functions per task

Quality Notes

  • Strength: skill is a well-structured router with clear decision tree (§1 model-type identification). Users are explicitly warned not to synthesize from SKILL.md alone but to copy production templates.
  • Strength: all production templates are complete, runnable, and include smoke-test mode, error handling, model-card generation, logging to file, TensorBoard/Trackio integration, and Hub push with try-except.
  • Strength: per-type reference docs (losses, evaluators, training-args, architecture, hardware, dataset formats) are comprehensive and cross-referenced. Decision tables map user intent to implementation details.
  • Strength: troubleshooting reference is symptom-indexed and covers 20+ failure modes with root causes and fixes (NaN loss, OOM, eval hangs, activation-fn mismatches, sparsity issues).
  • Strength: gotchas are documented extensively (e.g., activation_fn=Identity() mandatory for non-BCE cross-encoder losses, CachedMNRL incompatible with gradient_checkpointing, ModernBERT 8192-token default breaks memory).
  • Strength: multi-GPU, LoRA, distillation, Matryoshka, multilingual, and hard-negative-mining variants all have dedicated scripts or coverage in references.
  • Strength: HF Jobs execution guide includes timeout sizing, secrets management, dataset caching, and monitoring via CLI.
  • Weakness: SEC-010 (curl | bash) appears in hf_jobs_execution.md as an installation command for the `hf` CLI. Context is clear (official HF tool installation), but the pattern is flagged by static analysis. The pattern is appropriate for its purpose (first-time tool setup outside the skill), but could be mitigated by suggesting `pip install hf-cli` as an alternative.
  • Weakness: SEC-011 (eval) patterns in references appear to be false positives from markdown formatting (docstring examples mentioning `eval(...)` Python semantics in text, not actual shell eval). The analyzer flags 'eval"' as a substring match. No actual command injection risk.
  • Weakness: Skill is very large (~70KB of references + 11 templates). No index or decision flowchart in the SKILL.md intro. A visual tree (Model type -> Loss -> Evaluator -> Script) would reduce cognitive load for new users.
  • Weakness: Pre-computed teacher embeddings for distillation (MSELoss) require an offline pass. The referenced distillation scripts don't explain how to populate the `label` column from scratch; they assume it's already cached or provide a snippet (e.g., `train_sentence_transformer_distillation_example.py`). Implied knowledge: users must run the teacher, save outputs, and populate the dataset before the trainer sees it.
  • Strength: All file references in SKILL.md are present and verified in the manifest. No dangling links.
Model: claude-haiku-4-5-20251001Analyzed: Sep 9, 2026

Reviews

Add this skill to your library to leave a review.

No reviews yet

Be the first to share your experience.

Version History

  1. v2.0

    Contract changed: description

    ✦ AIAdds support for MultiVectorEncoder (ColBERT / late-interaction) model type with dedicated references and executable training script.

    triggeringnew script2026-09-09

    LATEST
  2. v1.0

    2026-07-11

    View This VersionInitial version

Use huggingface/train-sentence-transformers in your dev environment

Command Palette

Search for a command to run...