Updated 30 September 2026: this tutorial is rewritten and every output on it now comes from a real run. The earlier version trained on STS-B scores in the 0-5 range, but the loss it used only accepts targets between 0 and 1. That is why its "expected" training log showed negative loss values and a correlation of 0.55. The dataset it loaded also no longer resolves on a clean install. The new version uses the current CrossEncoderTrainer API and a dataset whose scores are already in the right range.
1. What you'll build: a cross-encoder that scores and reranks sentence pairs
$ python rerank.py0.602 Click 'Forgot password' on the sign-in page to get a reset link.0.541 You can change your password under Settings > Security.0.353 Passwords must be at least 12 characters long.0.083 Our office is closed on public holidays.That is a distilroberta-base model you fine-tune yourself on the STS Benchmark (STS-B), ranking four answers to a support question. Used this way, to reorder search results, a cross-encoder is usually called a reranker. Before training, the same model scores the benchmark's validation pairs at a Spearman correlation of 0.0156: its ordering of the pairs is no better than chance. After about 43 minutes of training on a laptop CPU, it reaches 0.8776, and it puts the two answers that actually explain a reset at the top.
By the end you will have a cross-encoder fine-tuned on STS-B with the Sentence Transformers trainer, a measured correlation that shows it learned, and two small scripts that use it to score sentence pairs and rerank candidates. This is an intermediate tutorial: you should know Python and have seen PyTorch or Hugging Face models before. The cross-encoder parts are explained in full. Budget about an hour on a CPU, most of it waiting for training.
Verified against Python 3.13.9, torch==2.14.0+cpu, sentence-transformers==6.1.0, transformers==5.17.0, datasets==5.0.1 and accelerate==1.15.0 on 2026-09-30, on a 4-thread Windows 11 CPU. Every step was run end to end, including the 43-minute training run, twice on that machine; both runs produced the same table and a byte-identical model. The macOS and Linux activation command was not run, because the test machine runs Windows.
2. Prerequisites: Python 3.10+, Sentence Transformers 6.1, and a CPU
You need:
- Python 3.10 or later. Sentence Transformers 6 requires it.
- About 4 GB of free memory and 2 GB of disk for PyTorch, the base model and the dataset.
- No GPU. Every step here was run on a 4-thread laptop CPU. A GPU cuts training from about 43 minutes to a few minutes, and the code does not change.
- No accounts and no API keys. The model and dataset download anonymously from the Hugging Face Hub.
Create a project folder and a virtual environment:
mkdir cross-encoder-stsbcd cross-encoder-stsbpython -m venv .venvActivate it. In Windows PowerShell (if it says running scripts is disabled, run Set-ExecutionPolicy -Scope CurrentUser RemoteSigned once, or use .venv\Scripts\activate.bat from cmd instead):
.venv\Scripts\Activate.ps1On macOS and Linux:
source .venv/bin/activateInstall PyTorch first. The CPU build is a smaller download. If you have an NVIDIA GPU and want the CUDA build, leave out the --index-url part. On a Mac, leave it out as well; the standard macOS wheel is the only one. On Apple silicon, the trainer uses the Mac's GPU automatically when PyTorch detects it, but this tutorial did not measure how long training takes there:
pip install torch==2.14.0 --index-url https://download.pytorch.org/whl/cpuThen create requirements.txt in the project folder:
File: requirements.txt
sentence-transformers[train]==6.1.0transformers==5.17.0datasets==5.0.1accelerate==1.15.0The [train] extra matters. The trainer needs datasets and accelerate, and without them it refuses to start. Install:
pip install -r requirements.txtCheck the install. The command imports accelerate too, so it fails if the [train] extra was skipped:
python -c "import accelerate, sentence_transformers as st; print(st.__version__)"6.1.03. How a cross-encoder scores a pair, and when you need one
Most tutorials will tell you to use bi-encoders for semantic search. They're fast, they scale, they work. And for 80% of use cases, that's the right answer.
But bi-encoders make a fundamental tradeoff. They encode each sentence independently, which means they miss the subtle interactions between words across sentences. When you need to catch those interactions - when accuracy matters more than millisecond latency - you need cross-encoders.
Think about how you judge if two sentences mean the same thing. You don't read sentence A, form an opinion, then read sentence B and form another opinion, then compare your two opinions. You read them together, constantly cross-referencing: "Oh, 'man' here corresponds to 'person' there. 'Playing' matches 'performing'. 'Guitar' and 'instrument' - close enough in this context."
That's what cross-encoders do. They take both sentences as a single input, <s> sentence A </s></s> sentence B </s> for a RoBERTa model, and let the transformer's attention draw connections between every word in sentence A and every word in sentence B. A small head on top turns the result into one number: how similar the pair is.
Bi-encoders can't do this. They process each sentence separately, create an embedding, then measure geometric similarity. It's like forming an opinion about two people by looking at their passport photos side by side instead of watching them interact.
The cost: you can't pre-compute anything. Every pair needs a fresh forward pass. For a million documents, that's a million forward passes per query instead of one. This is why the usual design uses a bi-encoder or BM25 (keyword-based ranking) for retrieval and a cross-encoder to rerank the top candidates.
The training you are about to run has three parts:
- The data. STS-B pairs, each with a human similarity score. The version you will load stores the score as a number from 0 to 1.
- The loss. With one output, a cross-encoder in Sentence Transformers trains with binary cross-entropy. The loss applies a sigmoid to the model's raw output, which turns it into a number between 0 and 1, and compares that with the target. A target like 0.76 is a soft label: the loss is lowest when the prediction is exactly 0.76, not 1. So the loss only makes sense for targets between 0 and 1, which is why the score range matters. At prediction time the same sigmoid is applied, so predicted scores are between 0 and 1 too. One consequence you will see in Step 3: even a perfect prediction leaves a loss above zero for any target that is not exactly 0 or 1. On the STS-B training scores, that floor is 0.483.
- The check. After each round of training, an evaluator scores the validation pairs and compares the model's scores with the human scores in two ways. Pearson correlation measures how closely the scores follow a straight-line relationship. Spearman correlation only asks whether the model puts the pairs in the same order as the humans did. A reranker only needs the order to be right, so Spearman is the number this tutorial tracks. For both, 1 is a perfect match and 0 is no relationship.
When you actually need this
Before you invest time fine-tuning a cross-encoder, make sure you're solving the right problem. It makes sense in these cases:
You have a re-ranking problem. You've already narrowed down candidates using BM25 or a bi-encoder. Now you need to score the top 10-100 results with high precision. Classic use case: search engines, question-answering systems, recommendation re-rankers. This RAG engineering piece makes the production case for adding a reranking stage.
Accuracy directly impacts your outcome. If you're building a duplicate detection system for support tickets and false positives waste hours of human time, the accuracy gain is worth the compute cost. If you're just doing approximate clustering, probably not.
You have labeled training data. Fine-tuning requires pairs of sentences with ground-truth similarity scores or classification labels. If you're starting from scratch, collecting this data is your main bottleneck, not the model training.
Your domain has specific similarity patterns. Legal documents, medical records, customer support conversations - these have domain-specific ways of expressing similarity that generic models miss. Fine-tuning teaches the model your domain's equivalences. Embedding models run into the same domain gap.
When not to use a cross-encoder
You need real-time retrieval from millions of documents. Cross-encoders can't pre-compute embeddings. For semantic search across large corpora, use bi-encoders for retrieval and cross-encoders for re-ranking the top 10-100 results.
Your task is symmetric similarity at scale. If you just need to cluster documents or find near-duplicates across a large set, bi-encoders are faster and work fine. Cross-encoders do best when you have a query-document asymmetry or a small number of pairs to judge precisely.
You don't have labeled training data. Pre-trained cross-encoders exist (cross-encoder/ms-marco-MiniLM-L6-v2 for search, cross-encoder/stsb-roberta-base for similarity), but if you need domain-specific fine-tuning and don't have labels, you'll need to collect or generate them first - and if you generate them with a hosted model, record which pairs came from it before they reach your training set.
Your accuracy requirements are loose. If you're building a "related posts" feature where approximate matches are fine, the extra complexity isn't worth it. Use a bi-encoder.
4. Steps: fine-tune and use the cross-encoder
Step 1: Load STS-B and check the score range
Goal: load the STS Benchmark and confirm its scores are already between 0 and 1.
Why this step: the score range decides whether training works at all. The loss you will use in Step 3 needs targets between 0 and 1. The sentence-transformers/stsb copy of the benchmark already divides the original 0-5 scores by 5, so you can use it directly. Checking it once here is cheaper than finding out from a strange training log.
Create data.py in the project folder:
File: data.py
from datasets import load_datasetdef load_stsb(): """STS-B sentence pairs with similarity scores already scaled to 0-1.""" return load_dataset("sentence-transformers/stsb")if __name__ == "__main__": stsb = load_stsb() for split in ("train", "validation", "test"): print(f"{split:<10} {stsb[split].num_rows:>5} pairs") print("columns:", stsb["train"].column_names) row = stsb["train"][0] print("first pair:", row["sentence1"], "|", row["sentence2"], "|", row["score"]) scores = stsb["train"]["score"] print(f"score range: {min(scores)} to {max(scores)}")Run it:
python data.pyExpected output:
train 5749 pairsvalidation 1500 pairstest 1379 pairscolumns: ['sentence1', 'sentence2', 'score']first pair: A plane is taking off. | An air plane is taking off. | 1.0score range: 0.0 to 1.0The first run also shows download progress, and it may print two warnings: an "unauthenticated requests" warning and, on Windows, a symlinks warning. Both are harmless; see When it breaks.
What just happened: you downloaded the benchmark, and it is now cached on disk. You train on train, check progress on validation, and keep test untouched for a final check. Every score is between 0.0 and 1.0, so the labels can go straight into the loss.
Step 2: Measure the untrained baseline
Goal: measure how well distilroberta-base scores the validation pairs before any training.
Why this step: a number after training means little without a number before it. distilroberta-base is a general language model, and its similarity head starts with random weights. Measuring it now gives you the floor that training has to beat. The training script also reuses the evaluation code you write here.
Create evaluate.py:
File: evaluate.py
import sysfrom sentence_transformers import CrossEncoderfrom sentence_transformers.cross_encoder.evaluation import ( CrossEncoderCorrelationEvaluator,)from transformers import set_seedfrom data import load_stsbdef make_evaluator(split="validation"): rows = load_stsb()[split] return CrossEncoderCorrelationEvaluator( sentence_pairs=list(zip(rows["sentence1"], rows["sentence2"])), scores=list(rows["score"]), name=f"stsb-{split}", )if __name__ == "__main__": model_path = sys.argv[1] split = sys.argv[2] if len(sys.argv) > 2 else "validation" set_seed(42) # fixes the random scoring head of an untrained model model = CrossEncoder(model_path, num_labels=1) results = make_evaluator(split)(model) for name, value in results.items(): print(f"{name}: {value:.4f}")make_evaluator builds a CrossEncoderCorrelationEvaluator from one split: the list of sentence pairs, and the human scores to compare against. The script loads whatever model path you pass on the command line. Point it at the base model now and at your trained model later. set_seed(42) fixes the random starting weights of the untrained head, so your baseline matches the one below.
Run it:
python evaluate.py distilroberta-baseExpected output:
stsb-validation_pearson: 0.0114stsb-validation_spearman: 0.0156A LOAD REPORT table also prints above the scores. Both of its warnings are expected. The pretrained model's pooler and language-modelling head weights show as UNEXPECTED, because a cross-encoder does not use them. Four classifier weights show as MISSING: that is the scoring head, created with random weights.
What just happened: the untrained model's scores have almost no relationship to the human scores. A correlation of 0.0156 is what random guessing gives.
Step 3: Train the cross-encoder with CrossEncoderTrainer
Goal: fine-tune distilroberta-base on the STS-B training pairs and save the result.
Why this step: you train with CrossEncoderTrainer, the training loop Sentence Transformers recommends since version 4. It is built on the Hugging Face Trainer, so it handles batching, the learning-rate schedule and periodic evaluation for you. The older model.fit() method still runs, but it is kept for backward compatibility and hides the settings you are about to see.
Create train.py:
File: train.py
from sentence_transformers import CrossEncoderfrom sentence_transformers.cross_encoder import ( CrossEncoderTrainer, CrossEncoderTrainingArguments,)from sentence_transformers.cross_encoder.losses import BinaryCrossEntropyLossfrom transformers import set_seedfrom transformers.trainer_callback import PrinterCallbackfrom data import load_stsbfrom evaluate import make_evaluatorBASE_MODEL = "distilroberta-base"OUTPUT_DIR = "models/stsb-distilroberta"METRIC = "eval_stsb-validation_spearman"set_seed(42)stsb = load_stsb()model = CrossEncoder(BASE_MODEL, num_labels=1)loss = BinaryCrossEntropyLoss(model)args = CrossEncoderTrainingArguments( output_dir=OUTPUT_DIR, num_train_epochs=2, per_device_train_batch_size=16, learning_rate=2e-5, warmup_steps=0.1, # a float below 1 is a fraction of all training steps eval_strategy="steps", eval_steps=180, logging_steps=180, save_strategy="no", disable_tqdm=True, seed=42,)trainer = CrossEncoderTrainer( model=model, args=args, train_dataset=stsb["train"], loss=loss, evaluator=make_evaluator("validation"),)# The default printer writes each log as one very long dict; print a table instead.trainer.remove_callback(PrinterCallback)trainer.train()history = trainer.state.log_historytrain_loss = {row["step"]: float(row["loss"]) for row in history if "loss" in row}print("step train loss validation spearman")for row in history: if METRIC in row: step = row["step"] print(f"{step:>4} {train_loss[step]:>10.4f} {float(row[METRIC]):>19.4f}")model.save_pretrained(f"{OUTPUT_DIR}/final")print("saved to", f"{OUTPUT_DIR}/final")The settings, in order of how much they matter:
BinaryCrossEntropyLoss(model)is the loss for pair scores between 0 and 1.num_train_epochs=2passes over the 5,749 training pairs twice. At a batch size of 16 that is 360 steps per epoch, 720 in total.learning_rate=2e-5withwarmup_steps=0.1starts the learning rate at zero and raises it over the first 10% of steps. This keeps the new random head from damaging the pretrained weights early on.eval_steps=180runs the evaluator from Step 2 four times during training, so you can watch the correlation climb.save_strategy="no"skips intermediate checkpoints; the script saves the final model itself.disable_tqdm=Trueturns off the progress bar, andremove_callback(PrinterCallback)stops the trainer from printing every log as one very long line. The script prints its own table instead, fromtrainer.state.log_history.
METRIC does not affect training; it only names the result the table reports. The trainer adds an eval_ prefix to every evaluator result, so stsb-validation_spearman from the evaluator becomes eval_stsb-validation_spearman.
The table pairs each evaluation with the training loss logged at the same step. So keep logging_steps equal to eval_steps, and choose a value that divides the total number of steps: 720 here, from 5,749 pairs at 16 per batch for two epochs. If you change the batch size or the number of epochs, adjust both so they still line up. If they do not, the last evaluation has no matching loss, and the table stops with a KeyError.
The terminal stays quiet while training runs: the table appears all at once when training ends, after about 43 minutes on a 4-thread CPU. A pin_memory warning at the start is expected on a CPU. The script saves the model only at the end, so stopping it early keeps nothing. If you want to watch progress, set disable_tqdm=False for a progress bar.
Run it:
python train.pyExpected output (43 minutes on a 4-thread CPU; your numbers can differ in the last digits on other hardware):
step train loss validation spearman 180 0.6184 0.8439 360 0.5582 0.8661 540 0.5382 0.8768 720 0.5383 0.8776saved to models/stsb-distilroberta/finalWhat just happened: the model went from a validation Spearman of 0.0156 to 0.8776 in 720 steps. Most of the gain came in the first epoch. The second added about 0.01, and the last two rows are almost level, so more epochs on this data would add little.
The training loss settles near 0.54 instead of falling toward zero. That is normal here. With targets between 0 and 1, binary cross-entropy has a floor above zero even for perfect predictions. On this training set the floor is 0.483, so a final loss of 0.538 is close to the best possible. A negative loss would mean the labels are outside 0-1 (see When it breaks). The trained model is now in models/stsb-distilroberta/final.
Step 4: Check the trained model on the test split
Goal: measure the trained model on test pairs it has never seen.
Why this step: you watched the validation score while training, so it could partly reflect your own choices. The test split was never used, so it gives the honest number. The Step 2 script already handles it.
Run it:
python evaluate.py models/stsb-distilroberta/final testExpected output:
stsb-test_pearson: 0.8451stsb-test_spearman: 0.8329What just happened: on pairs it never saw, the model scores a Spearman of 0.8329, a little below its validation score, as expected. For comparison, the Sentence Transformers team's own distilroberta-base run with the same trainer (4 epochs, on a GPU) reports 0.8389 on this test split. Two epochs on a CPU land within 0.006 of it.
Step 5: Score sentence pairs with the trained model
Goal: get a similarity score for any pair of sentences.
Why this step: duplicate detection and similarity checks use the model in exactly this way. You give it pairs and get one score per pair. predict applies the sigmoid after the model runs, so the scores are between 0 and 1, on the same scale as the training labels.
Create score_pairs.py:
File: score_pairs.py
from sentence_transformers import CrossEncodermodel = CrossEncoder("models/stsb-distilroberta/final")pairs = [ ("A man is playing a guitar.", "A person is playing a guitar."), ("A dog is running in the park.", "A cat is sleeping on the couch."), ("Python is a programming language.", "Python is a type of snake."),]scores = model.predict(pairs)for (first, second), score in zip(pairs, scores): print(f"{score:.3f} {first} | {second}")Run it:
python score_pairs.pyExpected output:
0.757 A man is playing a guitar. | A person is playing a guitar.0.042 A dog is running in the park. | A cat is sleeping on the couch.0.366 Python is a programming language. | Python is a type of snake.What just happened: the scores follow the STS-B scale the model learned from. The guitar pair gets 0.757 rather than 1.0 because "a man" and "a person" are close but not identical, and STS-B annotators score pairs like this around 4 out of 5. The dog and cat sentences share no meaning and get 0.042. The two Python sentences share a word but not a topic. They land in between, at 0.366.
Step 6: Rerank search candidates with rank()
Goal: order a list of candidate answers by how well each matches a query.
Why this step: reranking is the main production use of a cross-encoder. A fast retriever returns candidates, and the cross-encoder reorders them by reading each one together with the query.
Create rerank.py:
File: rerank.py
from sentence_transformers import CrossEncodermodel = CrossEncoder("models/stsb-distilroberta/final")query = "How do I reset my password?"candidates = [ "Click 'Forgot password' on the sign-in page to get a reset link.", "Our office is closed on public holidays.", "You can change your password under Settings > Security.", "Passwords must be at least 12 characters long.",]for hit in model.rank(query, candidates, return_documents=True): print(f"{hit['score']:.3f} {hit['text']}")rank returns one result per candidate, best first, each with a score and the candidate's position in the list. return_documents=True adds the candidate text itself, which is what the loop prints as hit['text'].
Run it:
python rerank.pyExpected output:
0.602 Click 'Forgot password' on the sign-in page to get a reset link.0.541 You can change your password under Settings > Security.0.353 Passwords must be at least 12 characters long.0.083 Our office is closed on public holidays.What just happened: rank paired the query with each candidate, scored the four pairs and returned them best-first. The two sentences that explain how to reset or change a password come first. The password-length rule is about passwords but does not answer the question, so it comes third. The office hours come last.
The model was never trained on support questions. It learned what similar meaning looks like from STS-B and applies that here. For a production reranker, fine-tune on query and answer pairs from your own domain, as described in Where to go next. To measure how much a reranker actually improves retrieval, this cross-encoder deep dive scores the gain in NDCG@10 on a real benchmark.
5. When it breaks: cross-encoder training errors and their fixes
Every error below was produced on purpose on the versions this tutorial pins, except the missing-accelerate message, which is quoted from the Sentence Transformers 6.1.0 source.
Negative training loss: loss -0.6496..., then -2.88..., -5.02...
Your labels are outside 0-1. Binary cross-entropy assumes every target is between 0 and 1. With a target above 1 it has no lower bound, so the loss falls below zero and keeps falling while the model's outputs grow without limit. These numbers come from training on this tutorial's data with the scores multiplied by 5, which is exactly what happens if you load the original 0-5 STS-B scores: the loss read -0.65, -2.88, -5.03 and -5.16 over the first 40 steps. The correlation such a run reports is not meaningful. Divide 0-5 scores by 5, or use sentence-transformers/stsb, whose scores are already scaled.
HfUriError: Invalid HF URI ... Repository id must be 'namespace/name', got 'stsb_multi_mt'.
Older tutorials load STS-B by passing the bare dataset name stsb_multi_mt, with the "en" configuration, to load_dataset. Current datasets releases no longer resolve dataset names without an owner, so on a clean install that call fails with this error. The full error also names a commit hash, which differs between machines. The same data is at PhilipMay/stsb_multi_mt, but its label column is similarity_score on a 0-5 scale, which causes the next two problems. Use sentence-transformers/stsb instead.
ValueError: BinaryCrossEntropyLoss expects a dataset with two non-label columns, but got a dataset with 3 columns.
The trainer treats the first column named label, labels, score or scores as the label, and every other column as model input. A dataset whose label column has another name, such as similarity_score, looks like three text columns and no label. Rename the column to score with dataset.rename_column("similarity_score", "score"), and remove any extra columns, such as an idx column, with dataset.remove_columns(...).
ValueError: Your setup doesn't support bf16/gpu. You need to assign use_cpu if you want to train the model on CPU.
You set bf16=True, usually because you copied it from the official Sentence Transformers training script, which targets GPUs. Remove it on a CPU. On a GPU that supports it, it speeds training up.
To train a CrossEncoder model, you need to install the `accelerate` and `datasets` modules. You can do so with the `train` extra:
The trainer needs two packages that plain pip install sentence-transformers does not install. Install the extra: pip install "sentence-transformers[train]==6.1.0". The install check in section 2 imports accelerate so that this shows up before training rather than after.
The `warmup_ratio` argument is deprecated in Transformers v5+ ...
Training continues after this warning, because Sentence Transformers converts the setting for you. Transformers 5 removed warmup_ratio. Pass the fraction to warmup_steps instead, as train.py does with warmup_steps=0.1.
The CECorrelationEvaluator has been renamed to CrossEncoderCorrelationEvaluator. Please use CrossEncoderCorrelationEvaluator instead.
Code written for Sentence Transformers 3 still runs, with this warning. Import CrossEncoderCorrelationEvaluator instead, as evaluate.py does. The old model.fit(...) training method also still runs, but it takes different arguments from the trainer. It has no learning_rate or gradient_accumulation_steps parameters, for example, so tips written for the trainer fail there with a TypeError.
UserWarning: 'pin_memory' argument is set as true but no accelerator is found, then device pinned memory won't be used.
You can ignore this on a CPU. The trainer asks for pinned memory, which only helps when copying batches to a GPU.
UserWarning: `huggingface_hub` cache-system uses symlinks by default to efficiently store duplicated files but your machine does not support them
This one appears only on Windows. Downloads still work; the cache just stores some files twice. Enable Windows Developer Mode, or set HF_HUB_DISABLE_SYMLINKS_WARNING=1 to hide the message.
Warning: You are sending unauthenticated requests to the HF Hub.
Harmless. The model and dataset are public, so no token is needed. Set HF_TOKEN only if downloads get rate-limited.
6. How the pieces connect
Two things come from the Hugging Face Hub: the pretrained distilroberta-base weights and the STS-B pairs. train.py combines them with a loss that expects 0-1 targets. It checks itself against the validation split through the same evaluate.py you used for the baseline. Everything after training reads one folder, models/stsb-distilroberta/final. Copy that folder to another machine and the three scripts at the bottom work unchanged.
direction: down
hub: "Hugging Face Hub" {
style: {fill: "#F4F6F7"; stroke: "#95A5A6"; font-color: "#2C2C2A"}
stsb: "sentence-transformers/stsb\ntrain | validation | test" {shape: cylinder; style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}}
base: "distilroberta-base\n(pretrained weights)" {shape: cylinder; style: {fill: "#98D8C8"; stroke: "#5FA897"; font-color: "#2C2C2A"}}
}
data: "data.py\nload_stsb()" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
evaluate: "evaluate.py\nmake_evaluator(split)" {style: {fill: "#4A90E2"; stroke: "#2C6FB0"; font-color: "#FFFFFF"}}
train: "train.py" {
style: {fill: "#EDEBFB"; stroke: "#7B68EE"; font-color: "#2C2C2A"}
model: "CrossEncoder\nnum_labels=1" {style: {fill: "#7B68EE"; stroke: "#5543C4"; font-color: "#FFFFFF"}}
loss: "BinaryCrossEntropyLoss\n(targets 0-1)" {style: {fill: "#FFD93D"; stroke: "#C9A800"; font-color: "#2C2C2A"}}
trainer: "CrossEncoderTrainer\n2 epochs, 720 steps" {style: {fill: "#7B68EE"; stroke: "#5543C4"; font-color: "#FFFFFF"}}
model -> trainer
loss -> trainer
}
saved: "models/stsb-distilroberta/final" {shape: cylinder; style: {fill: "#6BCF7F"; stroke: "#3E9E52"; font-color: "#2C2C2A"}}
use: "Using the model" {
label.near: bottom-center
style: {fill: "#F4F6F7"; stroke: "#95A5A6"; font-color: "#2C2C2A"}
test: "evaluate.py ... test\ncorrelation on unseen pairs" {style: {fill: "#FFA07A"; stroke: "#D9774F"; font-color: "#2C2C2A"}}
score: "score_pairs.py\npredict(pairs)" {style: {fill: "#FFA07A"; stroke: "#D9774F"; font-color: "#2C2C2A"}}
rerank: "rerank.py\nrank(query, candidates)" {style: {fill: "#FFA07A"; stroke: "#D9774F"; font-color: "#2C2C2A"}}
}
hub.stsb -> data
hub.base -> train.model
data -> train.trainer: "train split"
data -> evaluate
evaluate -> train.trainer: "validation evaluator\nevery 180 steps"
train.trainer -> saved: "save_pretrained"
saved -> use.test
saved -> use.score
saved -> use.rerank
7. The complete project
cross-encoder-stsb/├── .venv/├── requirements.txt├── data.py├── evaluate.py├── train.py├── score_pairs.py├── rerank.py└── models/ └── stsb-distilroberta/ └── final/ (written by train.py)File: requirements.txt
sentence-transformers[train]==6.1.0transformers==5.17.0datasets==5.0.1accelerate==1.15.0File: data.py
from datasets import load_datasetdef load_stsb(): """STS-B sentence pairs with similarity scores already scaled to 0-1.""" return load_dataset("sentence-transformers/stsb")if __name__ == "__main__": stsb = load_stsb() for split in ("train", "validation", "test"): print(f"{split:<10} {stsb[split].num_rows:>5} pairs") print("columns:", stsb["train"].column_names) row = stsb["train"][0] print("first pair:", row["sentence1"], "|", row["sentence2"], "|", row["score"]) scores = stsb["train"]["score"] print(f"score range: {min(scores)} to {max(scores)}")File: evaluate.py
import sysfrom sentence_transformers import CrossEncoderfrom sentence_transformers.cross_encoder.evaluation import ( CrossEncoderCorrelationEvaluator,)from transformers import set_seedfrom data import load_stsbdef make_evaluator(split="validation"): rows = load_stsb()[split] return CrossEncoderCorrelationEvaluator( sentence_pairs=list(zip(rows["sentence1"], rows["sentence2"])), scores=list(rows["score"]), name=f"stsb-{split}", )if __name__ == "__main__": model_path = sys.argv[1] split = sys.argv[2] if len(sys.argv) > 2 else "validation" set_seed(42) # fixes the random scoring head of an untrained model model = CrossEncoder(model_path, num_labels=1) results = make_evaluator(split)(model) for name, value in results.items(): print(f"{name}: {value:.4f}")File: train.py
from sentence_transformers import CrossEncoderfrom sentence_transformers.cross_encoder import ( CrossEncoderTrainer, CrossEncoderTrainingArguments,)from sentence_transformers.cross_encoder.losses import BinaryCrossEntropyLossfrom transformers import set_seedfrom transformers.trainer_callback import PrinterCallbackfrom data import load_stsbfrom evaluate import make_evaluatorBASE_MODEL = "distilroberta-base"OUTPUT_DIR = "models/stsb-distilroberta"METRIC = "eval_stsb-validation_spearman"set_seed(42)stsb = load_stsb()model = CrossEncoder(BASE_MODEL, num_labels=1)loss = BinaryCrossEntropyLoss(model)args = CrossEncoderTrainingArguments( output_dir=OUTPUT_DIR, num_train_epochs=2, per_device_train_batch_size=16, learning_rate=2e-5, warmup_steps=0.1, # a float below 1 is a fraction of all training steps eval_strategy="steps", eval_steps=180, logging_steps=180, save_strategy="no", disable_tqdm=True, seed=42,)trainer = CrossEncoderTrainer( model=model, args=args, train_dataset=stsb["train"], loss=loss, evaluator=make_evaluator("validation"),)# The default printer writes each log as one very long dict; print a table instead.trainer.remove_callback(PrinterCallback)trainer.train()history = trainer.state.log_historytrain_loss = {row["step"]: float(row["loss"]) for row in history if "loss" in row}print("step train loss validation spearman")for row in history: if METRIC in row: step = row["step"] print(f"{step:>4} {train_loss[step]:>10.4f} {float(row[METRIC]):>19.4f}")model.save_pretrained(f"{OUTPUT_DIR}/final")print("saved to", f"{OUTPUT_DIR}/final")File: score_pairs.py
from sentence_transformers import CrossEncodermodel = CrossEncoder("models/stsb-distilroberta/final")pairs = [ ("A man is playing a guitar.", "A person is playing a guitar."), ("A dog is running in the park.", "A cat is sleeping on the couch."), ("Python is a programming language.", "Python is a type of snake."),]scores = model.predict(pairs)for (first, second), score in zip(pairs, scores): print(f"{score:.3f} {first} | {second}")File: rerank.py
from sentence_transformers import CrossEncodermodel = CrossEncoder("models/stsb-distilroberta/final")query = "How do I reset my password?"candidates = [ "Click 'Forgot password' on the sign-in page to get a reset link.", "Our office is closed on public holidays.", "You can change your password under Settings > Security.", "Passwords must be at least 12 characters long.",]for hit in model.rank(query, candidates, return_documents=True): print(f"{hit['score']:.3f} {hit['text']}")8. Where to go next
- Train the bigger model. Change
BASE_MODELintrain.pyto"bert-base-uncased"or"roberta-base". Nothing else changes. In a short benchmark on the same CPU, each training step ofbert-base-uncasedtook about 1.5 times as long as adistilroberta-basestep, so plan for a little over an hour instead of 43 minutes. A base-size RoBERTa model is the one the Sentence Transformers team reports around 0.90 on STS-B. - Train on your own pairs. Any dataset with two text columns and a
scorecolumn between 0 and 1 works with the sametrain.py: load it indata.pyinstead of STS-B. Keep the column namessentence1,sentence2andscore, becauseevaluate.pyreads them, and give it the same three splits:trainfortrain.py,validationfor the evaluator during training, andtestfor Step 4. For duplicate detection with yes/no labels, use 0 and 1 as the scores. - Put it behind a retriever. Retrieve the top 50 candidates with BM25 or a bi-encoder, then call
rankon only those. This bake-off measures BM25, dense and hybrid retrieval with a cross-encoder reranking stage, and this hybrid search build wires the same pattern into a service.
9. References
- Sentence Transformers. v6.1.0 release notes (18 September 2026). https://github.com/huggingface/sentence-transformers/releases/tag/v6.1.0
- Sentence Transformers. Cross Encoder training overview. https://sbert.net/docs/cross_encoder/training_overview.html
- Sentence Transformers. Cross Encoder loss overview. https://sbert.net/docs/cross_encoder/loss_overview.html
- Sentence Transformers. Semantic Textual Similarity training example. https://sbert.net/examples/cross_encoder/training/sts/README.html
- Sentence Transformers. Pretrained cross-encoder models. https://sbert.net/docs/cross_encoder/pretrained_models.html
- Sentence Transformers. Migration guide (from
fitto the trainer). https://sbert.net/docs/migration_guide.html - Hugging Face. Transformers v5 migration guide. https://github.com/huggingface/transformers/blob/main/MIGRATION_GUIDE_V5.md
- Hugging Face. sentence-transformers/stsb dataset. https://huggingface.co/datasets/sentence-transformers/stsb
- Aarsen, T. reranker-distilroberta-base-stsb model card (reference scores for the same recipe). https://huggingface.co/tomaarsen/reranker-distilroberta-base-stsb
- Cer, D., Diab, M., Agirre, E., Lopez-Gazpio, I., Specia, L. (2017). SemEval-2017 Task 1: Semantic Textual Similarity. https://arxiv.org/abs/1708.00055
- Reimers, N., Gurevych, I. (2019). Sentence-BERT. EMNLP. https://arxiv.org/abs/1908.10084
Related Articles
- Text Classification using Large Language Models (LLMs)
- LLM Text Clustering and Topic Modeling: HDBSCAN and BERTopic Tutorial
More Articles




