Skip to main content

Training

Fine-tune a GLiFormer checkpoint on annotated data with train_model. This example adapts the model for named entity recognition.

Prepare Training Data​

NER records contain a text and an extraction list. Each extraction group supplies candidate labels through all_labels and entity mention/label pairs through ner:

train_data = [
{
"text": "Alice works at Acme.",
"extraction": [
{
"name": "entities",
"all_labels": ["person", "organization", "location"],
"ner": [["Alice", "person"], ["Acme", "organization"]],
}
],
},
{
"text": "Bob lives in Berlin.",
"extraction": [
{
"name": "entities",
"all_labels": ["person", "organization", "location"],
"ner": [["Bob", "person"], ["Berlin", "location"]],
}
],
},
]

This small dataset illustrates the format. For your task, use a representative training set and separate held-out examples for evaluation.

Fine-Tune and Save​

from gliformer import GLiFormer

model = GLiFormer.from_pretrained(
"knowledgator/gliformer-base-v1",
load_tokenizer=True,
)

trainer = model.train_model(
train_dataset=train_data,
output_dir="outputs/ner",
max_steps=100,
per_device_train_batch_size=2,
learning_rate=1e-5,
)
model.save_pretrained("outputs/ner/final")

Reload the saved model with GLiFormer.from_pretrained("outputs/ner/final", load_tokenizer=True).

Training Options​

OptionPurpose
eval_datasetHeld-out annotated examples for evaluation
freeze_components=["text_encoder"]Freeze the text encoder during training
train_head_only=TrueTrain task heads while freezing other components
resume_from_checkpoint="path/to/training-checkpoint"Resume a saved training run

Annotations for other tasks must match their processors and enabled heads. See the task processors, sample records, and training configurations for additional formats and settings.

Evaluation​

Run task-specific evaluators from the GLiFormer repository. For example, after preparing CrossNER files:

python gliformer_eval/eval_ner.py \
--model outputs/ner/final \
--data data/NER \
--datasets CrossNER_AI CrossNER_literature CrossNER_music CrossNER_politics CrossNER_science \
--output eval_results/gliformer_ner.json

Each dataset directory must contain labels.json and test.json. See the evaluation guide for classification, relation extraction, structuring, and similarity evaluation.