| Label | Examples |
|---|---|
| Asset | "mental health", "water resources", "raw material" |
| Body Part | "plant leaves", "deep tissue compartment", "leaves" |
| Body of Water | "Dhaleshwari river", "rivers", "peripheral rivers" |
| Chemical | "marine algal toxin", "domoic acid", "cathode materials" |
| Disease | "acute neurologic signs", "chronic epileptic syndrome", "seizures" |
| Ecosystem | "cloud forests", "Tropical montane cloud forest", "polluted environment" |
| Energy Source | "battery cells", "fossil fuels", "12-cell series battery-pack prototype" |
| Field of Study | "study", "veterinary medicine", "reference laboratory" |
| Geographical Feature | "heterogenous topography", "low point", "mountainous regions" |
| Intellectual Artefact | "Veterinary medical records", "data", "Daily husbandry records" |
| Location | "wild", "Westbrook", "beaches" |
| Mathematical Expression | "Stepwise machine hour constraints", "difference", "gradient" |
| Measuring Device | "MRI scan", "station", "EEG" |
| Meteorological Phenomenon | "climatic variability", "climate change", "rainfall" |
| Method | "serum monitoring", "clinical efficacy", "dosing" |
| Natural Disaster | "seasonal air pollution", "heavy metal contamination", "environmental pollution" |
| Natural Phenomenon | "biochemical changes", "changing ocean conditions", "algal blooms" |
| Organism | "Zalophus californianus", "California sea lions", "species" |
| Organization | "NOAA National Marine Fisheries Service", "reference laboratory", "long-term care facility" |
| Other | "reports", "normal eating", "marine mammal health" |
| Person | "staff", "clinicians", "Clinicians" |
| Physical Artefact | "electric vehicle", "paved east – west road", "EVs" |
| Physical Phenomenon | "normal food intake", "seasonal changes", "structural abnormalities" |
| Policy | "safety", "pollution", "energy security" |
| Quantity | "energy density", "200 mAhg − 1", ">" |
| Satellite | "satellites", "TRMM", "Tropical Rainfall Measuring Mission" |
| System | "climate", "system structure", "global overturning circulation" |
| Time Period | "101 days", "periods of prolonged anorexia", "several decades" |
| Metric | Score |
|---|---|
| Precision | 56.49 |
| Recall | 49.63 |
| F1 | 52.84 |
This checkpoint corresponds to the seed with the highest strict F1 on the gold evaluation set (Seed 3 - 3012).
| Seed | Precision | Recall | Strict F1 |
|---|---|---|---|
| 1 | 54.41 | 45.14 | 49.34 |
| 2 | 45.58 | 37.05 | 40.87 |
| 3 | 56.49 | 49.63 | 52.84 |
| 4 | 53.84 | 48.69 | 51.14 |
| 5 | 53.31 | 45.34 | 49.01 |
span_marker library for inference.pip install span_marker1from span_marker import SpanMarkerModel
2
3# Download from the 🤗 Hub
4model = SpanMarkerModel.from_pretrained("P0L3/CliReNER-sciclimatebert")
5
6# Run inference
7text = "The effectiveness of these approaches has been demonstrated in high-resource domains, including biomedicine and chemistry (Lee et al. 2019; Fries et al. 2022; Morin et al. 2023; Wang et al. 2021). "
8entities = model.predict(text)
9
10for entity in entities:
11 print(f"Entity: {entity['span']} | Label: {entity['label']} | Score: {entity['score']:.4f}")
12
13# Entity: effectiveness | Label: Quantity | Score: 0.7880
14# Entity: approach | Label: Method | Score: 0.8180
15# Entity: high-resource domains | Label: Location | Score: 0.2892
16# Entity: biomedicine | Label: Field of Study | Score: 0.2845
17# Entity: chemistry | Label: Field of Study | Score: 0.7580
18# Entity: Lee et al. | Label: Person | Score: 0.9177
19# Entity: Fries et al. | Label: Person | Score: 0.9175
20# Entity: 2022 | Label: Time Period | Score: 0.9332
21# Entity: Morin et al. | Label: Person | Score: 0.9583
22# Entity: Wang et al. | Label: Person | Score: 0.9435
23# Entity: 2021 | Label: Time Period | Score: 0.81851from span_marker import SpanMarkerModel, Trainer
2from datasets import load_dataset
3
4# Download from the 🤗 Hub
5model = SpanMarkerModel.from_pretrained("your-huggingface-username/your-model-name")
6
7# Specify a Dataset with "tokens" and "ner_tags" columns
8dataset = load_dataset("your_custom_dataset")
9
10# Initialize a Trainer using the pretrained model & dataset
11trainer = Trainer(
12 model=model,
13 train_dataset=dataset["train"],
14 eval_dataset=dataset["validation"],
15)
16trainer.train()
17trainer.save_model("span_marker_model_id-finetuned")| Training set | Min | Median | Max |
|---|---|---|---|
| Sentence length | 3 | 31.4819 | 97 |
| Entities per sentence | 1 | 7.0100 | 22 |
| Epoch | Step | Validation Loss | Validation Precision | Validation Recall | Validation F1 | Validation Accuracy |
|---|---|---|---|---|---|---|
| 1.0 | 62 | 0.1522 | 0.0 | 0.0 | 0.0 | 0.6075 |
| 2.0 | 124 | 0.1065 | 0.0 | 0.0 | 0.0 | 0.6075 |
| 3.0 | 186 | 0.0703 | 0.4503 | 0.2209 | 0.2964 | 0.6975 |
| 4.0 | 248 | 0.0539 | 0.5494 | 0.3831 | 0.4514 | 0.7647 |
| 5.0 | 310 | 0.0499 | 0.5369 | 0.5222 | 0.5295 | 0.8056 |
| 6.0 | 372 | 0.0453 | 0.5947 | 0.5452 | 0.5689 | 0.8153 |
| 7.0 | 434 | 0.0461 | 0.6125 | 0.5897 | 0.6009 | 0.8316 |
| 8.0 | 496 | 0.0452 | 0.6033 | 0.5739 | 0.5882 | 0.8256 |
| 9.0 | 558 | 0.0483 | 0.5882 | 0.5882 | 0.5882 | 0.8283 |
| 10.0 | 620 | 0.0486 | 0.6175 | 0.5882 | 0.6025 | 0.8268 |
| 11.0 | 682 | 0.0491 | 0.5860 | 0.5868 | 0.5864 | 0.8234 |
1@misc{poleksic2026named,
2 author = {Poleksić, Andrija and Martinčić-Ipšić, Sanda},
3 title = {Named Entity Recognition for Climate Change Research},
4 year = {2026},
5 howpublished = {Research Square},
6 note = {Preprint}
7}1@software{Aarsen_SpanMarker,
2 author = {Aarsen, Tom},
3 license = {Apache-2.0},
4 title = {{SpanMarker for Named Entity Recognition}},
5 url = {https://github.com/tomaarsen/SpanMarkerNER}
6}