gliner_datause_extended

Fine-tune of urchade/gliner_large-v2.1 for data-use mention extraction (dataset / survey / census / registry mentions in economics research papers).

Labels

  • NAMED_DATA โ€” a proper name, title, or acronym of a specific data source
  • DESCRIPTIVE_DATA โ€” a source described in words but not named
  • VAGUE_DATA โ€” generic data wording with no identifiable source

Training

  • base model: urchade/gliner_large-v2.1
  • dataset: rafmacalaba/data-use-mentions-extended (gliner config)
  • epochs: 5
  • learning rate: 5e-06
  • batch size: 16
  • precision: bf16

Evaluation (holdout)

thr tp fp fn precision recall f0.5 f1
0.10 12275 8853 289 0.5810 0.9770 0.6322 0.7287
0.20 12153 6445 411 0.6535 0.9673 0.6988 0.7800
0.30 12013 5218 551 0.6972 0.9561 0.7371 0.8064
0.40 11844 4143 720 0.7409 0.9427 0.7740 0.8297
0.50 11489 3002 1075 0.7928 0.9144 0.8145 0.8493
0.60 10542 1935 2022 0.8449 0.8391 0.8437 0.8420
0.70 8396 940 4168 0.8993 0.6683 0.8411 0.7668

Best F0.5: 0.8437 (thr=0.6) Best F1: 0.8493 (thr=0.5)

NER holdout comparison

device: NVIDIA H100 NVL

rafmacalaba/data-use-mentions-extended (n=9249)

model backend best F0.5 thr best F1 thr wall-clock (s) texts/s
ai4data/gliner_datause gliner 0.8437 0.6 0.8493 0.5 200.9 46.0
ai4data/gliner2_datause gliner2 0.8634 0.7 0.8624 0.6 181.0 51.1

F0.5 by threshold (sweet spots side-by-side):

thr ai4data/gliner_datause ai4data/gliner2_datause
0.1 0.6321 0.7321
0.2 0.6987 0.7712
0.3 0.7371 0.7979
0.4 0.7740 0.8201
0.5 0.8145 0.8363
0.6 0.8437 0.8523
0.7 0.8411 0.8634
Downloads last month
7
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support