GLiNER4j ONNX — GLiNER2 Base
ONNX export of fastino/gliner2-base-v1 for Java inference via ONNX Runtime.
Part of the GLiNER4j project.
Supported Tasks
| Task | Description |
|---|---|
| Named Entity Recognition | Extract typed entity spans from text with confidence scores |
| Text Classification | Assign labels to text with multi-label support and confidence scores |
Both tasks support entity/label descriptions for improved accuracy and per-call overrides without model reloading.
Repository Structure
├── gliner4j_config.json # Shared model configuration
├── tokenizer.json # Shared HuggingFace tokenizer
├── tokenizer_config.json
├── onnx/ # Base FP32 (~830 MB)
│ ├── ner_full.onnx
│ └── classifier_full.onnx
├── onnx_fp16/ # FP16 (~416 MB, ~50% smaller)
│ ├── ner_full.onnx
│ └── classifier_full.onnx
└── onnx_quantized/ # INT8 dynamic quantization (~208 MB, ~75% smaller)
├── ner_full.onnx
└── classifier_full.onnx
An onnx_optimized_cpu/ folder with the same two files may also be present (ONNX Runtime graph-optimized for CPU).
Model Architecture
Each variant ships two merged, self-contained ONNX graphs — one per task:
| Graph | Description |
|---|---|
ner_full.onnx |
Transformer encoder + span representation + count-aware scoring head (NER) |
classifier_full.onnx |
Transformer encoder + classifier head MLP (Classification) |
The graphs are fused at export time from the encoder and task heads; the intermediate split modules are not published.
Variants
| Variant | Folder | Precision | Size | Use case |
|---|---|---|---|---|
| Base | onnx/ |
FP32 | ~830 MB | Maximum accuracy |
| FP16 | onnx_fp16/ |
FP16 | ~416 MB, ~50% smaller | Good accuracy/size trade-off |
| Quantized | onnx_quantized/ |
INT8 (QUInt8, per-channel) | ~208 MB, ~75% smaller | Smallest footprint, fastest on CPU |
To download a specific variant only:
huggingface-cli download <repo> --include "onnx_fp16/*" "*.json"
Configuration
| Parameter | Value |
|---|---|
| Hidden size | 768 |
| Max span width | 8 |
| Max count | 20 |
| Span mode | SpanMarkerV0 |
| Token pooling | first |
| ONNX opset | 17 |
Usage
Use with GLiNER4j, a Java library for GLiNER2 inference via ONNX Runtime.
Named Entity Recognition
var entities = List.of(
new EntityDefinition("person", "Names of individuals"),
new EntityDefinition("organization", "Company or institution names")
);
var gliner = GLiNER4jNER.load(modelDir, entities);
Map<String, List<EntitySpan>> results = gliner.extract("John works at Google.");
Text Classification
var labels = List.of(
new ClassificationLabel("positive", "Expresses positive sentiment"),
new ClassificationLabel("negative", "Expresses negative sentiment")
);
var classifier = GLiNER4jClassifier.load(modelDir, labels);
List<ClassificationResult> results = classifier.classify("Great product!");
Model Variants
// FP16 variant
var gliner = GLiNER4jNER.load(modelDir, entities, "onnx_fp16");
// Quantized variant
var gliner = GLiNER4jNER.load(modelDir, entities, "onnx_quantized");
License
Apache License 2.0