Instructions to use onnx-community/gliner2-multi-v1-agent-ONNX with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers.js
How to use onnx-community/gliner2-multi-v1-agent-ONNX with Transformers.js:
// npm i @huggingface/transformers import { pipeline } from '@huggingface/transformers'; // Allocate pipeline const pipe = await pipeline('token-classification', 'onnx-community/gliner2-multi-v1-agent-ONNX'); - GLiNER2
How to use onnx-community/gliner2-multi-v1-agent-ONNX with GLiNER2:
from gliner2 import GLiNER2 model = GLiNER2.from_pretrained("onnx-community/gliner2-multi-v1-agent-ONNX") # Extract entities text = "Apple CEO Tim Cook announced iPhone 15 in Cupertino yesterday." result = extractor.extract_entities(text, ["company", "person", "product", "location"]) print(result) - Notebooks
- Google Colab
- Kaggle
GLiNER2 multi v1, one-graph ONNX for the browser
ONNX export of fastino/gliner2-multi-v1 (mDeBERTa-v3-base, multilingual) as a single graph that serves both entity extraction and label classification, for Transformers.js on WebGPU. It was made for Zipline, a browser agent that runs entirely in Chrome, and covers the two calls such an agent makes: pulling values out of a goal (extract_entities) and scoring page controls against it (classify with a softmax over labels).
| file | size |
|---|---|
onnx/model_fp16.onnx + _data |
614 MB (WebGPU) |
onnx/model.onnx + _data |
1.23 GB (fp32, reference) |
Inputs and outputs
| name | shape | meaning |
|---|---|---|
input_ids, attention_mask |
[1, seq] |
the GLiNER2 processor's prompt: ( [P] task [DESCRIPTION] label: desc … ( [E]/[L] label … ) ) [SEP_TEXT] words |
word_positions |
[1, words] |
index of the first sub-token of each text word |
schema_positions |
[1, 1 + labels] |
index of [P], then of each [E] (extraction) or [L] (classification) marker |
cls_logits |
[1, labels] |
classifier head on the markers; softmax for a single-label task |
count_logits |
[1, 20] |
instance count from [P]; extraction returns nothing when the argmax is 0 |
span_logits |
[1, labels, words, 8] |
instance-0 span scores (pre-sigmoid) for spans of 1–8 words |
Entity decoding is the library's default: sigmoid ≥ 0.5, map word spans to characters, then keep spans greedily by confidence without character overlap. Entity extraction only reads instance 0, so the count GRU is unrolled once (export_onnx.py asserts that instance 0 does not depend on the unroll length).
Verification
conversion/reference.json holds 14 calls recorded from the Python gliner2 library (six goals across English and French, eight control-scoring calls on flight-search and maps pages), with their exact token ids. Against it:
| dtype | same entities / top label | worst Δ (classify p / entity confidence) |
|---|---|---|
| fp32 | 14/14 | 1e-5 / 5e-7 |
| fp16 | 14/14 | 3.3e-3 / 1.3e-4 |
The JavaScript runtime reproduces the token ids exactly. On WebGPU (fp16, Apple M3 Pro): 35–37 ms per extraction with nine entity types, 43–44 ms per classification with twelve labels.
Rebuilding
pip install torch "gliner2[local]" onnx onnxruntime onnxscript transformers
python conversion/reference.py # Python reference outputs
python conversion/export_onnx.py # fp32 graph
python conversion/convert_fp16.py # fp16 graph
python conversion/verify_onnx.py # parity check
(The scripts write to dist-model/; external data is renamed to the .onnx_data convention afterwards.)
Credits
Model by Fastino (GLiNER2), Apache-2.0. The agent calls these outputs serve come from gliner2-ultrafast (MIT).
- Downloads last month
- 20
Model tree for onnx-community/gliner2-multi-v1-agent-ONNX
Base model
fastino/gliner2-multi-v1