Qefro Decision-Jef C1 (PyTorch & ONNX FP32)
Qefro Decision-Jef C1 is a high-performance, enterprise-grade decision and tool-routing transformer fine-tuned on the Qefro multi-intent marketplace distribution. It serves as the authoritative reranking decider in a two-stage hierarchical agent router:
User Request
โ
โผ
multilingual-e5-base
โ (Top-3 Functional Candidates)
โผ
Decision-Jef C1 + "unknown" Fallback (K=4 Options)
โ
โผ
Final Chosen Capability / Fallback
This repository provides both the PyTorch checkpoint (c1_best.pt) and the production-certified ONNX FP32 model (c1_decision_jef.onnx + c1_decision_jef.onnx.data) with 100% exact semantic and behavioral parity.
Key Performance Highlights
- Frozen Production Benchmark (525 Cases):
- Overall Accuracy: 92.38% (485 / 525 cases)
- Macro F1 Score: 0.9161
- Micro F1 Score: 0.9238
- Candidate Pool Recall (E5 Top-3 + unknown): 94.48% (496 / 525 cases)
- Semantic Parity (PyTorch vs ONNX):
- Prediction Mismatches: 0 / 525 cases (100.00% agreement)
- Max Logit Difference: 1.64e-04 (mean: 1.97e-05)
- Max Probability Difference: 1.06e-05 (mean: 1.26e-07)
- Order Stability: Identical flip rates across PyTorch, ONNX CPU, and ONNX CUDA.
- Inference Latency & Throughput:
- Tesla T4 GPU (CUDAExecutionProvider): 8.94 ms (p50), 10.81 ms (p99) | 111.4 req/s throughput
- Host CPU (CPUExecutionProvider): 152.00 ms (p50) | 6.2 req/s throughput
- Memory Footprint: 1.67 GB VRAM (GPU), 7.7 GB RSS (CPU)
Per-Capability Performance Breakdown (525 Frozen Cases)
Evaluated across all 14 enterprise capabilities in the Qefro platform:
| Capability | Description | Support | Precision | Recall | F1 Score | Accuracy |
|---|---|---|---|---|---|---|
orders.list |
List recent or historical customer orders | 47 | 0.9020 | 0.9787 | 0.9388 | 97.9% |
orders.status |
Check live status/tracking of an order | 44 | 1.0000 | 0.9091 | 0.9524 | 90.9% |
orders.cancel |
Request cancellation of an order | 43 | 0.8511 | 0.9302 | 0.8889 | 93.0% |
orders.refund |
Request refund for an order or item | 41 | 1.0000 | 0.9268 | 0.9620 | 92.7% |
catalog.search |
Search product catalog for items/stock | 42 | 0.9737 | 0.8810 | 0.9250 | 88.1% |
customers.search |
Look up customer profiles and CRM data | 38 | 0.9444 | 0.8947 | 0.9189 | 89.5% |
invoices.get |
Retrieve invoices, tax documents, or receipts | 38 | 1.0000 | 0.9474 | 0.9730 | 94.7% |
messages.send |
Send individual messages or customer notifications | 35 | 0.8919 | 0.9429 | 0.9167 | 94.3% |
campaigns.send |
Trigger bulk marketing/promotional campaigns | 35 | 0.9688 | 0.8857 | 0.9254 | 88.6% |
records.delete |
Delete database entries or customer records | 35 | 1.0000 | 1.0000 | 1.0000 | 100.0% |
sales.summary |
Aggregate analytics, GMV, and revenue metrics | 35 | 0.9722 | 1.0000 | 0.9859 | 100.0% |
tickets.create |
Create customer support / escalation tickets | 38 | 0.8837 | 1.0000 | 0.9383 | 100.0% |
ambiguous |
Mark underspecified or conflicting requests | 25 | 0.9333 | 0.5600 | 0.7000 | 56.0% |
unknown |
Fallback for out-of-domain requests | 29 | 0.6829 | 0.9655 | 0.8000 | 96.6% |
| Macro Average | โ | 525 | 0.9288 | 0.9209 | 0.9161 | โ |
| Micro / Overall | โ | 525 | 0.9238 | 0.9238 | 0.9238 | 92.38% |
Architectural Details: Segment Isolation & Sliding Attention
Decision-Jef employs a ModernBERT backbone with learned decision (q_proj) and option (o_proj) projection heads. The ONNX model natively computes two 4D boolean attention masks:
- Segment-Isolated Full Attention: State tokens (Segment 0) are mathematically prevented from cross-attending to Question/Option tokens (Segment 1+). This preserves state contextualization independently of candidate ordering.
- Position-ID Sliding Attention: Sliding window layers ($w = 128$, $ ext{half} = 64$) calculate distance using restarted
position_idsrather than absolute token indices, ensuring single-question geometric consistency.
Quickstart: Python Inference with ONNX Runtime
1. Install Dependencies
pip install onnxruntime-gpu transformers numpy
2. Run Inference
import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer
# Load tokenizer and ONNX session
model_id = "developerabu/qefro-decision-jef-c1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
session = ort.InferenceSession("c1_decision_jef.onnx", providers=["CUDAExecutionProvider", "CPUExecutionProvider"])
# Example input query and candidates
user_query = "Please cancel order ORD-99214 immediately"
candidates = ["orders.cancel", "orders.status", "orders.refund", "unknown"]
descriptions = {
"orders.cancel": "Request cancellation of an order",
"orders.status": "Check live status/tracking of an order",
"orders.refund": "Request refund for an order or item",
"unknown": "Fallback when request is out-of-domain or unsupported"
}
# Format question block
# Segment 0: <s> user query </s>
# Segment 1: <unused2> instructions <unused0> opt0 ... <unused1>
bos_id = tokenizer.bos_token_id or tokenizer.cls_token_id
eos_id = tokenizer.eos_token_id or tokenizer.sep_token_id
opt_id = tokenizer.convert_tokens_to_ids("<unused0>")
dec_id = tokenizer.convert_tokens_to_ids("<unused1>")
q_id = tokenizer.convert_tokens_to_ids("<unused2>")
state_ids = [bos_id] + tokenizer(user_query, add_special_tokens=False)["input_ids"] + [eos_id]
seg_ids = [0] * len(state_ids)
pos_ids = list(range(len(state_ids)))
# Question block
q_base = len(state_ids)
q_text = "choice question: Which Qefro capability should handle this request?"
q_tokens = [q_id] + tokenizer(q_text, add_special_tokens=False)["input_ids"]
opt_positions = []
for cand in candidates:
opt_positions.append(len(state_ids) + len(q_tokens))
q_tokens.append(opt_id)
q_tokens.extend(tokenizer(descriptions[cand], add_special_tokens=False)["input_ids"][:48])
dec_position = len(state_ids) + len(q_tokens)
q_tokens.append(dec_id)
all_ids = state_ids + q_tokens
all_segs = seg_ids + [1] * len(q_tokens)
all_poss = pos_ids + list(range(q_base, q_base + len(q_tokens)))
# Prepare ONNX tensors
ort_inputs = {
"input_ids": np.array([all_ids], dtype=np.int64),
"attention_mask": np.ones((1, len(all_ids)), dtype=np.int64),
"segment_ids": np.array([all_segs], dtype=np.int64),
"position_ids": np.array([all_poss], dtype=np.int64),
"batch_idx": np.array([0], dtype=np.int64),
"dec_pos": np.array([dec_position], dtype=np.int64),
"opt_pos": np.array([opt_positions], dtype=np.int64),
"opt_mask": np.ones((1, len(candidates)), dtype=bool),
}
logits, probabilities = session.run(["logits", "probabilities"], ort_inputs)
best_idx = int(np.argmax(probabilities[0]))
print(f"Selected Capability: {candidates[best_idx]} (Confidence: {probabilities[0][best_idx]:.4f})")
Benchmark Artifacts & Reproducibility
c1_best.pt: PyTorch weights checkpoint (Epoch 3, Validation Accuracy: 85.19%, Test Accuracy: 88.89%).c1_decision_jef.onnx: Standalone FP32 ONNX computational graph.c1_decision_jef.onnx.data: External model weights (1.23 GB).evaluation_results_525.json: Full per-case evaluation results across all 525 frozen benchmark instances.capabilities.json: Canonical dictionary of Qefro capability criteria.
Citation & License
Developed and released by Qefro AI for enterprise agent routing.
Licensed under the Apache 2.0 License.
- Downloads last month
- 15