Qefro Decision-Jef C1 (PyTorch & ONNX FP32)

Qefro Decision-Jef C1 is a high-performance, enterprise-grade decision and tool-routing transformer fine-tuned on the Qefro multi-intent marketplace distribution. It serves as the authoritative reranking decider in a two-stage hierarchical agent router:

User Request
     โ”‚
     โ–ผ
multilingual-e5-base
     โ”‚ (Top-3 Functional Candidates)
     โ–ผ
Decision-Jef C1 + "unknown" Fallback (K=4 Options)
     โ”‚
     โ–ผ
Final Chosen Capability / Fallback

This repository provides both the PyTorch checkpoint (c1_best.pt) and the production-certified ONNX FP32 model (c1_decision_jef.onnx + c1_decision_jef.onnx.data) with 100% exact semantic and behavioral parity.


Key Performance Highlights

  • Frozen Production Benchmark (525 Cases):
    • Overall Accuracy: 92.38% (485 / 525 cases)
    • Macro F1 Score: 0.9161
    • Micro F1 Score: 0.9238
    • Candidate Pool Recall (E5 Top-3 + unknown): 94.48% (496 / 525 cases)
  • Semantic Parity (PyTorch vs ONNX):
    • Prediction Mismatches: 0 / 525 cases (100.00% agreement)
    • Max Logit Difference: 1.64e-04 (mean: 1.97e-05)
    • Max Probability Difference: 1.06e-05 (mean: 1.26e-07)
    • Order Stability: Identical flip rates across PyTorch, ONNX CPU, and ONNX CUDA.
  • Inference Latency & Throughput:
    • Tesla T4 GPU (CUDAExecutionProvider): 8.94 ms (p50), 10.81 ms (p99) | 111.4 req/s throughput
    • Host CPU (CPUExecutionProvider): 152.00 ms (p50) | 6.2 req/s throughput
    • Memory Footprint: 1.67 GB VRAM (GPU), 7.7 GB RSS (CPU)

Per-Capability Performance Breakdown (525 Frozen Cases)

Evaluated across all 14 enterprise capabilities in the Qefro platform:

Capability Description Support Precision Recall F1 Score Accuracy
orders.list List recent or historical customer orders 47 0.9020 0.9787 0.9388 97.9%
orders.status Check live status/tracking of an order 44 1.0000 0.9091 0.9524 90.9%
orders.cancel Request cancellation of an order 43 0.8511 0.9302 0.8889 93.0%
orders.refund Request refund for an order or item 41 1.0000 0.9268 0.9620 92.7%
catalog.search Search product catalog for items/stock 42 0.9737 0.8810 0.9250 88.1%
customers.search Look up customer profiles and CRM data 38 0.9444 0.8947 0.9189 89.5%
invoices.get Retrieve invoices, tax documents, or receipts 38 1.0000 0.9474 0.9730 94.7%
messages.send Send individual messages or customer notifications 35 0.8919 0.9429 0.9167 94.3%
campaigns.send Trigger bulk marketing/promotional campaigns 35 0.9688 0.8857 0.9254 88.6%
records.delete Delete database entries or customer records 35 1.0000 1.0000 1.0000 100.0%
sales.summary Aggregate analytics, GMV, and revenue metrics 35 0.9722 1.0000 0.9859 100.0%
tickets.create Create customer support / escalation tickets 38 0.8837 1.0000 0.9383 100.0%
ambiguous Mark underspecified or conflicting requests 25 0.9333 0.5600 0.7000 56.0%
unknown Fallback for out-of-domain requests 29 0.6829 0.9655 0.8000 96.6%
Macro Average โ€” 525 0.9288 0.9209 0.9161 โ€”
Micro / Overall โ€” 525 0.9238 0.9238 0.9238 92.38%

Architectural Details: Segment Isolation & Sliding Attention

Decision-Jef employs a ModernBERT backbone with learned decision (q_proj) and option (o_proj) projection heads. The ONNX model natively computes two 4D boolean attention masks:

  1. Segment-Isolated Full Attention: State tokens (Segment 0) are mathematically prevented from cross-attending to Question/Option tokens (Segment 1+). This preserves state contextualization independently of candidate ordering.
  2. Position-ID Sliding Attention: Sliding window layers ($w = 128$, $ ext{half} = 64$) calculate distance using restarted position_ids rather than absolute token indices, ensuring single-question geometric consistency.

Quickstart: Python Inference with ONNX Runtime

1. Install Dependencies

pip install onnxruntime-gpu transformers numpy

2. Run Inference

import numpy as np
import onnxruntime as ort
from transformers import AutoTokenizer

# Load tokenizer and ONNX session
model_id = "developerabu/qefro-decision-jef-c1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
session = ort.InferenceSession("c1_decision_jef.onnx", providers=["CUDAExecutionProvider", "CPUExecutionProvider"])

# Example input query and candidates
user_query = "Please cancel order ORD-99214 immediately"
candidates = ["orders.cancel", "orders.status", "orders.refund", "unknown"]
descriptions = {
    "orders.cancel": "Request cancellation of an order",
    "orders.status": "Check live status/tracking of an order",
    "orders.refund": "Request refund for an order or item",
    "unknown": "Fallback when request is out-of-domain or unsupported"
}

# Format question block
# Segment 0: <s> user query </s>
# Segment 1: <unused2> instructions <unused0> opt0 ... <unused1>
bos_id = tokenizer.bos_token_id or tokenizer.cls_token_id
eos_id = tokenizer.eos_token_id or tokenizer.sep_token_id
opt_id = tokenizer.convert_tokens_to_ids("<unused0>")
dec_id = tokenizer.convert_tokens_to_ids("<unused1>")
q_id = tokenizer.convert_tokens_to_ids("<unused2>")

state_ids = [bos_id] + tokenizer(user_query, add_special_tokens=False)["input_ids"] + [eos_id]
seg_ids = [0] * len(state_ids)
pos_ids = list(range(len(state_ids)))

# Question block
q_base = len(state_ids)
q_text = "choice question: Which Qefro capability should handle this request?"
q_tokens = [q_id] + tokenizer(q_text, add_special_tokens=False)["input_ids"]

opt_positions = []
for cand in candidates:
    opt_positions.append(len(state_ids) + len(q_tokens))
    q_tokens.append(opt_id)
    q_tokens.extend(tokenizer(descriptions[cand], add_special_tokens=False)["input_ids"][:48])

dec_position = len(state_ids) + len(q_tokens)
q_tokens.append(dec_id)

all_ids = state_ids + q_tokens
all_segs = seg_ids + [1] * len(q_tokens)
all_poss = pos_ids + list(range(q_base, q_base + len(q_tokens)))

# Prepare ONNX tensors
ort_inputs = {
    "input_ids":      np.array([all_ids], dtype=np.int64),
    "attention_mask": np.ones((1, len(all_ids)), dtype=np.int64),
    "segment_ids":    np.array([all_segs], dtype=np.int64),
    "position_ids":   np.array([all_poss], dtype=np.int64),
    "batch_idx":      np.array([0], dtype=np.int64),
    "dec_pos":        np.array([dec_position], dtype=np.int64),
    "opt_pos":        np.array([opt_positions], dtype=np.int64),
    "opt_mask":       np.ones((1, len(candidates)), dtype=bool),
}

logits, probabilities = session.run(["logits", "probabilities"], ort_inputs)
best_idx = int(np.argmax(probabilities[0]))
print(f"Selected Capability: {candidates[best_idx]} (Confidence: {probabilities[0][best_idx]:.4f})")

Benchmark Artifacts & Reproducibility

  • c1_best.pt: PyTorch weights checkpoint (Epoch 3, Validation Accuracy: 85.19%, Test Accuracy: 88.89%).
  • c1_decision_jef.onnx: Standalone FP32 ONNX computational graph.
  • c1_decision_jef.onnx.data: External model weights (1.23 GB).
  • evaluation_results_525.json: Full per-case evaluation results across all 525 frozen benchmark instances.
  • capabilities.json: Canonical dictionary of Qefro capability criteria.

Citation & License

Developed and released by Qefro AI for enterprise agent routing.
Licensed under the Apache 2.0 License.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support