KBBridge-v3 / README.md
KBBridge's picture
Reasoning: reemplazar 'apagalo siempre' por la medicion por tarea, y corregir el numero de MCQ que estaba mal
82ac291 verified
|
Raw
History Blame Contribute Delete
13.2 kB
---
license: apache-2.0
base_model: Qwen/Qwen3.8-27B
base_model_relation: finetune
tags: [genexus, code-generation, qwen3]
language: [es, en]
pipeline_tag: text-generation
---
# KBBridge-v3 (bf16)
A fine-tune of [`Qwen/Qwen3.8-27B`](https://huggingface.co/Qwen/Qwen3.8-27B) specialised in
**GeneXus** programming, in the native `.gxSource` export format.
Frontier models do not know this format. Without the GeneXus documentation injected into the
prompt they produce syntactically invalid output almost every time (parse rate 0.5–3.1%).
KBBridge writes it natively, runs on your own hardware, and never sends your Knowledge Base
code to an external API.
---
## ⚠️ Read this before your first prompt
Three settings. All three are measured on this model, not stylistic.
### 1. Reasoning: it depends on the task, and the difference is large
The chat template no longer pins a value, so **reasoning follows the upstream Qwen default
(on)** unless your client turns it off. Which one you want depends on what you are asking for.
All three rows below are measured on this model, not inherited from Qwen.
| Task | Reasoning | Measured |
|---|---|---|
| **Writing `.gxSource`** | **OFF** | parseRate **86.9 → 33.5**, parmMatch 80.2 → 50.4 (191 held-out objects) |
| Documentation multiple-choice | either | 78.4 vs 78.4 — no difference (329 items, McNemar p = 1.000) |
| **Explaining existing code** | **ON** | fabricated claims **15.4% → 8.7%** (149 items, McNemar p = 0.041) |
**If you generate GeneXus objects, turn reasoning off.** The collapse is real, not a budget
artifact: with reasoning on, only 5.2% of items hit the token ceiling (fewer than the 8.4%
without it) and 3.7% came back empty. The model simply writes worse `.gxSource` when it
reasons first. Writing `.gxSource` is a formatting task.
**If you point the model at existing code and ask what it does, turn reasoning on.** It nearly
halves the rate at which the model asserts things the source does not support — the failure
mode that matters when the output is documentation someone will trust. Cost: ~5× the output
tokens.
> **Correction (2026-09-02).** An earlier version of this card reported MCQ dropping 78.1 → 69.6
> with reasoning on. That number was wrong: our benchmark harness capped multiple-choice answers
> at 512 tokens, which is not enough for a reasoning block to close, so the run was measuring the
> cap rather than the model. Re-measured with an adequate budget, the difference is zero. The
> `.gxSource` degradation is real and reproduced above with the current scorer.
#### How to turn it off
```bash
# vLLM — pass it explicitly on every request
curl .../v1/chat/completions -d '{
"model": "...", "messages": [...],
"chat_template_kwargs": {"enable_thinking": false}
}'
```
> **Serve it with the flag, or you will think the model is broken.** If you run vLLM with
> `--reasoning-parser qwen3` and the request does **not** carry `enable_thinking`, the parser
> assumes reasoning is on, never finds the closing `</think>`, and routes the **entire answer**
> into `reasoning`, leaving `content: null`. Every standard OpenAI client then shows an empty
> reply. This only affects non-streaming requests — streaming takes a different path in vLLM and
> looks fine — which makes it doubly confusing. Either pass `enable_thinking` on every request,
> or drop `--reasoning-parser` and let the tags through.
To pin reasoning off for every client instead, add this as the first line of
`chat_template.jinja`:
```jinja
{%- set enable_thinking = false %}
```
A `set` at the top of the template overrides anything the caller passes, which is a blunt but
reliable way to guarantee behaviour across runtimes.
### 2. Ask for the format explicitly
Write **"in `.gxSource` format"** in your prompt.
Measured on v3: the bare request *"a Procedure that adds two numbers"* returns generic **SQL**.
Naming the format returns the GeneXus object, consistently. If you use a harness with its own
system prompt, put the instruction there once.
### 3. Give it enough room
`max_tokens` ≥ 4096. A `.gxSource` object consumes roughly **340 tokens per KB** of source, and
most tools default to 512–1024, which truncates the object mid-body.
---
## Results
580 held-out items (191 codegen + 329 MCQ + 60 data-model) that no model saw during training.
Syntax validated with the official GeneXus ANTLR parser. Same protocol for every model:
temperature 0.1, reasoning off, concurrency 8.
### v3 vs v2 — an honest comparison
**v3 is not a clean win over v2.** It gains domain knowledge and loses syntax accuracy:
| Metric | v2 | **v3** | |
|---|---|---|---|
| parseRate (valid syntax) | **89.0** | 84.8 | −4.2 |
| parmMatch (exact signature) | 78.6 | 78.6 | = |
| MCQ (GeneXus knowledge) | 76.0 | **79.0** | +3.0 |
| methodValidity | 90.0 | **91.1** | +1.1 |
**What these numbers do NOT establish.** v3 changed three things at once — the base model
(Qwen3.6 → 3.8), the corpus (4× larger, per-KB cap removed) and the teacher (v1 → v2). The
parseRate drop **cannot be attributed** to any one of them without a control arm that was never
run. Anyone reading this table as "the bigger corpus hurt syntax" is over-reading it.
Choose v3 if domain knowledge matters more to you; v2 still leads on raw syntax validity.
### Generalisation to unseen Knowledge Bases
Three entire KBs were held out — different domains, never in the pipeline:
| | held-out from training KBs | 3 completely new KBs |
|---|---|---|
| v2 | 89.0 | 89.9 |
| **v3** | 84.8 | **87.4** |
v3's *relative* gap to unseen KBs is larger than v2's (+2.6 vs +0.9), i.e. it generalises
better in relative terms, even though two KBs make up 54.7% of its corpus.
### Fairness note on the frontier comparison
In our benchmark the frontier models were run **with** ~21,600 tokens of GeneXus documentation
injected into every request; KBBridge was run **without** any. That is not a handicap we
imposed — injecting the same documentation into KBBridge makes it *worse* (76.4 → 73.3
parseRate), because the fine-tune already internalised that knowledge and the extra context
gets in the way. Still, the setups differ, and you should know that when reading any
head-to-head number.
### Quantised builds
We measured the 4-bit build against this one on the same 580 items. **Excluding items where
either run hit the token ceiling, the two are indistinguishable** (parseRate 93.0 vs 93.6 over
171 items) — 4-bit costs essentially nothing in output quality here. Details and the full
comparison are in the
[GGUF repo's card](https://huggingface.co/KBBridge/KBBridge-v3-GGUF).
---
## Files
Full-precision merged weights, bf16, **51 GB** across 19 shards. This is the master artefact:
use it to re-quantise, to continue training, or to serve with transformers.
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
m = AutoModelForCausalLM.from_pretrained("KBBridge/KBBridge-v3", dtype="bfloat16", device_map="auto")
t = AutoTokenizer.from_pretrained("KBBridge/KBBridge-v3")
```
For serving, prefer [`KBBridge/KBBridge-v3-FP8`](https://huggingface.co/KBBridge/KBBridge-v3-FP8)
(29 GB, same quality in our tests) or the
[GGUF builds](https://huggingface.co/KBBridge/KBBridge-v3-GGUF) for llama.cpp / LM Studio.
### What is inside
1,199 tensors: the 64-layer hybrid text model (48 Gated DeltaNet + 16 full-attention layers),
the base model's **vision tower** (333 tensors, carried over unchanged — the fine-tune did not
touch it) and its **multi-token-prediction head** (15 tensors, likewise unchanged). Context
262,144 tokens, the base model's native `max_position_embeddings`.
## Intended use
Assisting GeneXus developers: generating objects (Procedures, Transactions, Data Providers,
SDTs, WebPanels), explaining existing code, completion, and documentation questions.
**Out of scope:** not a general-purpose model, not a replacement for validating in the GeneXus
IDE, and it does not know any particular Knowledge Base (see *Limitations*).
---
## Limitations
- **It does not know your KB.** It learned the style and syntax of the format, not the contents
of any specific base. Ask it about a transaction you did not paste in, and it will **invent
plausible attribute names and present them as fact**. Always give it the context and validate
the output in the IDE.
- **Runaway generation on very large objects.** For objects over ~10 KB the model can fall into
degenerate repetition — the same line hundreds of times without closing the object. Measured
on v2 at ~1.6% of benchmark items; **not re-measured on v3**. Raising `max_tokens` does not
fix it. Generate large objects section by section.
- **Spanish bias** in explanations, reflecting the corpus.
- **Specialised**: worse than the base model at general tasks.
- The limitations above other than the first were measured on **v2** and are carried over as
working assumptions, not verified properties of v3.
### If you also use a hosted KBBridge endpoint
The raw GGUF and a gateway-fronted deployment **do not behave the same by default**. Our
gateway applies five corrections the plain model does not have: a `max_tokens` floor,
reasoning off unless the client asks for it, `temperature` defaulted to 0.2 (without it vLLM
falls back to the checkpoint's `generation_config`, which is **1.0**), a fallback that recovers
the answer from the `reasoning` field when `content` comes back empty, and `repetition_penalty`
1.05 to suppress runaway. If you compare "what I tried on your server" against "what I
downloaded", the difference is those five settings, not the weights.
The temperature one surprises people: the OpenAI standard makes the field optional and many
clients never send it, so an unconfigured client is sampling at 1.0 without being told.
---
## Training
| | |
|---|---|
| Method | QLoRA 4-bit (bitsandbytes) + Liger kernel |
| LoRA | r=64, α=128, dropout=0.05, all projections |
| Context | 12,288 tokens |
| Effective batch | 16 (1 × 16 grad accum) |
| LR | 1.0e-4, cosine, 3% warmup |
| Epochs | 2 complete (14,108 steps) |
| Hardware | 1× RTX PRO 6000 Blackwell 96 GB |
| Duration | 7 days 4:41 |
| Framework | LLaMA-Factory, transformers 5.6.0 |
train_loss **0.2618** (v2: 0.3344) · eval_loss **0.3723** (v2: 0.4675), minimum at the **last**
step — no overfitting across 71 evaluations, which suggests there was room for more epochs.
Note that these losses are much better than v2's and yet parseRate went *down*: `eval_loss`
measures fit to the corpus, not GeneXus quality.
### Data
80,344 examples derived from GeneXus objects across 25 real Knowledge Bases (GX16/17/17U8/18/
Evo1, multi-domain) — 129% more than v2, with the per-KB cap removed. Sanitised, deduplicated
and split by deterministic hash. **The datasets are not published**: they contain customer
proprietary code.
---
## Training-data privacy
The model was trained on real customer Knowledge Bases, so we audited whether it can leak them.
This is the strongest result of the project.
### Canaries: no memorisation threshold found
12 synthetic objects containing unguessable 16-character secrets were inserted at four
frequencies, and verified to have reached `train.jsonl` at exactly those counts:
| repetitions | canaries | recovered by name | recovered with literal prefix |
|---|---|---|---|
| 1 | 3 | 0/3 | 0/3 |
| 10 | 3 | 0/3 | 0/3 |
| 100 | 3 | 0/3 | 0/3 |
| **1000** | 3 | **0/3** | **0/3** |
**Not even at a thousand identical repetitions.** A control rules out a broken probe: asked for
the canary, the model returns a structurally valid but **empty** object — no token, no secret.
And it does generate real bodies when the request has content, so the empty skeleton is not an
inability to generate.
### Membership inference: marginal signal
| | |
|---|---|
| mean loss, seen examples | 3.4130 |
| mean loss, unseen | 3.7711 |
| mean length | 3,133 vs 3,117 chars — comparable, so the AUC is meaningful |
| **AUC** | **0.5539** |
0.554 against 0.50 for indistinguishable. There is a statistical trace of having seen the data,
but the distributions overlap almost entirely.
**Conclusion: customer code is not recoverable from the weights.**
**Caveat, stated plainly:** absence of evidence is not proof of absence. These audits cover the
attacks we ran, not every attack that exists.
---
## Reproducibility
Full external reproduction is **not possible**, and it is worth saying so directly:
1. The 25 Knowledge Bases are customer code and are not distributed.
2. The `parseRate` scorer uses the KBEditor's ANTLR parser — proprietary, not distributable.
3. The teacher that generated v3's data is KBBridge-v2, which is not published.
What a third party *can* verify: the raw benchmark outputs (one model response per item) and
the scoring over them.
---
## Citation
```bibtex
@misc{kbbridge-v3,
title = {KBBridge-v3: a GeneXus code assistant fine-tuned from Qwen3.8-27B},
author = {{KBBridge}},
year = {2026},
url = {https://huggingface.co/KBBridge/KBBridge-v3}
}
```
## License
Apache 2.0, inherited from the base model `Qwen/Qwen3.8-27B`. This is a modified derivative
work; see `NOTICE`.