coding-agents
tool-selection
bialy / README.md
paintedwolf-ai's picture
Release open1-b5-b7g-e4
49119f9 verified
|
Raw History Blame Contribute Delete
3.05 kB
metadata
license: apache-2.0
base_model: convaiinnovations/laya-multilingual
datasets:
  - paintedwolfcode/bialy-dataset
tags:
  - coding-agents
  - tool-selection

Bialy decision heads open1-b5-b7g-e4

Tuned heads built on Laya by Convai Innovations, over its frozen multilingual encoder that Painted Wolf Code asks as it works: which loadable tool schemas a turn will need, which instruction units it can leave out, and what kind of work it is (turn-load); how relevant each skill and tool card is to a request (unit-rank); and which code units best answer a request, blended into the text-match order of code summaries, repository maps, and project search (code-rank). The engine (bialy) loads them beside the backbone; a head trained over another backbone is refused.

Trained on paintedwolfcode/bialy-dataset open1-b7g-e4: turn-load and unit-rank on its session rows, labeled by what open-weights models did in coding-agent sessions and scored by an open-weights judge; code-rank on its code-rank pairs, requests open-weights models wrote for code units in the same repositories. Trained and replayed with Painted Wolf Code at commit 68e19e34e7b85a459a2a007af52f3611d6996089; packaged with https://github.com/paintedwolf-ai/bialy.

Heads

  • code-rank.safetensors: open1-code-rank, over jhu-clsp/mmBERT-base, sha256 abe235c347a223e957a0f8b7b979e3a584aebab76ee239e5f62117e839b9c40a
  • guide-load.safetensors: open1-turn-load-B7G-release-independent, over jhu-clsp/mmBERT-base, sha256 e0305aa8b62ddc144774785f2e0d3710604af65cd21015662428870f8d34b76e
  • turn-load.safetensors: open1-turn-load-B5-release-independent, over jhu-clsp/mmBERT-base, sha256 550f30b94d5c6a8680d2cb2a607ebc9ee33e709c7679f72606d1e9fda17f2e90
  • unit-rank.safetensors: open1-unit-rank-e4-dense1, over jhu-clsp/mmBERT-base, sha256 c62674993eabba94660eccfafbb6e04d1549adc7335656c0ddefa7f068fa58f7

Results

Replayed through the shipped engine on sets neither head trained on. Tool columns are precision / recall / F1 of the loadable tools a turn used, at the catalog's load threshold; guide omission precision is the share of omitted instruction units the turn did not need; need MRR ranks the tools a request_tools need went on to use.

Set Heads tools P/R/F1 macro R loads/turn guide omit P kind acc need MRR

code-rank, on repositories it never trained on: the rank of the code unit a request was written for, in the site's text-match order and blended with the head.

Report Pairs MRR text / blended hit@1 text / blended improved / regressed

Check it yourself

Each row of the table comes from a replay report shipped in eval/. bialy audit heads (from https://github.com/paintedwolf-ai/bialy) replays these heads through the Painted Wolf Code engine on a CPU over the dataset's held-out split and compares every metric with the shipped report.