PureByte 1.0.0 model card for secrets-bin
Browse files
README.md
ADDED
|
@@ -0,0 +1,176 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: other
|
| 3 |
+
license_name: purebyte-model-license
|
| 4 |
+
license_link: https://github.com/purebyte-ai/purebyte/blob/secrets-bin-1.0.0/MODEL_LICENSE.md
|
| 5 |
+
pipeline_tag: token-classification
|
| 6 |
+
tags:
|
| 7 |
+
- secret-detection
|
| 8 |
+
- credentials
|
| 9 |
+
- binary-analysis
|
| 10 |
+
- byte-level
|
| 11 |
+
- state-space-model
|
| 12 |
+
- ternary
|
| 13 |
+
- cpu
|
| 14 |
+
---
|
| 15 |
+
# Model card: secrets-bin
|
| 16 |
+
|
| 17 |
+
> The model file is `purebyte-secrets-bin-1.0.0.gguf` in this repository; its SHA-256 is next to it. It runs with the PureByte runtime ([github.com/purebyte-ai/purebyte](https://github.com/purebyte-ai/purebyte)): `pip install purebyte`, then `purebyte models pull secrets-bin`, which downloads the same file from the runtime's release and checks its checksum.
|
| 18 |
+
|
| 19 |
+
An **AI Specialist** that finds credentials inside binary files: executables and libraries, bytecode, WebAssembly
|
| 20 |
+
modules, firmware images and any other file that a text scanner would skip. It reads the raw bytes, with no
|
| 21 |
+
disassembler, no `strings` pass and no tokenizer, and reports byte offsets.
|
| 22 |
+
|
| 23 |
+
It is one of the three example specialists of PureByte 1.0, which show what one byte-level architecture does on different
|
| 24 |
+
jobs; the architecture, the runtime and the training stack are described in the [README](https://github.com/purebyte-ai/purebyte/blob/secrets-bin-1.0.0/README.md).
|
| 25 |
+
|
| 26 |
+
| | |
|
| 27 |
+
|---|---|
|
| 28 |
+
| Version | 1.0.0 |
|
| 29 |
+
| Task | Credential detection as byte spans in arbitrary binary files |
|
| 30 |
+
| Input | Any file up to 64 MiB by default; zip (JAR, APK...), gzip and tar containers are opened and their members scanned (`--no-archives` scans them as they are); inputs under 24 bytes are not analyzed |
|
| 31 |
+
| Output | Findings: file, byte offsets (also in hex), the surrounding printable string with the credential masked (see [spec/OUTPUT.md](https://github.com/purebyte-ai/purebyte/blob/secrets-bin-1.0.0/spec/OUTPUT.md)) |
|
| 32 |
+
| Size | One file of 30,163,488 bytes (28.8 MiB); 54,170,048 parameters (54.2 M) in the network: the byte embedding, the blocks, the final norm, and the n-gram tables with their projection, 50.3 M of them in the tables, stored at 4 bits per value. The two heads add 34,439 (54,204,487 in all); the per-group scales of the ternary weights are derived, not counted. 12.6 M of the table values can never be read: the format gives every order's table the same number of rows (2^18), and the 65,536 byte bigrams reach only 65,536 rows of the order-2 table, so its other 196,608 rows (7.1 MB of the file, 23 %) are stored but unused; the addressable parameters number 41.6 M |
|
| 33 |
+
| Architecture | Byte-level state-space model with ternary weights (eight `ssm_v2` blocks of width 256), hashed n-gram memory tables (byte 2-, 3- and 4-grams, 2^18 buckets each) held in RAM, and a span-tagging head gated by a window head with max pooling ([docs/architecture.md](https://github.com/purebyte-ai/purebyte/blob/secrets-bin-1.0.0/docs/architecture.md)) |
|
| 34 |
+
| Default decision rule | One model, at the operating point stored in its file (`--bias` moves it); no ensemble |
|
| 35 |
+
| Post-processing | The `secrets-binary` profile: it merges spans, reports byte offsets and the surrounding printable string, masks the values, and drops the spans in the metadata of JAR manifests and signature files (digest lines such as `SHA-256-Digest: <base64>` and entry names, `Name: <path>`). By default the CLI reports only the findings that touch a printable string (16 or more ASCII or UTF-16LE characters); `--all-bytes` reports the others too |
|
| 36 |
+
| Runtime | `purebyte` 1.0.0 or newer, on an x86-64 or arm64 CPU (other little-endian CPUs through the portable build); no GPU, no network |
|
| 37 |
+
| License | [PureByte Model License](https://github.com/purebyte-ai/purebyte/blob/secrets-bin-1.0.0/MODEL_LICENSE.md) (the code of the runtime is Apache-2.0) |
|
| 38 |
+
| Files | `purebyte-secrets-bin-1.0.0.gguf`; checksum in [models.json](https://github.com/purebyte-ai/purebyte/blob/secrets-bin-1.0.0/models/models.json) |
|
| 39 |
+
| Recipe | [specialists/secrets-bin](https://github.com/purebyte-ai/purebyte-train/blob/main/specialists/secrets-bin/README.md) in purebyte-train: the recipe and the evaluation plan, to rebuild or improve it |
|
| 40 |
+
|
| 41 |
+
## Intended use
|
| 42 |
+
|
| 43 |
+
- Checking build artifacts, containers' binaries, mobile and desktop packages and firmware before they ship.
|
| 44 |
+
- Auditing third-party binaries you are authorized to assess.
|
| 45 |
+
- Triage: pointing a human at the few byte ranges worth a look in megabytes of machine code.
|
| 46 |
+
|
| 47 |
+
## Out of scope
|
| 48 |
+
|
| 49 |
+
- **Proving that a binary is free of secrets.** Treat findings as leads and absence of findings as weak evidence.
|
| 50 |
+
- **Encrypted or compressed content.** Bytes inside compressed streams are not readable text; zip, gzip and tar
|
| 51 |
+
containers are opened by the runtime (see [docs/cli.md](https://github.com/purebyte-ai/purebyte/blob/secrets-bin-1.0.0/docs/cli.md)), other compression is not.
|
| 52 |
+
- **Checking whether a credential is live.** It never contacts any service.
|
| 53 |
+
- **Text files.** Use [secrets-code](https://github.com/purebyte-ai/purebyte/blob/secrets-bin-1.0.0/models/secrets-code.md), which is smaller and more precise on text.
|
| 54 |
+
|
| 55 |
+
## Evaluation
|
| 56 |
+
|
| 57 |
+
There is no public, labeled benchmark of credentials inside compiled binaries yet. These figures come from our own
|
| 58 |
+
exams, run on the released file with the `purebyte` CLI and its default options. **They cannot be reproduced as they
|
| 59 |
+
are**: part of the binaries are system files and vendor libraries that cannot be redistributed, and the probe is cut
|
| 60 |
+
from our internal corpus. The files that come from public packages are pinned in
|
| 61 |
+
[data/exams/binary-fp.yaml](https://github.com/purebyte-ai/purebyte-train/blob/main/data/exams/binary-fp.yaml) of
|
| 62 |
+
purebyte-train.
|
| 63 |
+
|
| 64 |
+
**False alarms.** Binaries that hold no credential, so every finding is a false alarm. `--all-bytes` also reports
|
| 65 |
+
the findings that touch no printable string; its column was measured on the files that had findings in an earlier
|
| 66 |
+
run that reported every finding (the other files had none there):
|
| 67 |
+
|
| 68 |
+
| Exam | Files | Findings | With `--all-bytes` |
|
| 69 |
+
|---|---|---:|---:|
|
| 70 |
+
| bin-fp-a: 24 compiled Python extension modules (scipy 1.18.1 and scikit-learn 1.9.1 wheels, Linux x86-64; public) | 24 (38 MiB) | 1 | 1 |
|
| 71 |
+
| bin-fp-b: 25 Linux system libraries and executables (Ubuntu 26.04 packages, PyTorch and NumPy wheels; public) | 25 (93 MiB) | 9, all in 4 of the 8 files that overlap the training corpus; 0 in the other 17 | the same 9 |
|
| 72 |
+
| bin-fp-b: 15 Windows DLLs (system, DirectX, MFC, a graphics driver; not redistributable) | 15 (80 MiB) | 12, in 3 files | the same 12 |
|
| 73 |
+
| bin-fp-c: 69 files in 14 formats absent from the training corpus of the first binary models (JAR, CLASS, PYC, WebAssembly, APK, firmware, Mach-O, ELF for ARM and AVR, object files, ar archives, PE for ARM64, Windows executables, Node.js addons; partly not redistributable) | 69 (109 MiB) | 4, in 2 class files inside JARs; 0 with `--no-archives` | 7; 1 with `--no-archives` |
|
| 74 |
+
| Held-out extension modules: 60 compiled extension modules of 18 Python packages, set apart before training (local builds) | 60 (34 MiB) | 2, in 1 file | the same 2 |
|
| 75 |
+
|
| 76 |
+
The 25 Linux files include `/usr/bin/perl`, a hard link to `perl5.40.1` (one file). bin-fp-c was measured on the
|
| 77 |
+
copies of two Node.js addons that desktop applications bundle; on 2026-09-25 the exam manifest was corrected to pin
|
| 78 |
+
the files npm publishes for the same package versions, and this model reports nothing on either version of them.
|
| 79 |
+
|
| 80 |
+
By default the CLI opens zip containers and scans their members: on bin-fp-c, 13,921 members of 6 JAR and 3 APK
|
| 81 |
+
files. The 4 findings are in two class files (guava's list of public suffixes, and Tomcat); `--all-bytes` adds 2 in
|
| 82 |
+
Scala class files and, outside the containers, 1 in a lookup table of a WebAssembly module. The manifests and
|
| 83 |
+
signature files of two signed JARs (`META-INF/*.MF`, `*.SF`) list a base64 digest for every class, and the model
|
| 84 |
+
flags many of those lines: 373 findings (371 of them on `SHA-256-Digest:` lines) before the profile had a rule for
|
| 85 |
+
them. The profile now drops every span in the digest lines and entry names of those files, with or without
|
| 86 |
+
`--all-bytes`.
|
| 87 |
+
|
| 88 |
+
**Recall on planted credentials (a synthetic probe).** 2,000 windows of 512 bytes cut from real binaries of our
|
| 89 |
+
corpus: 1,200 left untouched, 400 with a planted credential (`NAME=value`, the value alone among other strings,
|
| 90 |
+
`--flag=value`, embedded JSON) and 400 look-alike traps (hashes, UUIDs, public keys, placeholders). The probe is
|
| 91 |
+
synthetic and in-distribution: its credential values, names and traps come from the generator the model was trained
|
| 92 |
+
with, in the layouts of its binary mode, and its background windows from an earlier corpus of binaries from the same
|
| 93 |
+
machine folders as the training corpus. It shows that the model learned what its generator makes, not how it does on
|
| 94 |
+
real leaks, and the rival was not built for the generator's formats. Of the 400 credentials, 360 can be told apart
|
| 95 |
+
from an identifier by looking at them. Each window is scanned by the CLI as a 512-byte file, with the default options
|
| 96 |
+
(`--all-bytes` gives the same figures):
|
| 97 |
+
|
| 98 |
+
| | secrets-bin 1.0.0 | `strings` + regular expressions |
|
| 99 |
+
|---|---:|---:|
|
| 100 |
+
| Recall on the 360 decidable credentials | **1.000** | 0.803 |
|
| 101 |
+
| Precision (traps flagged count as false alarms) | 0.989 | 0.876 |
|
| 102 |
+
| Findings on the 1,200 untouched windows | 0 | 0 |
|
| 103 |
+
|
| 104 |
+
The remaining 40 are generic values with no name next to them, which nothing can decide from the bytes alone: the
|
| 105 |
+
model flags 1 of them. 360 of 360 is a recall of 1.000 with a 95 % Clopper-Pearson interval of 0.990 to 1.000; over
|
| 106 |
+
the recipe's six seeds, 0.989 to 1.000 (two seeds found 356 of the 360).
|
| 107 |
+
|
| 108 |
+
**The recipe's six seeds.** The recipe trains six models from different random seeds and releases the best one on its
|
| 109 |
+
own validation data, before any exam is read. Measured with the research code (the same weights, n-gram tables at 32
|
| 110 |
+
bits instead of 4), mean and maximum over the six seeds: bin-fp-a 2.7 and 6; bin-fp-b Windows DLLs 13.2 and 27;
|
| 111 |
+
bin-fp-b Linux files, without the files that overlap the training corpus, 3.5 and 8; bin-fp-c, idem, 13.0 and 32;
|
| 112 |
+
held-out extension modules 0.8 and 2; probe recall on decidable credentials, lowest seed 0.989. The released seed
|
| 113 |
+
scored 2, 12, 0, 0, 2 and 1.000 there. The research code reports every span and scans containers as they are, like
|
| 114 |
+
`--all-bytes --no-archives`, and storing the tables at 4 bits changes the spans of about 1 % of windows: that is why
|
| 115 |
+
the released file's `--all-bytes` counts above differ by one finding here and there.
|
| 116 |
+
|
| 117 |
+
**Pre-registered criteria.** Every false-alarm criterion of the recipe passes. The probe criterion, a recall of at
|
| 118 |
+
least 0.99 for every seed, fails: two of the six seeds reach 0.989 (4 of the 360 missed), and the released seed 1.000.
|
| 119 |
+
The model was released with that criterion failed
|
| 120 |
+
([details](https://github.com/purebyte-ai/purebyte/blob/secrets-bin-1.0.0/benchmarks/RESULTS.md#the-pre-registered-criteria-of-secrets-bin)).
|
| 121 |
+
|
| 122 |
+
## Known limitations
|
| 123 |
+
|
| 124 |
+
- **False alarms depend on the file format.** Data regions that look structured (bitmaps, lookup tables of
|
| 125 |
+
WebAssembly modules or firmware, runs of a repeated byte, compressed bytes) can produce findings that are not
|
| 126 |
+
credentials. Formats that are rare in the wild are the most exposed.
|
| 127 |
+
- **Credentials need readable bytes.** A credential split across non-contiguous bytes, encoded, encrypted or inside a
|
| 128 |
+
compressed stream the runtime does not open is not found; generic values with no name next to them are mostly
|
| 129 |
+
missed (they cannot be told apart from identifiers).
|
| 130 |
+
- **Short strings are left out by default.** A finding must touch a printable string of 16 or more characters, so a
|
| 131 |
+
credential stored in a shorter string, with no longer string around it, is reported only with `--all-bytes`. The
|
| 132 |
+
probe plants every credential in a longer string, so it does not measure this cost.
|
| 133 |
+
- **Digest lists are handled by a rule, not by the model.** The manifests and signature files of signed JARs list a
|
| 134 |
+
base64 digest per class, and the model flags many of those lines (373 findings in two signed JARs of bin-fp-c
|
| 135 |
+
without the rule). The profile drops every span in their digest lines and entry names, with or without
|
| 136 |
+
`--all-bytes`; lists of digests in other formats have no such rule. Containers are opened by default; with
|
| 137 |
+
`--no-archives` their compressed bytes are scanned as they are, which also misses any credential inside the members.
|
| 138 |
+
- **The optional string prefilter skips windows without printable text** (runs of at least 16 ASCII or UTF-16
|
| 139 |
+
characters). It speeds up real binaries several times; what it can miss lies outside printable strings, which the
|
| 140 |
+
default rule leaves out anyway (`--all-bytes` does not). It is off by default.
|
| 141 |
+
- **Confidence values are not calibrated yet.**
|
| 142 |
+
- **Throughput is modest** (about 65 KB/s on a 12-core desktop, [docs/performance.md](https://github.com/purebyte-ai/purebyte/blob/secrets-bin-1.0.0/docs/performance.md)): scan
|
| 143 |
+
release artifacts, not whole disks.
|
| 144 |
+
|
| 145 |
+
## Training data
|
| 146 |
+
|
| 147 |
+
- **Generated examples in real binaries.** Windows of real binaries into which the generator of
|
| 148 |
+
[secrets-code](https://github.com/purebyte-ai/purebyte/blob/secrets-bin-1.0.0/models/secrets-code.md) injects credentials and look-alike negatives in binary layouts (between NUL bytes
|
| 149 |
+
like a string table, the value alone, long hex digests as negatives). For the released weights, the binaries came
|
| 150 |
+
from the Linux system folders of the training machine (`/usr/bin`, `/usr/lib`, `/opt` and the like); the data
|
| 151 |
+
scripts rebuild a comparable corpus from public packages pinned by version and SHA-256, so a retraining lands close,
|
| 152 |
+
not bit for bit.
|
| 153 |
+
- **Real negatives.** 7 windows of every training step (45 % of the nominal batch of 16, added to it: 7 of the 23
|
| 154 |
+
windows of a step) come from 78,283 windows of binaries without credentials: Windows executables and libraries of
|
| 155 |
+
the system folders, files of twelve other formats (JAR, APK, PYC, WebAssembly, Mach-O, Node addons, firmware, ar
|
| 156 |
+
archives, object files, ELF shared libraries, and Windows executables and DLLs from outside the system folders) and
|
| 157 |
+
compiled extension modules of 63 Python packages. None of them is an exam file; many are system files that cannot
|
| 158 |
+
be redistributed, so the builder of this stream is not in purebyte-train yet.
|
| 159 |
+
- **Overlap.** Parts of 10 exam files (8 of the 25 Linux files of bin-fp-b and 2 files of bin-fp-c) are also in the
|
| 160 |
+
binary corpus the generated examples were cut from; the six-seed figures above leave them out.
|
| 161 |
+
|
| 162 |
+
The recipe and the scripts that rebuild the corpus from public sources are in the training repository,
|
| 163 |
+
[purebyte-train](https://github.com/purebyte-ai/purebyte-train):
|
| 164 |
+
[specialists/secrets-bin](https://github.com/purebyte-ai/purebyte-train/blob/main/specialists/secrets-bin/README.md) and
|
| 165 |
+
[data/](https://github.com/purebyte-ai/purebyte-train/blob/main/data/README.md).
|
| 166 |
+
|
| 167 |
+
## Responsible use
|
| 168 |
+
|
| 169 |
+
Scan binaries you own or are authorized to assess. Findings are masked by default and credentials are never verified
|
| 170 |
+
against any service. If you find a live credential in someone else's software, report it to its vendor privately.
|
| 171 |
+
|
| 172 |
+
## Versions
|
| 173 |
+
|
| 174 |
+
| Version | Date | Changes |
|
| 175 |
+
|---|---|---|
|
| 176 |
+
| 1.0.0 | 2026-09-24 | First public release |
|