codegeist-llm / README.md
codegeist's picture
Record independent CPU reload test
d1d21f9 verified
|
Raw
History Blame
6.34 kB
metadata
base_model: Qwen/Qwen3-1.7B
base_model_relation: adapter
library_name: peft
pipeline_tag: text-generation
inference: false
language:
  - en
license: other
license_name: 0bsd
license_link: https://github.com/codegeist-ai/codegeist-ai/blob/main/LICENSE
tags:
  - peft
  - lora
  - sft
  - transformers
  - unsloth
  - non-production
  - identity-smoke

Codegeist Qwen3-1.7B Identity Smoke Adapter

This is a non-production LoRA adapter created to validate the Codegeist training pipeline. It teaches one response only:

User: What is Codegeist?
Assistant: Codegeist is a coding agent.

It is not evidence of coding ability, reasoning, generalization, safe tool use, Codegeist OS integration, GGUF conversion, Vulkan deployment, or production model quality.

Artifact Identity

Field Value
Base model Qwen/Qwen3-1.7B
Base revision 70d244cc86ccca08cf5af4e1e306ecf908b1ad5e
Adapter format PEFT LoRA, Safetensors
Adapter weight SHA-256 19d424106ef88ffeac4c26c22cebfb13ae1d5f309e1dcccf2da708727bec10a8
Training Job 6a75f25a3e1f34a7e32bd646
Training date 2026-08-07

evidence.json contains the sanitized run chronology, configuration, package versions, hashes, cost estimate, verification status, and known gaps. It does not contain model weights, private logs, or credentials.

Intended Use

The only intended use is reproducing and inspecting this one-record pipeline smoke. Use the immutable base revision above and pin this adapter repository to a specific Hub commit when loading it.

Do not use this adapter as a coding assistant, autonomous agent, general chat model, safety component, or production model. It was not evaluated for those purposes.

Loading

This CPU-compatible example prioritizes portability over speed. Replace ADAPTER_REVISION with an immutable commit from this repository:

import torch
from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

BASE_MODEL = "Qwen/Qwen3-1.7B"
BASE_REVISION = "70d244cc86ccca08cf5af4e1e306ecf908b1ad5e"
ADAPTER_MODEL = "codegeist/qwen3-1.7b-codegeist-identity-smoke"
ADAPTER_REVISION = "04d51edac56c6f1e068c644bfa8d014cadcecf9f"

tokenizer = AutoTokenizer.from_pretrained(
    BASE_MODEL,
    revision=BASE_REVISION,
    trust_remote_code=False,
)
base_model = AutoModelForCausalLM.from_pretrained(
    BASE_MODEL,
    revision=BASE_REVISION,
    trust_remote_code=False,
    torch_dtype=torch.float32,
    low_cpu_mem_usage=True,
)
model = PeftModel.from_pretrained(
    base_model,
    ADAPTER_MODEL,
    revision=ADAPTER_REVISION,
    is_trainable=False,
)

prompt = tokenizer.apply_chat_template(
    [{"role": "user", "content": "What is Codegeist?"}],
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=False,
)
inputs = tokenizer(prompt, return_tensors="pt", add_special_tokens=False)
with torch.inference_mode():
    output = model.generate(
        **inputs,
        do_sample=False,
        temperature=None,
        top_p=None,
        top_k=None,
        max_new_tokens=64,
        pad_token_id=tokenizer.eos_token_id,
        eos_token_id=tokenizer.eos_token_id,
    )

response = tokenizer.decode(
    output[0, inputs["input_ids"].shape[1]:],
    skip_special_tokens=True,
).strip()
print(response)

Expected whitespace-normalized response:

Codegeist is a coding agent.

Training Data

The complete project-authored synthetic dataset is one public record:

{
  "instruction": "What is Codegeist?",
  "response": "Codegeist is a coding agent."
}

The record ID is codegeist-identity-v1-001. It contains no private data, personal information, or credentials. Training and evaluation deliberately use the same record to test memorization; there is no held-out evaluation set.

Training

  • Python 3.12
  • PyTorch 2.6.0 with CUDA 12.4
  • Unsloth 2026.8.7
  • Transformers 5.5.0
  • TRL 0.24.0
  • PEFT 0.20.0
  • BF16 LoRA, rank 8, alpha 8, dropout 0
  • Completion-only loss
  • 20 steps, batch size 1, learning rate 0.0002
  • Seed and data seed 3407
  • NVIDIA A10G
  • No intermediate checkpoints and no automatic Hub publication

The aggregate training loss was 1.6867698234826094. The final logged step loss was approximately 0.0003.

Evaluation

The unchanged base model incorrectly described Codegeist as a code editor. After training, the adapter was loaded onto a fresh instance of the exact base revision in a child process. One greedy generation produced the expected answer after leading and trailing whitespace normalization.

The raw decoded continuation before .strip() was not retained. Training and inference repeatability, deterministic PyTorch algorithms, coding benchmarks, safety evaluation, and generalization were not tested.

Before public release, an independent local CPU reload used PyTorch 2.6.0+cpu, Transformers 5.5.0, PEFT 0.20.0, the immutable base revision, and adapter commit 04d51edac56c6f1e068c644bfa8d014cadcecf9f. In that publication test, both the raw and whitespace-normalized responses were exactly Codegeist is a coding agent.. This confirms public-artifact loading and the single memorized response only; it does not broaden the interpretation boundary.

Licenses And Provenance

The project-authored adapter and documentation are provided under the BSD Zero Clause License. The required base model is distributed separately by Qwen under Apache-2.0. This repository does not redistribute base-model weights. Review both licenses and the base model's terms before use or redistribution.

See THIRD_PARTY_NOTICES.md for the exact upstream model reference. The Codegeist source repository is codegeist-ai/codegeist-llm.

Publication Limitations

  • The successful training source was not committed when the paid Job launched; exact source bytes are anchored by SHA-256 in evidence.json.
  • Downloaded model and tokenizer cache bytes were not independently rehashed inside the Job against the upstream manifest.
  • The generated adapter configuration originally omitted the base revision; the publication copy sets it to the immutable revision used by the Job.
  • This publication does not change the experiment's non-production status.