Acknowledge the Responsible Use Agreement to access this repository

Access is granted automatically after you agree to the terms below and submit the form.

Responsible Use Agreement

This model has had safety refusals removed. That makes it useful for red-teaming, security research, evaluation, and unfiltered assistant tasks — and also removes guardrails a user must therefore supply themselves.

Prohibited uses (you must agree before access is granted):

  • Anything involving the sexual exploitation or endangerment of minors.
  • You must be of age 18 years or older to use and download this model.
  • You agree any information generated that can cause harm in terms of generating recipe, knowledge to make any materials/substances is your own input and responsibility. You will be accountable for any harm/damage caused by your action/input.
  • Content promoting self-harm or suicide.
  • Generation of material that is illegal in your jurisdiction, or that targets real individuals for harassment, doxxing, or fraud.
  • Any use prohibited by the upstream DeepSeek license.

You are responsible for adding appropriate safety filtering, human review, and access controls for your deployment. The weights are provided as-is, with no warranty. The license is inherited from the upstream DeepSeek base model — review and comply with it before use or redistribution.

Log in or Sign Up to review the conditions and access this model content.

keys-DeepSeekV4Flash-Vision-EXP-ablit

Abliterated DeepSeek-V4-Flash-Vision-Exp. Public repo, gated with automatic approval after you agree to RESPONSIBLE_USE.md.

0731-style wo_b projection on L10–35 (λ=3.5, 26 tensors). L0–9, L36–42, MTP/DSpark draft, and the full vision tower stay stock.

HF https://huggingface.co/drowzeys/keys-DeepSeekV4Flash-Vision-EXP-ablit
Parent deepseek-ai/DeepSeek-V4-Flash-Vision-Exp
0731 ancestor (same ablit idea) keys-DeepSeekV4-Flash-GA-0731-Dspark-Abliterated-Anchored-Tensors
Ablit L10–35 attn.wo_b · λ=3.5 · mean Δrel ≈ 0.057 · MTP not edited
Anchors L0–9 stock · L36–42 stock · vision / aligner / image markers stock
Layout 48 shards, FP8 block-scale (quant_method=fp8), ~157 GiB
This fleet (text serve) TP2 · DSpark k=6 · ctx 262144 · util 0.835 · native VL skipped by vLLM (tensors still in the files)

See METHOD.md. Access is gated with automatic approval after you agree.


What was edited

Keys 0731 lesson: smash early layers or the in-checkpoint drafter and DSpark accept dies. Vision-Exp adds a vision tower on top of 0731 text. This dest:

  • does project residual wo_b on L10–35
  • does not edit MTP / DSpark draft tensors
  • does not touch vision.*, aligner.*, image markers, or ffn.gate.bias_vl

Artifacts: ABLIT_META.json.


Vision vs this vLLM serve

The weights include the Vision-Exp tower (~309 vision tensors). Current Anemll DSpark vLLM (DeepseekV4ForCausalLM) is text-only and skips those tensors at load, including on the DSpark drafter (shared checkpoint, no vision params). Sending an image to that serve returns is not a multimodal model.

Native pixels on this checkpoint need a multimodal vLLM class for DeepSeek-V4-Flash-Vision, not a different model (Aeon / Qwen3.8 Omni is a separate stack). The designed workaround on the DSpark recipe is an optional Qwen3-VL sidecar (ENABLE_VL_SIDECAR=1, :8889), which does not use this tower.


Download

# after you agree to the gate (automatic approval)
hf download drowzeys/keys-DeepSeekV4Flash-Vision-EXP-ablit \
  --local-dir ~/models/keys-DeepSeekV4Flash-Vision-EXP-ablit

GPU memory utilization ≤ 0.85.


License

Inherited from DeepSeek-V4-Flash-Vision-Exp. You must still comply with the Responsible Use gate above.

Downloads last month
-
Safetensors
Model size
305B params
Tensor type
BF16
·
F32
·
F8_E4M3
·
I8
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for drowzeys/keys-DeepSeekV4Flash-Vision-EXP-ablit

Finetuned
(2)
this model
Quantizations
1 model