Add image_server_kernels 0.4.2 (Windows, cu128/torch2.8)
Browse filesSol-Attn sparse attention, W4A8 linear, in-tree llama.cpp GGUF CUDA kernels (_C_gguf).
Built on Windows 11 / CUDA 12.8 / torch 2.8.0+cu128, MSVC 14.38 toolset.
Verified on RTX 5090 (SM120): smoke_test_rebuild ALL PASS, validate_gguf_kernels ALL PASS on real Q4_K/Q5_K/Q6_K tensors, sol_attn 2.22x vs SageAttention2 at cos 0.9805 on real MiniMax-H3 activations.
sha256 8089f1dacb6ea89634d5f046d6a5e33d266d85b4c4b886199ace5a0b74a16bea
.gitattributes
CHANGED
|
@@ -46,3 +46,4 @@ image_server_kernels-0.3.0-cp311-cp311-win_amd64.whl filter=lfs diff=lfs merge=l
|
|
| 46 |
sageattention-2.2.0+cu128torch2.8.0-cp311-cp311-win_amd64.whl filter=lfs diff=lfs merge=lfs -text
|
| 47 |
image_server_kernels-0.3.0+cu129torch2.8-cp311-cp311-linux_x86_64.whl filter=lfs diff=lfs merge=lfs -text
|
| 48 |
sageattention-2.2.0+cu129torch2.8-cp311-cp311-linux_x86_64.whl filter=lfs diff=lfs merge=lfs -text
|
|
|
|
|
|
| 46 |
sageattention-2.2.0+cu128torch2.8.0-cp311-cp311-win_amd64.whl filter=lfs diff=lfs merge=lfs -text
|
| 47 |
image_server_kernels-0.3.0+cu129torch2.8-cp311-cp311-linux_x86_64.whl filter=lfs diff=lfs merge=lfs -text
|
| 48 |
sageattention-2.2.0+cu129torch2.8-cp311-cp311-linux_x86_64.whl filter=lfs diff=lfs merge=lfs -text
|
| 49 |
+
image_server_kernels-0.4.2+cu128torch2.8-cp311-cp311-win_amd64.whl filter=lfs diff=lfs merge=lfs -text
|
image_server_kernels-0.4.2+cu128torch2.8-cp311-cp311-win_amd64.whl
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:8089f1dacb6ea89634d5f046d6a5e33d266d85b4c4b886199ace5a0b74a16bea
|
| 3 |
+
size 91692698
|