debugdll commited on
Commit
f5b664d
·
verified ·
1 Parent(s): 2ad4a1a

Upload README.md with huggingface_hub

Browse files
Files changed (1) hide show
  1. README.md +28 -1
README.md CHANGED
@@ -4,12 +4,35 @@ language:
4
  - ru
5
  - en
6
  tags:
7
- - llama.cpp
8
  - gguf
 
 
 
9
  - moe
 
10
  - conversational
 
 
 
 
 
 
 
 
 
 
 
11
  - russian
 
 
 
 
 
 
 
12
  pipeline_tag: text-generation
 
13
  ---
14
 
15
  # Blind Text Models
@@ -42,6 +65,10 @@ Only one model is in the collection for now. New versions will be added to this
42
 
43
  > The model introduces itself as **Blind 1** — that is how it presents itself when asked. This is a build feature.
44
 
 
 
 
 
45
  ## Benchmarks
46
 
47
  Instrumental metrics (MMLU and similar) are still being measured and will be added here. Generation speed is already benchmarked:
 
4
  - ru
5
  - en
6
  tags:
7
+ - text-generation
8
  - gguf
9
+ - llama.cpp
10
+ - llama-cpp-python
11
+ - ollama
12
  - moe
13
+ - mixture-of-experts
14
  - conversational
15
+ - chat
16
+ - assistant
17
+ - instruction-following
18
+ - large-language-model
19
+ - llm
20
+ - quantized
21
+ - mxfp4
22
+ - q8_0
23
+ - 4bit
24
+ - multimodal-text
25
+ - multilingual
26
  - russian
27
+ - english
28
+ - local
29
+ - offline
30
+ - free
31
+ - inference
32
+ - deployment
33
+ - transformers
34
  pipeline_tag: text-generation
35
+ library_name: llama.cpp
36
  ---
37
 
38
  # Blind Text Models
 
65
 
66
  > The model introduces itself as **Blind 1** — that is how it presents itself when asked. This is a build feature.
67
 
68
+ ## Hardware / VRAM
69
+
70
+ Runs fully on GPU in ~11.5 GB — fits comfortably in a 12 GB VRAM card, and easily on 16 GB+. CPU-only inference works too (slower). No external API keys or cloud required — fully local and private.
71
+
72
  ## Benchmarks
73
 
74
  Instrumental metrics (MMLU and similar) are still being measured and will be added here. Generation speed is already benchmarked: