How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf NIM-AI/NIM-2-Coder-7B:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf NIM-AI/NIM-2-Coder-7B:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf NIM-AI/NIM-2-Coder-7B:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf NIM-AI/NIM-2-Coder-7B:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf NIM-AI/NIM-2-Coder-7B:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf NIM-AI/NIM-2-Coder-7B:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf NIM-AI/NIM-2-Coder-7B:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf NIM-AI/NIM-2-Coder-7B:Q4_K_M
Use Docker
docker model run hf.co/NIM-AI/NIM-2-Coder-7B:Q4_K_M
Quick Links

NIM-2 Coder (7B)

NIM-2 Coder is a specialized, high-density 7-billion parameter language model engineered by NIM AI for advanced software engineering, algorithmic design, and full-stack development.

Engineered specifically to punch above its weight class on consumer hardware, NIM-2 Coder delivers complete, type-safe, production-ready code with deep architectural reasoning while running fully locally within 8 GB VRAM.

GitHub License: Apache-2.0


Key Highlights

  • Autonomous Code Synthesis: Writes idiomatic, complete code across Python, TypeScript/JavaScript, Rust, Go, C++, and Bash with zero placeholders.
  • Deterministic Logic & Edge Cases: Trained on multi-stage algorithmic problem decomposition, cyclic graph traversals, and strict type constraints.
  • Hardware Optimized: Packaged in high-throughput Q4_K_M GGUF quantization (~4.6 GB), allowing full offloading to consumer GPUs like the NVIDIA RTX 4060 (8 GB) and Apple Silicon.
  • Agentic Precision: Minimal conversational fluff—outputs immediate technical rationale followed by runnable implementations.

Technical Specifications

Parameter Specification
Model Name NIM-2 Coder
Organization NIM AI (N-I-M-AI)
Architecture Dense Auto-regressive Transformer
Parameters 7.6 Billion
Context Length 4,096 tokens (dynamically extendable)
Format Q4_K_M GGUF (~4.6 GB) / LoRA FP16
Prompt Template ChatML (`<

Quickstart Guide

Run Directly via Ollama (Recommended)

Pull and execute directly from Hugging Face:

ollama run hf.co/N-I-M-AI/NIM-2-Coder-7B:NIM-2-Coder-7B-Q4_K_M.gguf
Downloads last month
-
GGUF
Model size
8B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support