Sleeping Agents Qwen2.5 Vl 3b Instruct Trl Sft ChartQA Trackio 🚀 Track and visualize project metrics
Sleeping Agents Qwen2.5 Vl 3b Instruct Trl Sft ChartQA Trackio 🚀 Track and visualize project metrics
vutran/Llama-3.2-1B-cuadqa-context-finetuned Feature Extraction • 1B • Updated May 12, 2025 • 6
vutran/Llama-3.2-1B-cuadqa-context-finetuned Feature Extraction • 1B • Updated May 12, 2025 • 6
view article Article DeepSeek-R1 Dissection: Understanding PPO & GRPO Without Any Prior Reinforcement Learning Knowledge NormalUhr • Feb 7, 2025 • 297
view article Article Open-R1: a fully open reproduction of DeepSeek-R1 +1 eliebak, lvwerra, lewtun • Jan 28, 2025 • 891
view article Article LLM Comparison/Test: Llama 3 Instruct 70B + 8B HF/GGUF/EXL2 (20 versions tested and compared!) wolfram • Apr 24, 2024 • 64