Qwen2.5-3B checkpoints for Defects4J unit-test generation with Online Policy Distillation and GRPO.
tomsawyer
tomhu
·
AI & ML interests
None yet
Recent Activity
updated a collection about 18 hours ago
RL4TG updated a collection about 18 hours ago
RL4TG updated a collection about 18 hours ago
RL4TGOrganizations
None yet
models 45
tomhu/RL4TG-Qwen3-8B-Offline-SFT-GRPO
Text Generation • 8B • Updated
tomhu/RL4TG-Qwen3-8B-OPD-GRPO
Text Generation • 8B • Updated
tomhu/RL4TG-Qwen3-8B-Direct-GRPO
Text Generation • 8B • Updated
tomhu/RL4TG-DeepSeek-Coder-1.3B-Coder6.7B-Offline-SFT-GRPO
Text Generation • 1B • Updated • 29
tomhu/RL4TG-DeepSeek-Coder-1.3B-GRPO
Text Generation • 1B • Updated • 35
tomhu/RL4TG-DeepSeek-Coder-1.3B-OPD6.7B-GRPO
Text Generation • 1B • Updated • 28
tomhu/RL4TG-DeepSeek-Coder-1.3B-OPD-33B-Teacher
Text Generation • 1B • Updated • 16
tomhu/RL4TG-DeepSeek-Coder-1.3B-OPD-6.7B-Teacher
Text Generation • 1B • Updated • 16
tomhu/RL4TG-DeepSeek-Coder-1.3B-Official-D4J-SFT
Text Generation • 1B • Updated • 16
tomhu/RL4TG-DeepSeek-Coder-1.3B-Coder6.7B-Offline-SFT
Text Generation • 0.7B • Updated • 29