PolicyGuide: From Guarding One Action to Guiding the Whole Workflow for Policy-Compliant LLM Agents Paper • 2608.19861 • Published 2 days ago • 7
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published 4 days ago • 155
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents Paper • 2608.18423 • Published 3 days ago • 16
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published 8 days ago • 277
nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4 Text Generation • 18B • Updated 5 days ago • 535k • 344
Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review Paper • 2608.12440 • Published 10 days ago • 10
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published 21 days ago • 261
DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF Image-Text-to-Text • 9B • Updated about 11 hours ago • 801k • 445