DMM: Decentralized Master-Mind

DMM is a decentralized multi-agent pathfinding policy that refines agents' action intents over several local communication rounds before committing to actions. The models are pretrained on expert solutions with imitation learning and optionally fine-tuned with MICPO, a critic-free reinforcement learning method.

Paper · Code

Checkpoint Imitation pretraining iterations MICPO optimizer updates
DMM-08M.pt 1,000,000 —
DMM-3M.pt 1,000,000 —
DMM-MICPO-08M.pt 1,000,000 96,000
DMM-MICPO-3M.pt 1,000,000 96,000

MICPO fine-tuning comprises 500 outer iterations (96,000 optimizer updates). Both model sizes use four communication rounds. Training details are reported in the paper.

See the GitHub repository for code, usage instructions, training, and evaluation.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for tviskaron/DMM