Unifying Symbolic and Audio Music Generation at Frontier Quality. Rivals Suno v5; covers & agentic editing. MERT2, SheetSage2, VAEs, WildSongBench.
AI & ML interests
None defined yet.
Recent Activity
Papers
OProver: A Unified Framework for Agentic Formal Theorem Proving
A Self-Evolving Framework for Efficient Terminal Agents via Observational Context Compression
Organization Card
Multimodal Art Projection (M-A-P) is an open-source AI research community.
The community members are working on research topics in a wide range of spectrum, including but not limited to pre-training paradigm of foundation models, large-scale data collection and processing, and the derived applciations on coding, reasoning and music creativity.
The community is open to researchers keen on any relevant topic. Welcome to join us!
- Discord Channel
- Our Full Paper List
- mail: contact@m-a-p.ai
The development log of our Multimodal Art Projection (m-a-p) model family:
- π₯10/09/2026: We release YuE2, an open-source frontier music generation model that rivals Suno v5, with support for cover song generation and agentic editing. Explore the models and benchmark and listen to the demos.
- π₯28/01/2025: We release YuE (δΉ), the most powerful open-source foundation models for music generation, specifically for transforming lyrics into full songs (lyrics2song), like Suno.ai. See demos.
- π₯08/05/2024: We release the fully transparent large language model MAP-Neo, series models for scaling law exploraltion and post-training alignment, and along with the training corpus Matrix.
- π₯11/04/2024: MuPT paper and demo are out. HF collection.
- π₯08/04/2024: Chinese Tiny LLM is out. HF collection.
- π₯28/02/2024: The release of ChatMusician's demo, code, model, data, and benchmark. π
- π₯23/02/2024: The release of OpenCodeInterpreter, beats GPT-4 code interpreter on HumanEval.
- 23/01/2024: we release CMMMU for better Chinese LMMs' Evaluation.
- 13/01/2024: we release a series of Music Pretrained Transformer (MuPT) checkpoints, with size up to 1.3B and 8192 context length. Our models are LLAMA2-based, pre-trained on world's largest 10B tokens symbolic music dataset (ABC notation format). We currently support Megatron-LM format and will release huggingface checkpoints soon.
- 02/06/2023: officially release the MERT pre-print paper and training codes.
- 17/03/2023: we release two advanced music understanding models, MERT-v1-95M and MERT-v1-330M , trained with new paradigm and dataset. They outperform the previous models and can better generalize to more tasks.
- 14/03/2023: we retrained the MERT-v0 model with open-source-only music dataset MERT-v0-public
- 29/12/2022: a music understanding model MERT-v0 trained with MLM paradigm, which performs better at downstream tasks.
- 29/10/2022: a pre-trained MIR model music2vec trained with BYOL paradigm.
models 218
m-a-p/YuE2-3B
Text-to-Audio β’ 4B β’ Updated β’ 81 β’ 61
m-a-p/SheetSage2
Feature Extraction β’ 57.2M β’ Updated β’ 41 β’ 12
m-a-p/MERT-v2-30s
Feature Extraction β’ 0.6B β’ Updated β’ 17 β’ 6
m-a-p/MERT-v2-FullSong
Feature Extraction β’ 0.6B β’ Updated β’ 52 β’ 7
m-a-p/YuE2-Vae-legacy
Feature Extraction β’ 0.1B β’ Updated β’ 16 β’ 3
m-a-p/YuE2-Vae
Feature Extraction β’ 0.1B β’ Updated β’ 74 β’ 3
m-a-p/OProver-8B
Text Generation β’ 8B β’ Updated β’ 259 β’ 1
m-a-p/OProver-32B
Text Generation β’ 33B β’ Updated β’ 104 β’ 4
m-a-p/OProver-32B-Round1
Text Generation β’ 33B β’ Updated β’ 11
m-a-p/OProver-8B-Round2
Text Generation β’ 8B β’ Updated β’ 11
datasets 77
m-a-p/WildSongBench
Viewer β’ Updated β’ 192 β’ 21 β’ 4
m-a-p/LPFQA
Viewer β’ Updated β’ 430 β’ 121 β’ 6
m-a-p/TerminalTraj-5k-instances
Updated β’ 127 β’ 6
m-a-p/MSQA
Updated β’ 53 β’ 2
m-a-p/bins
Updated
m-a-p/ClinPath
Updated β’ 4
m-a-p/OProofs
Viewer β’ Updated β’ 6.8M β’ 5.44k β’ 4
m-a-p/PIN-200M
Viewer β’ Updated β’ 68.1k β’ 163k β’ 24
m-a-p/TerminalTraj
Viewer β’ Updated β’ 20k β’ 252 β’ 10
m-a-p/Encyclo-K
Viewer β’ Updated β’ 5.04k β’ 67 β’ 4