Please specify attention mechanism (how KV cache is treated)

#2
by Reverger - opened

Hi, thank you for release!

Please write if you use MLA (Multi-Head Latent Attention) to save up KV cache.

UPD:
Surprisingly, https://arxiv.org/pdf/2607.05471 you attached here doesn't mention attention either.

Sign up or log in to comment