FlexMoE: One-for-All Nested Intra-Expert Pruning for MoE Language Models Paper • 2606.27866 • Published Jun 26 • 1
Llama-Mobile: Efficient 2.7-Bit Quantization of VLMs Paper • 2608.21134 • Published about 1 month ago • 8
Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices Paper • 2607.10183 • Published Jul 14 • 2
BitsMoE: Efficient Spectral Energy-Guided Bit Allocation for MoE LLM Quantization Paper • 2606.00079 • Published May 22 • 1
CCQ: Convolutional Code for Extreme Low-bit Quantization in LLMs Paper • 2507.07145 • Published Jul 9, 2025 • 1
XORTRON - Criminal Computing Collection Release Quality Xortron Apps & Models • 6 items • Updated Apr 6 • 25