How to use from
Docker Model Runner
docker model run hf.co/prithivMLmods/Infinity-Parser2-Flash-GGUF:
Quick Links

Infinity-Parser2-Flash-GGUF

Infinity-Parser2-Flash is a low-latency document understanding model from infly-ai, one of two variants in the Infinity-Parser2 flagship family (alongside the accuracy-optimized Infinity-Parser2-Pro), engineered for fast inference while consolidating robust multi-modal parsing into a unified architecture trained via an upgraded synthetic data engine spanning nearly 5 million diverse document samples and a novel multi-task reinforcement learning approach with verifiable rewards across document parsing, element parsing, chart parsing, chemical formula parsing, document VQA, and general multimodal understanding. It delivers a 3.68x speedup over the previous Infinity-Parser-7B model (increasing throughput from 441 to 1,624 tokens/sec) while still posting strong benchmark results — 86.0% on olmOCR-Bench, 72.2% on ParseBench, and 91.98% on OmniDocBench-v1.6 — outperforming frontier models like DeepSeek-OCR-2 and MinerU2.5 on several document-parsing tasks, though trailing its larger Pro sibling on layout analysis, chart/chemical formula parsing, and general multimodal benchmarks (e.g., MMMU, AI2D, MathVista). It extracts structured layout with bounding boxes, category labels, and per-element text (LaTeX for formulas, HTML for tables, Markdown for text), supports command-line and Python API usage via the infinity_parser2 package with vLLM, transformers, or vLLM-server backends, and is released under Apache-2.0, with known limitations primarily around English/Chinese-only support, degraded accuracy on complex charts and rotated table elements, and no fine-grained text formatting (bold, italic, strikethrough) capture.

Model Files

File Name Quant Type File Size File Link
Infinity-Parser2-Flash.BF16.gguf BF16 3.78 GB Download
Infinity-Parser2-Flash.F16.gguf F16 3.78 GB Download
Infinity-Parser2-Flash.F32.gguf F32 7.54 GB Download
Infinity-Parser2-Flash.Q2_K.gguf Q2_K 969 MB Download
Infinity-Parser2-Flash.Q3_K_L.gguf Q3_K_L 1.16 GB Download
Infinity-Parser2-Flash.Q3_K_M.gguf Q3_K_M 1.1 GB Download
Infinity-Parser2-Flash.Q3_K_S.gguf Q3_K_S 1.02 GB Download
Infinity-Parser2-Flash.Q4_0.gguf Q4_0 1.2 GB Download
Infinity-Parser2-Flash.Q4_K_M.gguf Q4_K_M 1.27 GB Download
Infinity-Parser2-Flash.Q4_K_S.gguf Q4_K_S 1.21 GB Download
Infinity-Parser2-Flash.Q5_0.gguf Q5_0 1.37 GB Download
Infinity-Parser2-Flash.Q5_K_M.gguf Q5_K_M 1.41 GB Download
Infinity-Parser2-Flash.Q5_K_S.gguf Q5_K_S 1.37 GB Download
Infinity-Parser2-Flash.Q6_K.gguf Q6_K 1.56 GB Download
Infinity-Parser2-Flash.Q8_0.gguf Q8_0 2.01 GB Download
Infinity-Parser2-Flash.mmproj-bf16.gguf mmproj-bf16 671 MB Download
Infinity-Parser2-Flash.mmproj-f16.gguf mmproj-f16 671 MB Download
Infinity-Parser2-Flash.mmproj-f32.gguf mmproj-f32 1.33 GB Download
Infinity-Parser2-Flash.mmproj-q8_0.gguf mmproj-q8_0 365 MB Download

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
14
GGUF
Model size
2B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

2-bit

3-bit

4-bit

5-bit

6-bit

8-bit

16-bit

32-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/Infinity-Parser2-Flash-GGUF

Quantized
(4)
this model

Collection including prithivMLmods/Infinity-Parser2-Flash-GGUF