File size: 4,028 Bytes
78362e3
 
 
 
 
60dba78
 
 
 
 
 
 
78362e3
60dba78
 
78362e3
 
 
 
 
799e302
78362e3
b1abaaa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
78362e3
b1abaaa
78362e3
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1190a43
78362e3
 
 
 
 
 
 
 
 
 
b1abaaa
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
---
license: apache-2.0
base_model: Qwen/Qwen3-VL-4B-Instruct
pipeline_tag: image-text-to-text
tags:
- document-understanding
- information-extraction
- vision-language
- qwen3-vl
- structured-data-extraction
- multimodal
- document-ai
library_name: transformers
language:
- en
---
# obj_v1
Vision-language model fine-tuned for structured data extraction from Indian
financial documents. Give it a page image and a JSON schema; it returns the
schema filled in from what is on the page.
A 4B vision-language model, LoRA fine-tuned and merged

## Authors

<p align="left">
  <!-- <a href="https://www.linkedin.com/in/ahmedzaweel/">
    <img src="https://img.shields.io/badge/LinkedIn-Ahmed%20Zaweel-0A66C2?style=for-the-badge&logo=linkedin&logoColor=white" alt="Ahmed Zaweel on LinkedIn" />
  </a> -->
  &nbsp;&nbsp;
  <a href="https://www.linkedin.com/in/rachit-kumar-b41299228/">
    <img src="https://img.shields.io/badge/LinkedIn-Rachit%20Kumar-0A66C2?style=for-the-badge&logo=linkedin&logoColor=white" alt="Rachit Kumar on LinkedIn" />
  </a>
  &nbsp;&nbsp;
  <a href="https://www.linkedin.com/in/ahmedzaweel/">
    <img src="https://img.shields.io/badge/LinkedIn-Ahmed%20Zaweel-0A66C2?style=for-the-badge&logo=linkedin&logoColor=white" alt="Ahmed Zaweel on LinkedIn" />
  </a>
</p>

## Serving with vLLM
```bash
vllm serve objectai/obj_v1 \
  --served-model-name obj_v1 \
  --max-model-len 16384 \
  --limit-mm-per-prompt '{"image":1}' \
  --mm-processor-kwargs '{"max_pixels":1003520}' \
  --trust-remote-code
```
`max_pixels` is 1280x28x28, the resolution the model was trained at. Raising it
wastes KV cache; lowering it makes small print unreadable.
## Calling it
The server is OpenAI-compatible, so an ordinary chat completion works:
```python
import base64, json, openai
client = openai.OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
image = base64.b64encode(open("cheque.jpg", "rb").read()).decode()
schema = {"cheque_details": {"amount": "number", "payee": "string",
                             "date": "string", "cheque_number": "string"}}
response = client.chat.completions.create(
    model="obj_v1",
    temperature=0.0,
    max_tokens=8192,
    messages=[
        {"role": "system", "content":
            "You are a document data extraction model. "
            "Extract only values present in the document. "
            "Use null for fields that are absent or illegible. "
            "Output a single compact JSON object matching the requested schema. "
            "No prose, no markdown, no explanation."},
        {"role": "user", "content": [
            {"type": "image_url",
             "image_url": {"url": f"data:image/jpeg;base64,{image}"}},
            {"type": "text",
             "text": f"document_type: cheque\nschema: {json.dumps(schema)}"},
        ]},
    ],
)
print(response.choices[0].message.content)
```
## Prompt format
Match training or accuracy drops. The system prompt above is verbatim, and the
user turn is the image followed by exactly two lines:
```
document_type: <type>
schema: <compact json>
```
Set `temperature=0.0` so the same page yields the same answer.
## Requirements
| | |
| --- | --- |
| Weights | 8.9 GB (bf16) |
| VRAM | 16 GB minimum, 24 GB comfortable |
| Precision | bf16 (Ampere or newer; use fp16 below that) |
| Context | 16384 covers the longest documents |
Runs on an L4, A10G, L40S, A100 or RTX 4090. On a T4 add `--dtype float16`.

## Output
Compact JSON matching the requested schema. Fields absent from the page come
back `null` rather than guessed. Values found on the page that the schema did
not ask for are placed under `extras` when that key is included in the schema.
## Limitations
- Trained on Indian financial documents; other domains and layouts are untested.
- Handwriting is the weakest case, particularly digits at low resolution.
- The model does not verify its own arithmetic. Totals that must reconcile
  should be checked by the caller.
## License
Apache 2.0. Fine-tuned from Qwen3-VL-4B-Instruct, which is Apache 2.0.