Instructions to use cmes-deepvision/ACR-instance-segmentaiton-refinement-v1.0.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use cmes-deepvision/ACR-instance-segmentaiton-refinement-v1.0.0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-segmentation", model="cmes-deepvision/ACR-instance-segmentaiton-refinement-v1.0.0")# Load model directly from transformers import AutoImageProcessor, OpenVocabMaskPromptMaskDINO processor = AutoImageProcessor.from_pretrained("cmes-deepvision/ACR-instance-segmentaiton-refinement-v1.0.0") model = OpenVocabMaskPromptMaskDINO.from_pretrained("cmes-deepvision/ACR-instance-segmentaiton-refinement-v1.0.0", device_map="auto") - TensorRT
How to use cmes-deepvision/ACR-instance-segmentaiton-refinement-v1.0.0 with TensorRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Refinement-only 1280 TensorRT FP16
exp5 refinement-only 1280 ๋ชจ๋ธ์ ์ต์ข
checkpoint๋ฅผ ONNX๋ก exportํ๊ณ TensorRT
FP16 mixed precision์ผ๋ก ๋ณํํ ๋ฐฐํฌ ํจํค์ง๋ค.
๋ชจ๋ธ
| ํญ๋ชฉ | ๊ฐ |
|---|---|
| source | acr_20260814_exp5_refinement_only_lr1e-4/model/checkpoint-66780 |
| input | pixel_values (1,3,1280,1280), pixel_mask (1,1280,1280) |
| prompts | impossible, possible |
| class embedding | (1,2,128) |
| outputs | logits (1,300,2), pred_boxes (1,300,4), mask_logits (1,300,320,320) |
| precision | TensorRT FP16 mixed precision |
| build GPU | NVIDIA A100-SXM4-80GB |
| TensorRT | 10.16.1.11 |
์์ FP16์ ์ด ๋ชจ๋ธ์ mask_logits์์ NaN์ด ๋ฐ์ํ ์ ์์ด ์ฌ์ฉํ์ง ์์๋ค. ํ์ฌ ์์ง์
8,764๊ฐ ๋ ์ด์ด๋ฅผ FP16 ๋์์ผ๋ก ๋๊ณ ์์น์ ์ผ๋ก ๋ฏผ๊ฐํ 262๊ฐ ๋ ์ด์ด๋ฅผ FP32๋ก ๊ณ ์ ํ๋ค.
์ฐ์ถ๋ฌผ
| ํ์ผ | ํฌ๊ธฐ | ์ฉ๋ |
|---|---|---|
onnx/vision_branch_k2.onnx |
161.2 MB | FP32 ONNX, ๋ค๋ฅธ GPU์ฉ ์์ง ์ฌ๋น๋ ์๋ณธ |
onnx/vision_branch_k2_fp16-mixed.engine |
99.0 MB | A100์ฉ TensorRT FP16 ์์ง |
onnx/vision_branch_k2_fp16-mixed_RTX5080.engine |
100.1 MB | RTX 5080์ฉ TensorRT FP16 ์์ง |
onnx/vision_branch_k2_fp16-mixed_RTX5070.engine |
100.2 MB | RTX 5070์ฉ TensorRT FP16 ์์ง |
class_embeddings_padded/refinement.pt |
4.8 KB | ๋ฐํ์ refinement text embedding |
model_src/model.safetensors |
235.2 MB | ์๋ณธ PyTorch checkpoint |
model_src/conversion_manifest.json |
2.0 KB | shape, precision, ๊ฒ์ฆ ๋ฐ ์์ง ๋ฉํ๋ฐ์ดํฐ |
benchmark_report.json |
3.9 KB | ์๋ยท์ ํ๋ ๊ฒ์ฆ ์๋ณธ ๊ฒฐ๊ณผ |
๊ฒ์ฆ ๊ฒฐ๊ณผ
์ค์ refinement ๊ณ ์ ํ๊ฐ์ 120์ฅ์ ๋ํด ๋์ผํ ์ ์ฒ๋ฆฌ์ ํ์ฒ๋ฆฌ๋ฅผ ์ฌ์ฉํ๋ค. ์ ํ๋ ์งํ๋ score threshold 0.5, mask IoU 0.5์์ ๊ณ์ฐํ๋ค.
| runtime | Precision | Recall | F1 | mask mIoU(TP) | TP / FP / FN |
|---|---|---|---|---|---|
| PyTorch FP32 | 0.8696 | 0.5435 | 0.6689 | 0.8933 | 100 / 15 / 84 |
| ONNX FP32 | 0.8621 | 0.5435 | 0.6667 | 0.8930 | 100 / 16 / 84 |
| TensorRT FP16 mixed | 0.8772 | 0.5435 | 0.6711 | 0.8925 | 100 / 14 / 84 |
FP16 ๋ณํ์ ๋ฐ๋ฅธ ์ต์ข ์ ํ๋ ์์ค์ ๊ด์ธก๋์ง ์์๋ค. FP16์ mask mIoU๋ PyTorch ๋๋น 0.0008 ๋ฎ๊ณ F1์ 0.0022 ๋์ ๋ชจ๋ ์ธก์ ๋ณ๋ ๋ฒ์๋ค. ์์ง ์ถ๋ ฅ์์ NaN/inf๋ ๊ฒ์ถ๋์ง ์์๋ค.
์๋
| runtime | Mean latency | FPS | PyTorch ๋๋น |
|---|---|---|---|
| PyTorch FP32 | 322.54 ms | 3.10 | 1.0x |
| ONNX FP32 | 154.79 ms | 6.46 | 2.08x |
| TensorRT FP16 mixed | 48.14 ms | 20.77 | 6.70x |
์๋ ์ธก์ ๋น์ ๊ฐ์ ์๋ฒ์ ๋ค๋ฅธ ํ์ต ์์ ์ด GPU๋ฅผ ์ ์ ํ๊ณ ์์์ผ๋ฏ๋ก ์ ๋ latency๋ ์ฐธ๊ณ ๊ฐ์ด๋ค. ๊ฐ์ ์กฐ๊ฑด ๋ด ์๋ ๋น๊ต์์๋ TensorRT FP16์ด PyTorch๋ณด๋ค ์ฝ 6.7๋ฐฐ ๋นจ๋๋ค.
RTX 5080 / 5070 ๊ต์ฐจ ๊ฒ์ฆ
๋ PC ๋ชจ๋ ๋์ผํ refinement ๊ณ ์ ํ๊ฐ์ 120์ฅ, ์ ๋ ฅ 1280x1280, batch 1์ ์ฌ์ฉํ๋ค. ์ ํ๋๋ score threshold 0.5์ mask IoU 0.5์์ ์ธก์ ํ์ผ๋ฉฐ, ์๋๋ warmup 10ํ ํ 50ํ ์ถ๋ก ์ ํ๊ท ์ด๋ค. ๊ฐ ์์ง์ ํด๋น PC์์ ONNX๋ก ๋ค์ ๋น๋ํ๋ค.
| ํ๊ฒฝ | runtime | Precision | Recall | F1 | mask mIoU(TP) | Mean latency | FPS |
|---|---|---|---|---|---|---|---|
| RTX 5080 16GB, driver 580.173.02, TRT 10.16.1.11 | PyTorch FP32 | 0.8696 | 0.5435 | 0.6689 | 0.8944 | 60.09 ms | 16.64 |
| RTX 5080 16GB, driver 580.173.02, TRT 10.16.1.11 | TensorRT FP16 mixed | 0.8696 | 0.5435 | 0.6689 | 0.8939 | 18.58 ms | 53.81 |
| RTX 5070 12GB, driver 580.173.02, TRT 10.16.1.11 | PyTorch FP32 | 0.8696 | 0.5435 | 0.6689 | 0.8944 | 90.69 ms | 11.03 |
| RTX 5070 12GB, driver 580.173.02, TRT 10.16.1.11 | TensorRT FP16 mixed | 0.8696 | 0.5435 | 0.6689 | 0.8937 | 28.62 ms | 34.94 |
๋ GPU ๋ชจ๋ TP/FP/FN์ 100/15/84๋ก FP32์ ์์ ํ ๋์ผํ๋ฉฐ F1 ๋ณํ๋ 0์ด๋ค.
FP16์ mIoU ๊ฐ์๋ 5080์์ 0.00056, 5070์์ 0.00074์๋ค. TensorRT ์๋ ํฅ์์
๊ฐ๊ฐ 3.23๋ฐฐ์ 3.17๋ฐฐ๋ค. ๋ฐ๋ผ์ ์ด ํ๊ฐ์
์์๋ ์๋ฏธ ์๋ ์ฑ๋ฅ ์ ํ๊ฐ ๊ด์ธก๋์ง ์์๋ค.
์๋ณธ ์ธก์ ๊ฐ์ benchmark_5080.json๊ณผ benchmark_5070.json์ ๋ณด์กดํ๋ค. 5070์ 12GB
VRAM ์ ์ฝ ๋๋ฌธ์ ONNX FP32 ๋ฐํ์์ ๋์์ ์ ์ฌํ์ง ์๊ณ PyTorch FP32์ TensorRT๋ง ๋น๊ตํ๋ค.
๋ฌด๊ฒฐ์ฑ
cbc90afa1968affcc801515dda7fc3d47be9ec653800255bc28a22fdc5528029 vision_branch_k2.onnx
efea40d7e0506d9b5684c4a415394b7a4c5b3916b3db12c6ea7945ba125b7a6c vision_branch_k2_fp16-mixed.engine
26b93e905f6ba736317d6394b2dc43440a8f11d4c60c70213d5a8522d8d20055 vision_branch_k2_fp16-mixed_RTX5080.engine
05bbf43dbebcfdc528eb045ca4f821dd4c3186b701f6fb0ced16d1d60e0ea056 vision_branch_k2_fp16-mixed_RTX5070.engine
์ฃผ์์ฌํญ
TensorRT ์์ง์ GPU ์ํคํ
์ฒ์ TensorRT ๋ฒ์ ์ ์ข
์๋๋ค. ํ์ฌ .engine์ A100๊ณผ TensorRT
10.16.1.11 ํ๊ฒฝ์์ ๋น๋๋๋ค. ๋ค๋ฅธ GPU์์๋ vision_branch_k2.onnx๋ก ์์ง์ ๋ค์ ๋น๋ํด์ผ
ํ๋ฉฐ, ๊ธฐ์กด A100 ์์ง ํ์ผ์ ๊ทธ๋๋ก ๋ฐฐํฌํ๋ฉด ์ ๋๋ค.
์ฌ๋ณํ๊ณผ ๊ฒ์ฆ์ ์ฌ์ฉํ ๋ช
๋ น ๋ฐ ์ ์ฒด ์ถ๋ ฅ์ ๊ฐ๊ฐ conversion.log, benchmark.log์ ๋ณด์กดํ๋ค.
- Downloads last month
- 16