| --- |
| language: |
| - fa |
| license: apache-2.0 |
| library_name: coreml |
| pipeline_tag: automatic-speech-recognition |
| base_model: |
| - Reza2kn/Shenava-Rizeh-v1.0 |
| base_model_relation: quantized |
| tags: |
| - coreml |
| - neuralnetwork |
| - ios |
| - ios15 |
| - ios14 |
| - persian |
| - farsi |
| - asr |
| - fastconformer |
| - streaming |
| - ctc |
| - fp16 |
| - on-device |
| - shenava |
| - shenava-1 |
| - visualears |
| --- |
| |
| # 🎙️ Shenava-Rizeh-v1.0-CoreML-iOS15-fp16 |
|
|
| > **English + فارسی** · Part of [Shenava 1.0](https://huggingface.co/collections/Reza2kn/shenava-10-open-streaming-persian-asr-and-captioning) · [Project hub](https://github.com/Reza2kn/shenava-1) · [SLT paper submission](https://openreview.net/forum?id=QTa6ax9PU3) |
|
|
| ## 🌟 At a glance | معرفی سریع |
|
|
| | | English | فارسی | |
| |---|---|---| |
| | 🎯 Role | Rizeh iOS 15 CoreML export. | مدل دانشآموز فشردهٔ شنوا ریزه؛ این مخزن یکی از مصنوعات رسمی خانوادهٔ Shenava-1 است. | |
| | 🧠 Family | Shenava Rizeh compact student | مدل دانشآموز فشردهٔ شنوا ریزه | |
| | 📦 Format | Core ML FP16 deployment export | خروجی FP16 برای Core ML | |
| | 📐 Scale | 32M parameters | اندازه: 32M parameters | |
| | 📥 Input | mono Persian speech resampled to 16 kHz | گفتار تککانالهٔ فارسی با نرخ نمونهبرداری ۱۶ کیلوهرتز | |
| | 📤 Output | Persian transcription; normalization and ITN belong in the display layer | رونویسی فارسی؛ نرمالسازی و تبدیل عدد گفتاری در لایهٔ نمایش انجام میشود | |
| | ⚖️ License | Apache License 2.0 | مجوز Apache 2.0 | |
|
|
| ## 🇬🇧 English documentation |
|
|
| ### 🧭 Overview |
|
|
| Rizeh iOS 15 CoreML export. This repository is an official Shenava-1 release artifact, not an isolated checkpoint. It belongs to a Persian-first stack covering training data, streaming ASR, semantic evaluation, on-device exports, captioning applications, and reproducible benchmarks. Use the collection link above to locate sibling model sizes, deployment formats, datasets, and evaluation assets. |
|
|
| The artifact is optimized for Persian speech and the conventions used by the Shenava/VisualEars pipeline. A model file alone is not the entire inference system: audio preparation, tokenizer assets, streaming state, decoding, Persian text normalization, and inverse text normalization can materially affect observed output. |
|
|
| ### ✅ Intended uses |
|
|
| - Persian ASR research, benchmarking, and reproducible comparison inside the Shenava-1 evaluation protocol. |
| - Offline or streaming transcription when the selected runtime and graph support that mode. |
| - On-device captioning, accessibility prototypes, and Persian speech interfaces. |
| - Conversion or runtime integration work that preserves the source model’s tokenizer, decoding assumptions, and numerical checks. |
|
|
| ### 🚫 Out-of-scope or unsafe uses |
|
|
| - Do not treat transcripts as guaranteed verbatim records for legal, medical, emergency, or other high-stakes decisions. |
| - Do not infer identity, health, ethnicity, intent, or other sensitive traits from speech or model errors. |
| - Do not compare formats using different text normalization, test subsets, or decoding settings and present the result as model quality. |
| - Do not assume robustness to every Persian accent, code-switching pattern, recording channel, or adversarial acoustic condition. |
|
|
| ### 📁 Repository contents |
|
|
| This snapshot contains **9 files** totaling approximately **52.19 MB**. Common file groups: `.json` × 4, `no extension` × 2, `.md` × 1, `.py` × 1, `.mlmodel` × 1. |
|
|
| Largest or representative artifacts: |
|
|
| - `shenava_rizeh_v1_0_ctc_streaming_att70_0_ios15_fp16.mlmodel` |
| - `mel_filters_slaney_80x257.json` |
| - `tokens.json` |
| - `export_koochik10_streaming_coreml.py` |
| - `preprocessor.json` |
|
|
| The repository card and `LICENSE` are part of the release. Runtime-specific configuration, tokenizer, vocabulary, metadata, and state files should be kept beside the main weights when present. |
|
|
| ### 🚀 Download and integration |
|
|
| ```python |
| from huggingface_hub import snapshot_download |
| |
| local_dir = snapshot_download( |
| repo_id="Reza2kn/Shenava-Rizeh-v1.0-CoreML-iOS15-fp16", |
| local_dir="./Shenava-Rizeh-v1.0-CoreML-iOS15-fp16", |
| ) |
| print(local_dir) |
| ``` |
|
|
| Use the runtime named by the artifact format. Inspect the exported graph signature before binding input and output tensors; deployment exports may expose cache/state tensors in addition to acoustic features. |
|
|
| For NeMo checkpoints, restore through `nemo.collections.asr.models.ASRModel.restore_from(...)` rather than assuming a CTC-only class. For converted artifacts, follow the graph metadata and the runtime-specific notes retained later in this card. Validate one known clip against the source checkpoint before shipping a conversion. |
|
|
| ### 📏 Evaluation |
|
|
| Report at least WER and CER using the same Persian normalization rules, plus S³ when semantic importance matters. Shenava’s public [Triple Threat leaderboard](https://huggingface.co/spaces/Reza2kn/PersianASR-TrippleThreat) combines the Golden6669 and FLEURS-fa splits. Record the exact repository revision, decoder settings, chunk/context configuration, precision, device, and normalization code. |
|
|
| Deployment exports should be checked for numerical and transcription parity against their parent repository, [ `Reza2kn/Shenava-Rizeh-v1.0` ](https://huggingface.co/Reza2kn/Shenava-Rizeh-v1.0). Runtime speed is hardware-specific; publish latency, real-time factor, warm-up policy, thread count, and audio duration together. |
|
|
| ### ⚠️ Limitations and responsible use |
|
|
| ASR quality varies with accent, age, speaking style, background noise, distance, clipping, reverberation, telephony bandwidth, overlapping speech, and code-switching. Persian orthography also permits multiple acceptable written forms. WER or CER can therefore penalize a semantically correct alternative, while a low aggregate score can still hide loss of a critical word. Review meaning-critical outputs and expose uncertainty in accessibility-facing products. |
|
|
| ### 🔁 Reproducibility checklist |
|
|
| 1. Pin the Hub revision and runtime/library versions. |
| 2. Resample audio deterministically and document channel mixing. |
| 3. Keep tokenizer and decoding assets from this repository together. |
| 4. Record streaming chunk, left/right context, cache reset, and endpointing behavior. |
| 5. Apply one documented Persian normalization/ITN pipeline to references and hypotheses. |
| 6. Publish failed cases and condition-level results, not only a single average. |
|
|
| ## 🇮🇷 مستندات فارسی |
|
|
| ### 🧭 معرفی |
|
|
| مدل دانشآموز فشردهٔ شنوا ریزه است. این مخزن یک مصنوع رسمی از انتشار Shenava-1 است و باید همراه با دادههای آموزشی، توکنایزر، روش رمزگشایی، نرمالسازی فارسی و تنظیمات اجرای جریانی دیده شود. پیوند مجموعه در بالای صفحه، نسخههای همخانواده، قالبهای استقرار، دادهها و معیارهای ارزیابی را یکجا نشان میدهد. |
|
|
| هدف پروژه ارائهٔ زیرساخت باز و قابل بازتولید برای بازشناسی گفتار و زیرنویس فارسی است. نتیجهٔ نهایی فقط به وزن مدل وابسته نیست؛ نرخ نمونهبرداری، کانال صوت، وضعیت کش، روش رمزگشایی، تبدیل اعداد گفتاری و یکسانسازی نیمفاصله نیز بر خروجی اثر دارند. |
|
|
| ### ✅ کاربردهای پیشنهادی |
|
|
| - پژوهش، بنچمارک و مقایسهٔ منصفانهٔ ASR فارسی با پروتکل یکسان. |
| - رونویسی آفلاین یا جریانی، در صورتی که قالب و زماناجرای انتخابی از آن پشتیبانی کند. |
| - زیرنویس روی دستگاه، ابزارهای دسترسپذیری و رابطهای گفتاری فارسی. |
| - تبدیل مدل و یکپارچهسازی با زماناجراهای مختلف همراه با آزمون برابری خروجی. |
|
|
| ### 🚫 کاربردهای نامناسب |
|
|
| - خروجی را در تصمیمهای پزشکی، حقوقی، اضطراری یا پرخطر بهعنوان سند قطعی به کار نبرید. |
| - از خطا یا صدای کاربر برای استنباط هویت، سلامت، قومیت، نیت یا ویژگی حساس استفاده نکنید. |
| - نتایجی را که با زیرمجموعه، نرمالسازی یا رمزگشایی متفاوت ساخته شدهاند مقایسهٔ مستقیم ننامید. |
| - پوشش کامل همهٔ لهجهها، گفتار آمیخته، کانالها و شرایط صوتی را فرض نکنید. |
|
|
| ### 📁 محتوای مخزن |
|
|
| این نسخه شامل **9 فایل** با حجم تقریبی **52.19 MB** است. گروههای رایج فایل: `.json` × 4, `no extension` × 2, `.md` × 1, `.py` × 1, `.mlmodel` × 1. |
|
|
| فایلهای شاخص: |
|
|
| - `shenava_rizeh_v1_0_ctc_streaming_att70_0_ios15_fp16.mlmodel` |
| - `mel_filters_slaney_80x257.json` |
| - `tokens.json` |
| - `export_koochik10_streaming_coreml.py` |
| - `preprocessor.json` |
|
|
| فایلهای توکنایزر، واژگان، پیکربندی، وضعیت جریانی و فراداده را در صورت وجود کنار وزن اصلی نگه دارید. |
|
|
| ### 🚀 دریافت و استفاده |
|
|
| ابتدا snapshot کامل مخزن را دریافت کنید، سپس از زماناجرای متناسب با قالب استفاده کنید. پیش از اتصال ورودی و خروجی، امضای گراف را بررسی کنید؛ خروجیهای جریانی ممکن است علاوه بر ویژگی صوتی، تنسورهای وضعیت و کش داشته باشند. |
|
|
| برای چکپوینت NeMo از `ASRModel.restore_from(...)` استفاده کنید و مدل را صرفاً CTC فرض نکنید. برای خروجیهای تبدیلشده، یک کلیپ مرجع را با مدل مبدأ مقایسه کنید و سپس استقرار را انجام دهید. |
|
|
| ### 📏 ارزیابی |
|
|
| حداقل WER و CER را با نرمالسازی فارسی یکسان گزارش کنید و در سناریوهای حساس به معنا، S³ را نیز بیاورید. در [جدول Triple Threat](https://huggingface.co/spaces/Reza2kn/PersianASR-TrippleThreat) دو بخش Golden6669 و FLEURS-fa با هم سنجیده میشوند. شناسهٔ دقیق نسخه، تنظیمات دیکودر، کانتکست، دقت عددی، سختافزار و کد نرمالسازی را ثبت کنید. |
|
|
| ### ⚠️ محدودیتها و استفادهٔ مسئولانه |
|
|
| لهجه، سن، سبک گفتار، نویز، فاصله، کلیپشدن، پژواک، کانال تلفنی، همپوشانی گویندگان و کدسوئیچینگ میتوانند کیفیت را تغییر دهند. چند نگارش فارسی ممکن است از نظر معنایی درست باشند، اما WER/CER یکی را خطا حساب کند. در محصولات دسترسپذیری، واژههای کلیدی را جداگانه بازبینی و عدم قطعیت را به کاربر نشان دهید. |
|
|
| ### 🔁 چکلیست بازتولید |
|
|
| ۱. نسخهٔ دقیق مخزن و کتابخانهها را ثابت کنید. ۲. تبدیل نرخ نمونه و کانال را مستند کنید. ۳. توکنایزر و داراییهای رمزگشایی همین مخزن را نگه دارید. ۴. اندازهٔ قطعه، کانتکست، بازنشانی کش و endpointing را ثبت کنید. ۵. یک خط لولهٔ نرمالسازی/ITN مشترک به مرجع و خروجی اعمال کنید. ۶. خطاهای نمونهای و نتایج هر شرایط را در کنار میانگین منتشر کنید. |
|
|
| ## 📚 Citation, links, and license | استناد، پیوندها و مجوز |
|
|
| - 🤗 [Shenava-1 collection](https://huggingface.co/collections/Reza2kn/shenava-10-open-streaming-persian-asr-and-captioning) |
| - 🧰 [Project repository](https://github.com/Reza2kn/shenava-1) |
| - 📄 [SLT paper submission](https://openreview.net/forum?id=QTa6ax9PU3) |
| - 📊 [Persian ASR Triple Threat](https://huggingface.co/spaces/Reza2kn/PersianASR-TrippleThreat) |
|
|
| ```bibtex |
| @misc{shenava1_shenava_rizeh_v1_0_coreml_ios15_fp16, |
| title = {Shenava-Rizeh-v1.0-CoreML-iOS15-fp16: a Shenava-1 Persian speech artifact}, |
| author = {Reza2kn}, |
| year = {2026}, |
| url = {https://huggingface.co/Reza2kn/Shenava-Rizeh-v1.0-CoreML-iOS15-fp16} |
| } |
| ``` |
|
|
| Released under the **Apache License 2.0**. این مخزن با **مجوز Apache 2.0** منتشر شده است. |
|
|
| --- |
|
|
| ## 📎 Retained technical notes | یادداشتهای فنی پیشین |
|
|
| The pre-existing technical card is retained below for revision-specific commands, measurements, and artifact details. The bilingual sections above define the common Shenava-1 documentation contract. |
|
|
| یادداشت فنی قبلی برای فرمانها، اندازهگیریها و جزئیات همان نسخه در ادامه حفظ شده است. بخشهای دوزبانهٔ بالا قرارداد مستندسازی مشترک Shenava-1 را تعریف میکنند. |
|
|
| # Shenava — Rizeh v1.0 (32M) · CoreML iOS15 NeuralNetwork fp16 |
|
|
| CoreML **NeuralNetwork** (not ML Program) fp16 export of [`Reza2kn/Shenava-Rizeh-v1.0`](https://huggingface.co/Reza2kn/Shenava-Rizeh-v1.0) — built so **older Apple devices capped at iOS 15** (e.g. iPad Air 2 / iOS 15.8) can load and run it. ML Program packages require iOS 16+; this targets **NeuralNetwork / CoreML spec v5 with iOS 14 availability**, so it runs on iOS 15. |
|
|
| This is the **cache-aware streaming** step (one 170 ms prediction), the same kind of artifact as [`shenava-fa-fastconformer-streaming-32m-coreml-ios15-fp16`](https://huggingface.co/Reza2kn/shenava-fa-fastconformer-streaming-32m-coreml-ios15-fp16). |
|
|
| ## The Shenava-1 family (CoreML iOS15) |
|
|
| - [`Shenava-Koochik-1.0-CoreML-iOS15-fp16`](https://huggingface.co/Reza2kn/Shenava-Koochik-1.0-CoreML-iOS15-fp16) — **Koochik 1.0** (114M) · teacher / flagship |
| - [`Shenava-Rizeh-v1.0-CoreML-iOS15-fp16`](https://huggingface.co/Reza2kn/Shenava-Rizeh-v1.0-CoreML-iOS15-fp16) — **Rizeh v1.0** (32M) · mid-tier |
| - [`Shenava-Rizeh-Pizeh-v1.0-CoreML-iOS15-fp16`](https://huggingface.co/Reza2kn/Shenava-Rizeh-Pizeh-v1.0-CoreML-iOS15-fp16) — **Rizeh Pizeh v1.0** (6.9M) · tiniest |
|
|
| ## Benchmark — fair WER/CER (parent model, decoded @ `[70,13]`) |
|
|
| | Member | golden-6669 WER | CER | FLEURS-fa WER | CER | |
| |---|---|---|---|---| |
| | **Rizeh v1.0 (32M)** | **12.11%** | 3.94% | **14.45%** | 5.10% | |
|
|
| ## CoreML contract (cache-aware streaming CTC step, att_context `[70,0]`) |
| |
| Inputs: |
| - `processed_signal`: `Float32 [1, 80, 17]` |
| - `cache_last_channel`: `Float32 [16, 1, 70, 256]` |
| - `cache_last_time`: `Float32 [16, 1, 256, 8]` |
|
|
| Outputs: |
| - `logits`: `Float32 [1, 1, 1025]` |
| - `cache_last_channel_next`: `Float32 [16, 1, 70, 256]` |
| - `cache_last_time_next`: `Float32 [16, 1, 256, 8]` |
|
|
| Streaming geometry: feature_frames per prediction = **17** (pre_encode_cache 9 + chunk 8), audio window **170 ms**, constant cache length **70**, `d_model=256`, `16` conformer layers, ×8 subsampling (80 ms/frame). |
|
|
| ## Compatibility (Xcode `coremlc`) |
|
|
| - model type: `MLModelType_neuralNetwork` |
| - storage precision: `Float16` |
| - specification version: `5` |
| - availability: `iOS 14.0`, `macOS 11.0` |
|
|
| ```bash |
| coremlc compile shenava_rizeh_v1_0_ctc_streaming_att70_0_ios15_fp16.mlmodel /tmp/out --deployment-target 15.0 --platform ios |
| ``` |
|
|
| ## Files |
|
|
| - `shenava_rizeh_v1_0_ctc_streaming_att70_0_ios15_fp16.mlmodel` — fp16 NeuralNetwork model (~52 MB) |
| - `tokens.json`, `preprocessor.json`, `mel_filters_slaney_80x257.json` — sidecars (ve_tok_v4, shared across the family) |
| - `shenava_rizeh_v1_0_ctc_streaming_att70_0_ios15_fp16_manifest.json` — export manifest |
| - `export_koochik10_streaming_coreml.py` — reproducible export script |
|
|
| Tokenizer: ve_tok_v4 (SentencePiece BPE-1024 +blank, digit/punct/«»-aware). Numbers are emitted in **spoken form**; apply Persian ITN at display for digits. Part of [VisualEars / Shenava](https://shenava.app). |
|
|
| Export stack: coremltools 9.0, torch 2.7.0, NeMo 2.7.3. fp16 vs fp32 argmax agreement: **1.000**. |
|
|
|
|