-
OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution
Paper • 2609.06490 • Published • 7 -
dipta007/OracleZoom
Image-to-Image • Updated • 2 -
dipta007/OracleZoom-4KLSDB-train
Viewer • Updated • 52.8k • 622 -
OracleZoom
🔎2Zoom any photo to 256x, one 4x step at a time
Shubhashis Roy Dipta PRO
dipta007
AI & ML interests
Multimodal Understanding, Reasoning, Generation
Recent Activity
new activity about 2 hours ago
JonesLin/writing-model-papers-2016-2021:add OCR markdown of the PDF corpus as parquet shards submitted a paper 1 day ago
OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution published a Space 1 day ago
dipta007/OracleZoomOrganizations
DecomposeRL
GanitLLM (ACL 2026 Findings)
-
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
Paper • 2601.06767 • Published • 1 -
dipta007/Ganit
Viewer • Updated • 32.3k • 154 -
dipta007/GanitLLM-4B_SFT_CGRPO
Text Generation • 196k • Updated • 72 -
dipta007/GanitLLM-4B_SFT_GRPO
Text Generation • 196k • Updated • 12 • 1
Q2E (AACL 2025 Main)
Datasets used in the paper: Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval
-
Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval
Paper • 2506.10202 • Published -
dipta007/Q2E_MultiVENT_LLAMA_3.3_70B_InternVL_38B_Funiform_16_noASR
Viewer • Updated • 2.39k • 10 -
dipta007/Q2E_MultiVENT_LLAMA_3.3_70B_InternVL_38B_Funiform_16_ASR
Viewer • Updated • 2.39k • 11 -
dipta007/Q2E_MSRVTT-1kA_LLAMA_3.3_70B_InternVL_38B_Funiform_16_ASR
Viewer • Updated • 1k • 9
DAGGER (EMNLP 2026 Findings)
A graph based CoT for math reasoning (DAGGER tokens <<<< CoT Tokens)
-
†DAGGER: Distractor-Aware Graph Generation for Executable Reasoning in Math Problems
Paper • 2601.06853 • Published • 1 -
dipta007/DistractMath-Bn
Viewer • Updated • 3.69k • 64 -
dipta007/dagger
Viewer • Updated • 6.48k • 68 -
dipta007/dagger-12B_SFT_GRPO
Text Generation • 12B • Updated • 708 • 1
VC-Inspector (ACL 2026 Main)
-
Advancing Reference-free Evaluation of Video Captions with Factual Analysis
Paper • 2509.16538 • Published • 2 -
dipta007/VCInspector-7B
Image-Text-to-Text • 8B • Updated • 64 • 1 -
dipta007/VCInspector-3B
Image-Text-to-Text • 4B • Updated • 32 • 1 -
dipta007/ActivityNet-FG-It
Viewer • Updated • 242k • 25
BIRD-Synthetic
Synthetically rephrased version of BIRD dataset
OracleZoom
-
OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution
Paper • 2609.06490 • Published • 7 -
dipta007/OracleZoom
Image-to-Image • Updated • 2 -
dipta007/OracleZoom-4KLSDB-train
Viewer • Updated • 52.8k • 622 - Running on ZeroAgents2
OracleZoom
🔎2Zoom any photo to 256x, one 4x step at a time
DAGGER (EMNLP 2026 Findings)
A graph based CoT for math reasoning (DAGGER tokens <<<< CoT Tokens)
-
†DAGGER: Distractor-Aware Graph Generation for Executable Reasoning in Math Problems
Paper • 2601.06853 • Published • 1 -
dipta007/DistractMath-Bn
Viewer • Updated • 3.69k • 64 -
dipta007/dagger
Viewer • Updated • 6.48k • 68 -
dipta007/dagger-12B_SFT_GRPO
Text Generation • 12B • Updated • 708 • 1
DecomposeRL
VC-Inspector (ACL 2026 Main)
-
Advancing Reference-free Evaluation of Video Captions with Factual Analysis
Paper • 2509.16538 • Published • 2 -
dipta007/VCInspector-7B
Image-Text-to-Text • 8B • Updated • 64 • 1 -
dipta007/VCInspector-3B
Image-Text-to-Text • 4B • Updated • 32 • 1 -
dipta007/ActivityNet-FG-It
Viewer • Updated • 242k • 25
GanitLLM (ACL 2026 Findings)
-
GanitLLM: Difficulty-Aware Bengali Mathematical Reasoning through Curriculum-GRPO
Paper • 2601.06767 • Published • 1 -
dipta007/Ganit
Viewer • Updated • 32.3k • 154 -
dipta007/GanitLLM-4B_SFT_CGRPO
Text Generation • 196k • Updated • 72 -
dipta007/GanitLLM-4B_SFT_GRPO
Text Generation • 196k • Updated • 12 • 1
BIRD-Synthetic
Synthetically rephrased version of BIRD dataset
Q2E (AACL 2025 Main)
Datasets used in the paper: Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval
-
Q2E: Query-to-Event Decomposition for Zero-Shot Multilingual Text-to-Video Retrieval
Paper • 2506.10202 • Published -
dipta007/Q2E_MultiVENT_LLAMA_3.3_70B_InternVL_38B_Funiform_16_noASR
Viewer • Updated • 2.39k • 10 -
dipta007/Q2E_MultiVENT_LLAMA_3.3_70B_InternVL_38B_Funiform_16_ASR
Viewer • Updated • 2.39k • 11 -
dipta007/Q2E_MSRVTT-1kA_LLAMA_3.3_70B_InternVL_38B_Funiform_16_ASR
Viewer • Updated • 1k • 9