Mariano Phielipp
mphielipp
AI & ML interests
DRL, Multimodality, Continual Learning, Large Scale RL, Embody Intelligence, Robotics, Agents, Advanced Machine Intelligence.
Recent Activity
updated a collection 2 days ago
Robot Learning updated a collection 2 days ago
Robot Learning updated a collection 2 days ago
Visual Reasoning and LLMsOrganizations
None yet
Computer Vision
RL for Autoregressive Tasks
Real2Sim2Real
Light TTS models
Diffusion and RL
Visual Reasoning and LLMs
-
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
Paper • 2501.06186 • Published • 67 -
Planning with Reasoning using Vision Language World Model
Paper • 2509.02722 • Published • 24 -
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning
Paper • 2607.01191 • Published • 19
Robot Learning
-
Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers
Paper • 2403.12943 • Published • 15 -
Rolling Diffusion Models
Paper • 2402.09470 • Published • 13 -
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
Paper • 2607.00678 • Published • 20 -
ASPIRE: Agentic /Skills Discovery for Robotics
Paper • 2607.00272 • Published • 24
SSMs and Diffusion
Self Pedicting Learning in RL
CV
Flow Matching
Agentic RL
CUDA Optimization
LLM Training
-
LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Paper • 2403.13372 • Published • 186 -
On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
Paper • 2508.05629 • Published • 190 -
Experiential Reinforcement Learning
Paper • 2602.13949 • Published • 76
Datasets for Robotic Learning
VLM
Diffusion Transformers
Conditional Diffusion
Grokking
LLMs Evaluation
VLA
-
OpenVLA: An Open-Source Vision-Language-Action Model
Paper • 2406.09246 • Published • 47 -
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
Paper • 2411.19650 • Published -
Octo: An Open-Source Generalist Robot Policy
Paper • 2405.12213 • Published • 29 -
Diffusion-VLA: Scaling Robot Foundation Models via Unified Diffusion and Autoregression
Paper • 2412.03293 • Published
Robot Learning from Physical World Model
Flow Matching
Computer Vision
Agentic RL
RL for Autoregressive Tasks
CUDA Optimization
Real2Sim2Real
LLM Training
-
LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models
Paper • 2403.13372 • Published • 186 -
On the Generalization of SFT: A Reinforcement Learning Perspective with Reward Rectification
Paper • 2508.05629 • Published • 190 -
Experiential Reinforcement Learning
Paper • 2602.13949 • Published • 76
Light TTS models
Datasets for Robotic Learning
Diffusion and RL
VLM
Visual Reasoning and LLMs
-
LlamaV-o1: Rethinking Step-by-step Visual Reasoning in LLMs
Paper • 2501.06186 • Published • 67 -
Planning with Reasoning using Vision Language World Model
Paper • 2509.02722 • Published • 24 -
Perceive-to-Reason: Decoupling Perception and Reasoning for Fine-Grained Visual Reasoning
Paper • 2607.01191 • Published • 19
Diffusion Transformers
Robot Learning
-
Vid2Robot: End-to-end Video-conditioned Policy Learning with Cross-Attention Transformers
Paper • 2403.12943 • Published • 15 -
Rolling Diffusion Models
Paper • 2402.09470 • Published • 13 -
ABot-M0.5: Unified Mobility-and-Manipulation World Action Model
Paper • 2607.00678 • Published • 20 -
ASPIRE: Agentic /Skills Discovery for Robotics
Paper • 2607.00272 • Published • 24
Conditional Diffusion
SSMs and Diffusion
Grokking
Self Pedicting Learning in RL
LLMs Evaluation
CV
VLA
-
OpenVLA: An Open-Source Vision-Language-Action Model
Paper • 2406.09246 • Published • 47 -
CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
Paper • 2411.19650 • Published -
Octo: An Open-Source Generalist Robot Policy
Paper • 2405.12213 • Published • 29 -
Diffusion-VLA: Scaling Robot Foundation Models via Unified Diffusion and Autoregression
Paper • 2412.03293 • Published