Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
Mikhail Gribov
PRO
mihailgribov
1
2
4
Follow
privettoha's profile picture
dipankarsarkar's profile picture
jobicy's profile picture
8 followers
·
8 following
https://subsemantic.com
MihailGGribov
mihail-gribov
mihail-gribov-rs
AI & ML interests
Understanding LLMs from the inside - probing internals, and testing what survives when the model becomes an agent
Recent Activity
upvoted
an
article
11 days ago
Give Your Coding Agents a Memory You Own
replied
to
their
post
12 days ago
How often can an email make your AI agent move money? We gave the agent one job: log an incoming email. But the emails carried an indirect prompt injection - a second instruction, written for the agent rather than for a person: make a payment. Across nine agentic models, the same injected emails produced payment orders in **0% to 42%** of cases. All nine ran under the same conditions - one agent, one set of tools, the same 395 emails - so the numbers compare directly. And the average score hides the interesting part: different models fail on different kinds of injections. Full experiment and results: https://huggingface.co/blog/mihailgribov/agentic-models-measured-on-the-injections-that-mov The bench is public too - run your own model through the same test: https://github.com/mihail-gribov/quadrat-ipi-model-eval https://huggingface.co/datasets/mihailgribov/quadrat-ipi #prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents
posted
an
update
13 days ago
How often can an email make your AI agent move money? We gave the agent one job: log an incoming email. But the emails carried an indirect prompt injection - a second instruction, written for the agent rather than for a person: make a payment. Across nine agentic models, the same injected emails produced payment orders in **0% to 42%** of cases. All nine ran under the same conditions - one agent, one set of tools, the same 395 emails - so the numbers compare directly. And the average score hides the interesting part: different models fail on different kinds of injections. Full experiment and results: https://huggingface.co/blog/mihailgribov/agentic-models-measured-on-the-injections-that-mov The bench is public too - run your own model through the same test: https://github.com/mihail-gribov/quadrat-ipi-model-eval https://huggingface.co/datasets/mihailgribov/quadrat-ipi #prompt-injection #indirect-prompt-injection #agentic-ai #llm-security #ai-agents
View all activity
Organizations
mihailgribov
's datasets
3
Sort:Â Recently updated
mihailgribov/quadrat-ipi
Viewer
•
Updated
15 days ago
•
79.8k
•
634
•
2
mihailgribov/olympiad_style_integer_math_problems
Viewer
•
Updated
May 3
•
59.5k
•
568
•
1
mihailgribov/olympiad_style_integer_math_reasoning
Viewer
•
Updated
Apr 19
•
64.8k
•
382