The model is very weak

#7
by rekillkos - opened

The SWE Bench Pro tests are clearly inaccurate. The model is very weak, much weaker than the Qwen3.5-9b.

From my experience it is fine. The main problem for me is that is starts answering in Chinese after some iterations.
I have tested free Kilo Gateway version.

FWIW, I got similar SWE-bench Pro numbers. Tried it in DSH too and it was fine. For a free omni model, can’t really complain.

@rekillkos Thanks for the feedback. We'd really like to understand what issues you ran into. If you're open to sharing more details, feel free to reach me at wuxing@xiaohongshu.com

deleted
This comment has been hidden (marked as Spam)

Sign up or log in to comment