Nice Jev benchmark results! I ran an independent evaluation of TypeSafe's Jev against 10+ LLMs (GPT-4o, Claude, Gemini, Llama...) with open-source reproducible code. Some interesting findings on where the benchmark's methodology holds up and where it doesn't.
Article: https://medium.com/@pravvich/typesafes-jev-beyond-the-hype-an-independent-benchmark-8bdc1c99d000
Code: https://github.com/PavelRavvich/jev-bench