Thanks for posting! Any chance for details of experiment, maybe some example code, so I can try on some different examples/task to gather more data? I have some tasks which, based on my previous testing, were not really solvable with just embeddings similarity, things like relevance of post/comment, prompt injection classification.