Your assumption in that LLMs are "just prediction engines" rather than having true intelligence is something I just cannot help to not comment here on "hugging face social media".
I hate to say that you might not know what you are talking about but the way you talk in high technical sense might resist you from actually looking into things seriously, so to the highest concern of my ego I strongly encourage you to take a deep dive into linear and non linear systems, transfer functions, and the math behind why LSTM deep neural networks works so well. What do you think your brain really is?
After that, this is when you should consider and read attention is all your need a few times to grasp why the KQV cache layer matters and how this performed so well while avoiding a lot of the issues from LSTM. Once you are here. I hope you may be a bit more skeptical, where you actually can have sufficent knowlege to say: These large networks might've abstracted critical thinking, though very different than us, but hold all hallmark of intelligence.