One thing to keep in mind is that, when it comes to LLMs, the models have not significantly changed in architecture.
There’s been new experiments and advancements in architecture on neural networks, and machine learning for specific applications. But LLM, as they are being commercialized by AI corporations to the general public, have stayed relatively the same. Except for one thing. Increasing in size. Larger datasets, or more specialized datasets like with coding, and larger number of tokens in memory. This is why it takes such large data centers. It’s all been just brute forcing greater capabilities by enlarging the models.
One of the things with LLM is that all the dataset influences the weighs and probabilities of the results. Even if the dataset includes a single event of a chain of words (think of the pizza with superglue incident), it can show up in the results eventually.
One thing to keep in mind is that, when it comes to LLMs, the models have not significantly changed in architecture.
There’s been new experiments and advancements in architecture on neural networks, and machine learning for specific applications. But LLM, as they are being commercialized by AI corporations to the general public, have stayed relatively the same. Except for one thing. Increasing in size. Larger datasets, or more specialized datasets like with coding, and larger number of tokens in memory. This is why it takes such large data centers. It’s all been just brute forcing greater capabilities by enlarging the models.
One of the things with LLM is that all the dataset influences the weighs and probabilities of the results. Even if the dataset includes a single event of a chain of words (think of the pizza with superglue incident), it can show up in the results eventually.