this post was submitted on 26 Nov 2023
1 points (100.0% liked)

LocalLLaMA

3 readers
1 users here now

Community to discuss about Llama, the family of large language models created by Meta AI.

founded 1 year ago
MODERATORS
 

Let’s say you spend an unholy amount of processing time training a 70b. You like history. You want a good LLM for historical info.

By the time you upload it the LLM is outdated. Now what?

If you want it to speak accurately about modern events you’d have to retrain it again. Repeating the process over and over, because time keeps moving on while your LLM does not.

This clearly could become more efficient. Optimally, each subject would probably need to be considered a separate file while the central “brain” of the LLM becomes its own structure.

As it stands, updating the entire LLM is very cost prohibitive and makes no sense if you’re trying to work out specific data points. Why, for example, would you want to update the entire Cantonese dictionary when you just want to fix the list of Alaskan donut shops?

I understand that the tech currently has to treat both the information and the “thinking” behind an LLM as one and the same. It seems more efficient, more effective, to separate the two.

you are viewing a single comment's thread
view the rest of the comments
[–] Mission_Revolution94@alien.top 1 points 11 months ago

its really about the data curation and normalization.

think yi-34B they are getting results the same and better than 70B LLM's due

to the quality of there data.

work on the data and you will more than likely be happy with the results.

the training is really the fast part when you think of what is required to really nail down quality input.