LocalLLaMA

11 readers

4 users here now

Community to discuss about Llama, the family of large language models created by Meta AI.

founded 2 years ago

MODERATORS

communick@poweruser.forum

Why is Mistral-7b so capable? Any ideas re: dataset? (alien.top)

submitted 2 years ago by Fun_Tangerine_1086@alien.top to c/localllama@poweruser.forum

24 comments fedilink hide all child comments

So Mistral-7b is a pretty impressive 7B param model ... but why is it so capable? Do we have any insights into its dataset? Was it trained very far beyond the scaling limit? Any attempts at open reproductions or merges to scale up # of params?

you are viewing a single comment's thread
view the rest of the comments

[–] Feztopia@alien.top 1 points 2 years ago

As far as I know (I might be wrong) it's partly the team that made llama1 (and maybe made the first steps for llama2?). So they already knew what they were doing. How llama could be improved* and so on.

*The dataset