LocalLLaMA

14 readers

1 users here now

Community to discuss about Llama, the family of large language models created by Meta AI.

founded 2 years ago

MODERATORS

communick@poweruser.forum

What kind of specs to run local llm and serve to say up to 20-50 users (alien.top)

submitted 2 years ago by Appropriate-Tax-9585@alien.top to c/localllama@poweruser.forum

10 comments fedilink hide all child comments

Hi all,

Just curious if anybody knows the power required to make a llama server which can serve multiple users at once.

Any discussion is welcome:)

you are viewing a single comment's thread
view the rest of the comments

[–] Aggressive-Drama-899@alien.top 1 points 2 years ago (1 children)

We run llama 2 70b for around 20-30 active users using TGI and 4xA100 80gb on Kubernetes. If 2 users send a request at the exact same time, there is about a 3-4 second delay for the second user. Never really had any complaints around speed from people as of yet. We do have the ability to spin up multiple new containers if it became a problem though. This is all on prem

[–] Appropriate-Tax-9585@alien.top 1 points 2 years ago

Thank you, this is really good to hear!