You are viewing a single thread.
View all comments View context
2 points
*

I regularly run llama3 70b unqantized on two P40s and CPU at like 7tokens/s. It’s usable but not very fast.

permalink
report
parent
reply
1 point

so there is no way a 24gb and 64gb can run thing?

permalink
report
parent
reply
2 points
*

My specs because you asked:

CPU: Intel(R) Xeon(R) E5-2699 v3 (72) @ 3.60 GHz
GPU 1: NVIDIA Tesla P40 [Discrete]
GPU 2: NVIDIA Tesla P40 [Discrete]
GPU 3: Matrox Electronics Systems Ltd. MGA G200EH
Memory: 66.75 GiB / 251.75 GiB (27%)
Swap: 75.50 MiB / 40.00 GiB (0%)
permalink
report
parent
reply
1 point

ok this is a server. 48gb cards and 67gb ram? for model alone?

permalink
report
parent
reply
1 point

What are you asking exactly?

What do you want to run? I assume you have a 24GB GPU and 64GB host RAM?

permalink
report
parent
reply
1 point

correct. and how ram speed work in this tbh

permalink
report
parent
reply

Technology

!technology@lemmy.world

Create post

This is a most excellent place for technology news and articles.


Our Rules


  1. Follow the lemmy.world rules.
  2. Only tech related content.
  3. Be excellent to each another!
  4. Mod approved content bots can post up to 10 articles per day.
  5. Threads asking for personal tech support may be deleted.
  6. Politics threads may be removed.
  7. No memes allowed as posts, OK to post as comments.
  8. Only approved bots from the list below, to ask if your bot can be added please contact us.
  9. Check for duplicates before posting, duplicates may be removed

Approved Bots


Community stats

  • 17K

    Monthly active users

  • 6K

    Posts

  • 128K

    Comments