35b variant?

#2
by dagbs - opened

Are there plans to release a 35B variant of this model?

I'd love to see a 35B too

same, I just tried this model on my toy project and it is really impressive! my fav so far by far! Thank you! Now i am wondering how better would 35 A3B be? will it be with only 3B active tho?

Tesslate org

Working writing the training pipeline for it right now. We were training on modal but since their gpus are usually spot instances, it keeps crashing us.

How does this model compare to Qwen-Coder-Next-80B MoE model in terms of real world usage?

Thank u smirki for the hard work , cant wait to see the end result of the 35b

How does this model compare to Qwen-Coder-Next-80B MoE model in terms of real world usage?

That would be a nice comparison to have!

9B is a dense model right? Wouldn't it be easier to train the 27B?

Upvote also for 35B variant because of cpu offloading makes it good for low vram. Also 4B would be huge for lapstops for small editing.

9B is a dense model right? Wouldn't it be easier to train the 27B?

but i can run 35b with 12g vram + 32g ram, not the 27b...

Sign up or log in to comment