AI & ML interests

A public archive of decoder-only models where anyone can archive their own models.

Recent Activity

Banaxi-Tech 
posted an update about 6 hours ago
view post
Post
659
Well GPT X3 takes the lead. As of right now.




We are now announcing BananaMind 3 🍌! (not ai for those emoji guys)

All models will use BGA (which is almost just NSA) and our BM3X architecture.
Its sizes will be:
BananaMind 3 Flash Lite, 3M parameters at a context of 8K context.
BananaMind 3 Lite, 10M parameters with 16K context.
BananaMind 3 Flash, 25M Parameters with 16K context.
BananaMind 3 Pro, 50M parameters with 24K context.
BananaMind 3 Ultra, 100M parameters with 32K context.
And lastly, BananaMind 3 Max with 150M parameters and 64K CONTEXT.

I can assure you BananaMind 3 Max WILL beat GPT X3 or match it, we won't release it otherwise. We hope for a 40+ INTELLIGENCE INDEX!

BananaMind 3 may also be partnered with dot labs.

We will cancel BananaMind 2.1 and BananaMind 2 Ultra.

As of the BETU SLM Leaderboard we may need to release it after October 11, im very busy right now (even though we said We will release BETU leaderboard before Oct 11 😟)


  • 3 replies
·
Banaxi-Tech 
posted an update 1 day ago
view post
Post
1813
We have released BGA!
And wow, It provides 256x (and 512x at the end of 1M) yes 256x LESS attention compute at 1M context window.
That means you can train a 1M context window at the compute of a ~4K context window.

Check IT OUT: BananaMind/blog

The Accuracy Should BE WAy better than DSA but untested yet.


And, now some updates on BananaMind 3:

BananaMind 3 Will start training Soon!
Sizes: 10M, 25M, 50M, 100M, 150M

And the context windows ARE INSANE: 10M, 16K context, 25M 16k context, 50M 32K context, 100M and 150M, 64K context!!!!

  • 8 replies
·
Banaxi-Tech 
posted an update 2 days ago
view post
Post
123
Hi! We're right now experimenting with so many new architectures!

We are also going to release BananaMind Gate Attention very soon, its a new attention that can make attention 128x cheaper at 1M Context! Check my account for HF blogs it will release there!
Banaxi-Tech 
posted an update 3 days ago
view post
Post
80
We're introducing ACR 1.0.
We trained this model on a 5070 Ti for weeks, here are some of the architecture details:
57M parameters, with one M and one G stream.
When we tested it on benchmarks, we got these results:
Benchmark Full G-Only Delta
PIQA 62.24% 53.43% +8.81
ARC-Easy 41.96% 32.28% +9.68
HellaSwag 33.19% 29.08% +4.11
Tiny ToM 40.65% 33.75% +6.90
ArithMark 3.0 33.40% 32.80% +0.60
Base Bench 1.1 51.71% 40.29% +11.42

Check it out at saicr/ACR-1.0
  • 3 replies
·
DedeProGames 
posted an update 4 days ago
view post
Post
3982
Im working on a 23M ASR model, trained on 100k hours of audio
  • 4 replies
·
Banaxi-Tech 
posted an update 4 days ago
view post
Post
162
Its SAICR time tomorrow.
Get ready!
saicr
  • 28 replies
·
Banaxi-Tech 
posted an update 5 days ago
view post
Post
3059
We will release the BEST SLM Leaderboard before Oct 11.

It will feature everything:
Easy to use model picker.
EXTREMELY Easy way to add your own models (2 click)
MULTIPLE leaderboard for different model types

And more!


So why don't you help us build it?

Join
betu-slm-leaderboard-testers
  • 3 replies
·
DedeProGames 
posted an update 8 days ago
view post
Post
5820
how is this possible
  • 13 replies
·
Banaxi-Tech 
posted an update 9 days ago
view post
Post
5160
ACR 1.0 launch is being prepared and researched now!
Also I'm going to vacation tomorrow but it should still be released!

saicr
  • 3 replies
·